Someone finally put a price on the frontier. A Mozilla report previewed this week measures the capability gap between the best closed models and the best open-weights models at 4.4 months — and the premium for closing that gap at roughly five times the cost per task. For a one-person business, that is the most useful number of the week, because it turns “which AI should I use” from a brand question into a per-task decision you can actually make.

The premium is real, narrow, and priced

Ars Technica previewed Mozilla’s State of Open Source AI report on 15 September. Two figures matter. The performance gap is 4.4 months. And on a neutral test harness, an open-weights model scored within a point of a top closed model at about five times less per completed task.

Mozilla CTO Raffi Krikorian’s framing is the part to keep: the decision to pay for a closed model is “workload-specific rather than organization-specific.” Pay the premium when a deadline lands before the cheaper option catches up. Do not pay it for “routine work you’ll still be doing next quarter, because you’ll be able to do it for a fifth of the cost soon, and the model won’t be the bottleneck anyway.”

Read that last clause twice. The model won’t be the bottleneck anyway. In a one-person business it almost never is. The bottleneck is that nobody wrote down what a good output looks like.

Most of what you automate is a decision, not an essay

The week’s clearest illustration came from a launch. Latent Space covered TypeSafe’s Jev, a model that only decides, classifies, routes and scores — it requires predefined output formats and cannot produce free-form text. The claimed advantage over small frontier models is >100x faster and >200x cheaper.

Set that against your own workflow. Which intake email is a real lead. Which support message is urgent. Which invoice is overdue. Which draft belongs to which client. Which of the fifty things in your inbox needs you today. Those are classifications, and most solo operators are paying a general-purpose conversational model to make every one of them, at conversational prices, on every single item, forever.

This is The Delegation Archetype Map applied to cost rather than quality. The map’s question is what role AI is playing in a task — is it deciding, drafting, checking, or advising? The billing consequence has been quietly sitting underneath that question the whole time: a decision step and a drafting step do not need the same class of model, and the decision steps are the ones you run hundreds of times. Sort your recurring AI steps into those two piles this month. You do not need a new tool to benefit — you need to stop sending a classification to the model you chose for your hardest writing task.

Your prospect’s first contact may now be someone else’s agent

OpenAI shipped a set of advertising changes on 16 September. It is testing Sponsored Agents with select US advertisers: after clicking an ad in ChatGPT, a user can start a clearly labelled conversation with a business-sponsored agent, separate from ChatGPT’s own answers. Advertisers can now create and analyse campaigns through natural-language prompts, get suggested copy and imagery from their landing page, and opt into AI-powered text customisation that adapts headlines to the context of a conversation. ChatGPT Ads also arrived inside HubSpot and Shopify, with the Shopify app going international on 23 September.

For a solo business, the interesting half is not the ad tools. It is that a buyer’s first substantive interaction with a category can now be a conversation an advertiser configured. That changes what your own expertise has to do to register.

The Trust Visibility Calculus asks how visible AI in your delivery changes what a client believes about your competence. This week extends the question upstream, to acquisition: if your prospect has already had a fluent, patient, well-informed conversation with a competitor’s sponsored agent before meeting you, fluency is no longer evidence of anything. The things that still separate you are specificity, judgment about cases the agent cannot have seen, and a point of view someone signed. Those were always the differentiators. They are now the only ones left.

Big companies are getting dashboards. You need one line.

OpenAI also published how to connect AI usage to business value — admin analytics that group AI usage into tasks, show which model settings are eating credits, and track whether AI-assisted code reaches merged commits. One suggestion in it is worth borrowing outright: a routine brief “may be worth testing with a faster or lower-cost setup, comparing quality and the time spent reviewing and correcting it.”

That is the whole measurement, and it costs nothing. Time spent reviewing and correcting is the number that tells you whether an AI step is helping. You will not get a console, so keep the crude version: for your three most-repeated AI tasks, note how many minutes of your own editing each output needs. A step whose correction time is not falling is not compounding — it is a fixed cost you never chose.

And reliability became something you can buy

AIUC raised a $40M Series A for what it calls confidence infrastructure: AIUC-1, a standard for agent security, safety and reliability, with 51 requirements and 130 controls across hallucinations, jailbreaks and data leakage, re-tested quarterly, paired with insurance underwritten through Lloyd’s of London.

You are not buying agent insurance. But notice what an insurer requires before it will price a risk: a written standard, and evidence the standard was met, re-checked on a schedule. The Reliance Calibration Dial asks you to set trust from what an AI output actually does, not from how confident it sounds. An insurance market is that dial with money behind it — and the fact that one now exists is a strong signal about which half of the AI conversation is maturing.

The throughline

Four separate stories, one shape: the industry is learning to price and verify AI instead of admiring it. A 4.4-month head start with a five-times premium. A model that only makes decisions, cheaply. A dashboard connecting spend to outcomes. A standard with an insurer behind it.

Your version needs no platform. Know which of your AI steps are decisions and which are drafts, know what each one costs you in correction time, and write down what a correct output looks like so the question can be answered at all. Rented fluency keeps getting cheaper. The part that compounds is the system you own around it.

Start with the audit, not the tooling: curiochat.ai/solopreneur/.