Skip to content

White papers

Shaped by evidence September 3, 2026 (rev C)

The One-Person AI Operating System

A human-led way of working that lets a small online business learn without swallowing its owner

Pierre Boutquin

For solo creators, teachers, consultants, coaches, course builders, and small software founders who sell online.

Short form: The proposition and the five-minute reading path

First published
September 1, 2026
Design status
Shaped by evidence, still being tested. It will change as research and product results improve.
Conformance
Draft target. No Curio Chat Academy release should claim it matches this paper until two things are true. It names a fixed version of the paper. And it publishes test results, requirement by requirement, that anyone can repeat.
Reading level
Written to Flesch–Kincaid Grade 6.0–7.5 (Reading Ease 65–75), measured on prose only with textstat.

The proposition

A business learns when past judgment governs the next task, and new evidence changes that judgment.

The reader this paper is written for already runs a solo practice with AI. They built the custom GPT (or wrote a few skills). They wrote the standing instructions. They organized the second brain. They keep a prompt file named something like "v7 FINAL(2)." They paid for a pack, a course, or a memory tool. It promised to hold on to their corrections. They did not fail to try. They watched the gain fade and the same correction come back.

That failed build is where this paper starts. Its reader needs no proof that AI forgets. They want to know why a system that remembers more still fails. It repeats January's tone error in March. It brings back an old price. It forces one more full reread before anything goes to a client.

If your first thought is "I tried that," the rest of this paper rests on one difference. A prompt can ask. Memory can inform. But to refuse a draft that breaks a standing correction, you need a governed way of working — one that can hand the draft back.

The owner is still the only part of the business that can say no. In the failed setup, nothing stands between a confident, mostly right draft and the client's inbox except the owner's attention. More prompting will not change that. The setup has no lasting place for a correction to govern later work, and no checkpoint where a draft can fail against it.

The model is not the operating system. It is one guessing part inside a larger system. That system also holds the harness, business knowledge, source evidence, workflow state, rules, test rules, permissions, and learning records. What makes work reliable is not one huge prompt. It is keeping those parts apart, and giving each one its own power and its own life. Then you write down how they met on a real piece of work.

That split also changes the buy-or-build choice. The owner can buy the hard, reusable machinery that governs the work. They do not have to sell the thing the business is actually building. That is its grasp of the customer, its offer, its voice, its evidence, its decisions, and its earned judgment. Curio Chat Academy's intended line keeps that business-specific asset in plain files the owner holds — outside the model, and outside any private app database. A model or a working layer can be swapped out while the built-up business stays.

A solo owner does not need a pretend company staffed by AI employees. They need a small operating system that does seven things well:

  1. remembers what is true about the business
  2. names the constraint that matters now
  3. turns the owner's intent into a brief you can test
  4. gives AI bounded work it can do well
  5. checks output in step with its risk
  6. puts useful work in front of real customers.
  7. turns results and corrections into better work next time.

The goal is a smaller approval pile, with the final yes still human. Routine candidates arrive with the business truth that applies, the checks that apply, and a verdict. The owner then reads the exceptions instead of rebuilding and rereading everything.

The Curio Chat Academy manifesto names the goal: build a business that learns. Its enemy is business amnesia — losing customer insight, hard-won judgment, working corrections, and evidence over and over. This paper turns that idea into a working mechanism. Remembering is needed, but it is not enough. The business must keep the right judgment and put it where work can be held to it. It must test that judgment against reality and change it when it fails. Each cycle then starts from what the last cycle earned.

No science shows that one AI workflow makes solo owners succeed. The evidence does show several narrower things. AI can lift speed and quality on some tasks. The gains swing sharply by task and by user. Work can get worse past the edge of what the model does well. Advice can hurt when the user cannot judge what to act on. And a human plus AI is not always better than the stronger of the two alone.

This paper offers a working model shaped by that evidence. It stops well short of calling it a proven recipe for success.

Five-minute reading path

If you read only seven sections, read these:

  • Buy the shared layer; build the edge: why buying the common layer is a sound default, where it breaks, and what Curio Chat Academy must prove.
  • The model is not the operating system: why model, harness, knowledge, sources, workflow, rules, test rules, permissions, and learning need clear lines between them.
  • Moving and running many models are conformance properties: how one master owner-held system serves many machines, apps, and model vendors without splitting the business's truth.
  • The seven-stage workflow: what to do on a real piece of work.
  • What must stay human: where handing off stops.
  • How to test the design: what to compare, what to measure, and what to change when a line does not earn its cost.
  • What the evidence supports: which claims are strong and which are still open.

Contents

  1. The claim nobody should make
  2. What success means for a one-person business
  3. The business must come before the workflow
  4. The model is not the operating system
  5. The seven-stage workflow
  6. What must stay human
  7. Use AI as a menu of bounded skills
  8. The work rhythm
  9. How to test the design
  10. Worked example: a customer-facing offer test
  11. The scoreboard
  12. Failure modes and fixes
  13. The hand-off contract and escalation path
  14. What the evidence supports
  15. Conclusion

Appendices:

  • Appendix A: Public design record list
  • Appendix B: Evidence register
  • Appendix C: Method and claim limits

1. The claim nobody should make

"Install these AI agents and they will run a winning business for you."

No sound evidence backs that claim.

An online business wins when it makes something a set customer values. It must reach that customer cheaply, earn trust, and turn demand into sales. Then it must deliver what it promised and keep enough value to go on. AI can help with many tasks inside that system. It cannot make weak demand strong. It cannot make a bland offer needed. It cannot turn a misleading claim into trust. And it cannot decide what kind of owner a person wants to be.

The closest direct evidence is a warning. A 2026 peer-reviewed field trial worked with 640 Kenyan business owners. Access to a GPT-4 business adviser was assigned at random. The researchers could not rule out a zero average effect on revenue and profit. In subgroup results, weaker performers did about 8% worse, while stronger performers may have gained just over 15%. The gap seemed to lie in which advice people chose to act on. They asked much the same questions and got much the same advice. The result belongs to its setting, but the lesson travels: getting advice is not the same as having the judgment to use it. Otis et al., Management Science (2026)

Handing over as much as possible is a poor goal for that system. A better split of thinking has three parts:

  • the human brings purpose, customer insight, priorities, taste, accountability, and the final call
  • AI brings search, reshaping, comparison, drafting, role-play, and steady work inside a set line
  • the market decides whether the work together deserves to go on.

2. What success means for a one-person business

"Successful" is too vague to work with. In this paper, a winning online business has four traits:

  1. Customer value: the business makes a result that customers can see and care about.
  2. Sound money: revenue, delivery cost, the cost to win a customer, refunds, repeat business, and owner time all support the business continuing.
  3. Repeatability: the business can find, win, and serve customers again and again without reinventing itself every week.
  4. Owner control: the owner can still understand, steer, question, and change the system.

Revenue with no customer value is extraction. Revenue with unsound money is a countdown. Repeatability with no control can leave the owner tied to a machine they no longer understand. Control with no market contact stays private craft, short of a business.

AI belongs inside this meaning of success. It does not get to redefine it.

Who this system is for

An AI operating system cannot be judged apart from how mature the person and the business are.

This system helps most an AI-experienced, expert-led solo owner with a real offer or customer workflow. They have already tried prompts, standing instructions, memory features, or a second brain. They still find that corrections fade and the final call stays stuck with them.

That line matters for four reasons:

  1. The system needs source material. A working business has customer language, offers, delivery records, decisions, near misses, and real drafts. An idea with no revenue has mostly guesses. Memory that lasts can compound either one. Much of what a seasoned owner knows will not be in that material. A document is the residue a decision leaves behind. It rarely keeps the reasoning that made it. Knowledge in that state has to be drawn out, not gathered — a separate job covered in Section 7 and named as a gap in Appendix C.
  2. The system needs judgment to encode. AI can extend an owner's standards. It cannot supply mature standards the owner has not yet formed.
  3. The system needs a loop it can measure. Existing work gives you a before-state, repeat failures, and customer or money results you can see. Without them, "better" is only a feeling.
  4. Risk shifts with the owner's skill. Field evidence on business owners suggests AI advice can help and harm different owners unevenly. A paper written for everyone would hide the very conditions that decide whether advice becomes an edge or a mistake.

Someone still shaping an offer should start with customer contact and small market tests. The fuller operating system pays off once there is something real to remember, govern, and improve.

The reader does not need to become an engineer. Editing a document and running a guided command are enough hands-on skill for a supported setup. Learning software engineering, or taking up a new technical hobby, is not the job.

The setup does need controlled access to files the owner holds, or to a governed data store much like it. A browser-only chat window that cannot read and write the owner's own governed records is the wrong tool for the design here. It can explain the ideas. It cannot run the file-based learning loop. That limit should be clear before anyone spends money on a build.

Buy the shared layer; build the edge

For the reader this paper is written for, buying a fit-for-purpose base layer is the sound place to start. Buy the shared plumbing that sets no one apart. Build the truth and the edge that belong to your business alone.

The One-Person AI Operating System covers both layers. A vendor can supply the shared layer. That means the store contracts, routing, test rules, gates, learning parts, and life-stage machinery listed in the table below. The owner's business vault is the layer that sets them apart. It holds customers, offers, voice, evidence, money, decisions, identity lines, and the situated learning no other buyer gets. The line between the two is a design line. The system exists to make that asset usable without swallowing it.

The core business vault is stored as plain Markdown. Learning records use plain text formats you can read, where structure is needed. The owner can read them, track versions, search them with common tools, and carry them into another file-capable AI setup. Swapping the model therefore replaces only a guessing part. The business's memory stays where it was. Leaving Curio Chat Academy does not mean leaving your built-up truth behind. Plain formats cannot promise an easy move. They cannot copy Curio Chat Academy's behavior elsewhere either. They do cut the dependence that buying creates.

The split between stores makes the ownership line concrete:

The owner builds, owns, and keeps Curio Chat Academy supplies and maintains
Source material and curated business truth Store contracts, templates, addressing, and the parts that pick context
Customer language, offer, voice, evidence, money, and decisions General craft base and shared workflow scaffolding
Working corrections, owner-specific expertise, results, and learning history Rule routing, locked test rules, paper trail, gates, and restart parts
The final call, and the power to approve high-stakes action Install, checking, upgrades, compatibility work, and safe removal

Curio Chat Academy supplies the parts that put the records to work, and takes none of them. The same split that makes work more reliable is what lets the owner switch models. It also lets them stop using the system without handing over the asset it helped organize.

That default is advice, not a measured fact. No study assigns solo owners to buy or build an AI system and then compares results. The advice rests on a narrower body of work: young firms, make-or-buy choices, software reuse, small-business buying, and owner time.

The closest evidence on firm life comes from a panel of about 240,000 U.S. factory plant-year records from 2006 to 2014. About 41,300 of them came from plants five years old or younger. Jin and McElheran found that young plants gained more than others by renting modern IT. Rented IT helped early output and survival. Building up owned IT went with a higher chance of shutting down in the fragile early years. They also found that owned IT could build firm-specific strengths over time. The study covers factory plants. It says nothing directly about solo owners or AI systems. It does support the core order: rent flexible, widely sold plumbing before making big ownership bets. Then bring in-house what turns strategic and worth keeping. Jin and McElheran, Management Science (2025)

Wider make-or-buy research does not say "always buy." A meta-analysis of make, buy, and ally choices found strong support for matching the deal to the setup. Deals with rare, custom assets favored in-house control. So did doubt about how a partner would act, or how much would be needed. Doubt about the technology itself favored buying on the market. Software studies name the same practical factors. Interviews with eight Thai small firms found seven factors that shaped the buying call. Fit to needs, cost, size and complexity, time, in-house skill, support, and daily practicalities. Geyskens, Steenkamp, and Kumar, Academy of Management Journal (2006); Daneshgar, Low, and Worasinchai, Information and Software Technology (2013)

The reuse evidence strengthens the case for starting from a layer someone else maintains. A 2024 review of 30 industry studies judged the evidence strong for better quality and higher output from software reuse. The cost side was studied too thinly to support a general payback claim. Only three of the studies looked at small companies. Reuse supports a direction. Any Curio Chat Academy savings figure would have to come from somewhere else. Chen, Usman, and Badampudi, Information and Software Technology (2024)

Owner time belongs in the comparison. Three deep cases of very small firms looked at how the owner spent time. It was a core mechanism. It either kept the firm able to adapt or damaged that ability. Time spent designing schemas, fixing installers, tending judges, testing upgrades, and repairing workflow state is time taken from customers, delivery, and strategy. The study supports that reading. It does not put an hourly price on the time. Custom work is paid for in those hours. Kevill et al., International Small Business Journal (2021)

The closest on-topic evidence is also the weakest. A 2026 report published by a vendor drew on 149 survey answers (p. 3). It added interviews with people building similar shared-context systems (p. 34). The report says only 17% had one running live (p. 29). Roughly half of those who tried had custom-built it on Claude or Codex (p. 29). The source does not say how many of the 149 had tried. The usual reason a working system was dropped was ongoing upkeep, not the first build (p. 34). About nine in ten of the publisher's own big-company pilots had already built something in-house before buying a maintained product. The number of pilots is not disclosed (p. 34). Read it with care. The publisher sells in this category. The sampling frame is not described. The questions were not registered in advance. The last figure describes that vendor's own pipeline, not the market. The direction is usable at that strength. The sizes are not. It sits below the paper's three warrant levels, which are the three levels of evidence strength set out in Appendix C. It is one non-peer-reviewed practitioner report, good for direction only. It supports one point the sources above cannot reach. For systems like this, the risk of being abandoned sits in upkeep rather than in the build. That is what a maintained shared layer is bought to absorb. It is also why any do-it-yourself test has to run over months of use. It must not stop the moment something first works. Slite, The Ontology of the Company Brain (2026), 41-page PDF edition retrieved August 28, 2026; access requires the email form at slite.com/ebooks/company-brain

Together, those findings support a default you can argue with:

Decision question Buy the base layer when… Build it yourself when…
Does it set the business apart? The job is common working care: context structure, workflow state, test rules, learning records, checking, install, and restart The job carries a distinct customer result, method, data edge, regulated control, or working insight rivals cannot simply buy
Does something on the market fit? It supports the workflow you need without bending your customer promise, your evidence standard, or your approval line Real needs stay unmet, and the cost of the mismatch beats the full lifetime cost of owning the thing
Who can maintain it? A credible maintainer can spread testing, docs, updates, and support across many users The owner has the skill and the lasting time to run, secure, test, document, move, and one day retire what they build
Can the owner leave? Records stay portable, readable, and owner-held; permissions are stated; uninstall and model change do not erase built-up judgment The vendor demands you give up needed data, control, audit access, or a workable way out

Buying, read this way, is a split of labor. Curio Chat Academy tends the shared parts. The owner supplies and keeps the part no vendor can supply, so authorship stays with the owner. Some bought systems trap what matters inside a private store you cannot reach. That can be the customer's language, the current offer, or the owner's voice. It can be the evidence they accept, their decisions, their situated fixes, or their final say. Those have crossed from edge into dependence.

The ability to move belongs inside product fit. A review of 78 selected studies looked at cloud lock-in. Gaps in interfaces, technology, and meaning get in the way of moving and joining systems. It also found that many proposed cures lacked real-world testing. The safe purchase keeps records you can read and a way out you can trust, whatever its ease on the way in. Costa Silva, Rose, and Calinescu, IEEE CloudCom (2013)

Why the shared layer is still hard to build well

"Sets no one apart" does not mean easy. A home-built version must do more than make folders and prompts. It must:

  • hold the power line between source, fact, rule, test rule, verdict, and permission
  • pick the right context without trusting the instructions it pulls in
  • lock tests before anything is written
  • keep the paper trail and the workflow state
  • catch failures without making the model the only judge of itself
  • protect owner edits during upgrades
  • restart interrupted work.
  • uninstall the product without removing what the owner has built up.

Each need is easy to grasp on its own. The burden comes from how they interact, how they age, cross-platform install, testing, docs, and constant fitting to changing models.

Carrying that burden alone is what makes a bought product worth its price. Curio Chat Academy can spread the cost of research, specs, build, testing, docs, and upkeep across many buyers. A solo owner building it alone carries the whole cost. They also carry the cost of becoming its long-term maintainer. The buyer still supplies the valuable custom part — the business truth and judgment. They do not have to rebuild the whole control system first, before that truth can govern any work.

The case for Curio Chat Academy over DIY, and what it must earn

The Curio Chat Academy bet is this hybrid line. The Foundation Pack is meant to supply shared working care and a path for fixes to last. The Vault Pack is meant to scaffold the owner-held business vault and the workflows that use it. The buyer supplies and keeps the business's truth, evidence, voice, decisions, and power. The packs can be replaced without making those records unreadable. Their uninstall rules are meant to leave owner-made knowledge whole.

The supported build carries these market names. The done-with-you work that installs the system is The VAULT Install. It runs five sessions, one per letter: Vault (Session 1, The Install), Assert (Session 2, The Extraction), Uphold (Session 3, The First Refusal), Loop (Session 4, The Live Run), and Transfer (Session 5, The Handover). The owner-held result on day 31 is named The Truth Layer. It is the filled vault, plus the rules, locked test rules, and run records that govern work against it. That is the owner's business, held to its own truth. These are product names, not design claims. They add nothing to the conformance evidence this section demands, and they stand in for none of it.

The design promises are specific. Research findings keep their scope and their counter-evidence. Design objects have stated power and stated life stages. Claims are kept apart from permissions. Test rules lock before judging. Installer writes follow no-clobber rules. Upgrade, checking, rollback, and uninstall are product needs in their own right. That is more care than an ad hoc prompt library carries. It gives Curio Chat Academy a fair basis for one argument: that a finished, checked v1 beats DIY for its intended buyer. To copy the value you must copy the whole tested system and take on its upkeep. That is a different job from copying a set of Markdown headings.

Rigor does not certify itself, however. "Built from current science" means research and standards shaped bounded design choices. It does not mean the papers tested Curio Chat Academy. "Well engineered" needs proof that a release matches a fixed version of this paper. That means tracing each need to a test, passing the acceptance and life-stage tests, stating the limits, and showing demos anyone can repeat. That is the main proof of build quality. Today's design intent is not proof that the packs already form the full system described here. Nor is it proof that buying them beats every capable do-it-yourself build. A credible v1 claim needs a pass on the conformance profile in Section 9. It also needs product comparisons on at least five results:

  1. time from install to a finished governed workflow
  2. repeat fixes, caught failures, escaped defects, and false holds on typical work
  3. owner review, upkeep, and restart burden against the buyer's earlier setup
  4. survival and readability of owner-held records through upgrade, uninstall, model change, moving, and restart.
  5. customer, quality, capacity, or money results where the workflow could plausibly move them.

Until that evidence exists, Curio Chat Academy can honestly sell two things. The first is a research-shaped starting design you can inspect. The second is relief from rebuilding its shared layer from scratch. It can argue that buying is the best default for its intended buyer. The business-specific asset stays owned and portable, while the hard shared work is supplied and maintained. It cannot claim that science has validated this product. It cannot claim that buying suits every solo owner. And it cannot claim that buying the system buys the judgment needed to use it well.

The smallest governed learning loop

The full system spans the whole business. Its smallest useful form fits inside one repeating correction:

Move What changes What makes it real
Write a rule that can fail Turn a repeated fix into a situated rule: what must or must not happen, when it applies, and why A later draft can be shown to pass or break it; "be better" cannot
Put a checkpoint in front of shipping Add a step between "looks done" and "ships" that is not one more full reread by the owner The checkpoint can return a fail, and a held draft proves it has
Retire the rule when its reason expires Remove or change limits that no longer protect the business The live rule set stays worth trusting, instead of becoming a graveyard of old reactions

A learning business opens with those three moves: rule, hold, retire. They do not install the full system. They do not promise a fix will last. They do not remove the owner's approval. They make the missing layer visible. A workspace full of context and instructions is informed. Work becomes governed only when three things are possible. A test rule can be broken, a draft can be held, and a wrong or stale rule can be changed on the record.

The loop also swaps comfort for evidence. Say a fix that applied was never loaded. Then the run exposes a context-picking failure. Say it was loaded, but the draft that broke it still passed. Then the verdict exposes a gate failure. Say the draft was held. Then the held work and the failed test rule show that the system refused it. "I must be prompting wrong" stops being a private worry. It becomes a diagnosis you can test.

3. The business must come before the workflow

A common AI mistake is to start with a tool: "What can I automate?" A better start is a business question: "What is blocking the next real customer or money result?"

A useful map of a one-person business is:

  1. Desire: Is there a specific person with an urgent or real problem, hope, or job to be done?
  2. Offer: Is the promised result valuable, believable, and worth its price and effort?
  3. Position: Can the right person quickly see why this is for them, and why it is truly different?
  4. Reach: Is there one dependable way to earn attention from likely buyers?
  5. Conversion: Does that attention turn into qualified talks, trials, or sales?
  6. Experience: Does delivery create the promised result, trust, repeat business, referrals, or useful learning?
  7. Money: Does the whole system make enough cash and margin to stay healthy?

The owner picks the current binding constraint from this map. AI then works on that constraint. Whatever task is easiest to generate can wait.

Constraint-led management and business tests supply the method. A randomized trial with 116 startups trained some founders to make predictions and test them strictly. Those founders were more precise about what to pursue or change than founders using more usual gut-led search. What it supports is a discipline of stated guesses, tests, and updates. It falls well short of proving any one solo-owner rhythm. Camuffo et al., Management Science (2020)

The constraint record

Before you start a big AI-assisted work cycle, write down:

  • the current constraint, and the evidence for naming it
  • the customer or money result you want
  • the smallest test that could change your belief
  • the success signal and the stop signal
  • the decision the result will inform.
  • the date that decision is due.

If you cannot fill in these fields, production waits. The next task is diagnosis.

4. The model is not the operating system

An AI model turns input into possible output. An AI system adds the app context, data, process, tests, permissions, people, and results around it. ISO/IEC 22989 defines an AI system as a built system that makes output for goals people set. It describes a model as something such a system may use to stand for data, knowledge, or process. NIST treats app context, data and input, the AI model, and task and output as separate parts of an AI system. These standards settle one thing only: model and system are not the same word. None of them ranks designs for solo owners. ISO/IEC 22989:2022; NIST AI RMF 1.0 (2023)

Nine objects that should not collapse into one prompt

Keeping objects apart does not need nine apps or nine databases. Plain files can do it. Apart means each object has its own job and its own power. It has its own paper trail, its own update path, and its own answer when it fails.

Object Its job What breaks when it collapses
Model Write, reshape, compare, sort, or suggest The model's current behavior gets mistaken for business truth, process state, or permission
Harness Run the loop around the model: choose which tools exist, pass the context in, call the tools, count the retries, and enforce the gates A limit is believed because it was written down, when nothing was there to apply it; swapping the app quietly changes what the business can enforce
Workflow state Record which stage the work is in, what ran, what is next, and which moves are allowed A smooth chat poses as a finished process; skipped steps leave no visible gap
Business knowledge Hold current, owner-held facts about customers, offers, voice, money, decisions, and limits Facts turn into prompt folklore: copied, unsourced, stale, and hard to update the same way twice
Sources and evidence Keep the notes and records that high-stakes claims come from A summary becomes its own proof; the paper trail and the ability to re-check both vanish
Rules State lasting musts and must-nots: what has to happen, where, and why Plain facts get treated as enforcement; a fix can be read and still stop nothing
Test rules Turn success or failure into something you can check on this piece of work, before the result is known The writer invents the test after seeing its own answer, and can redefine "good" until it passes
Permissions and gates Set who or what may read, write, publish, send, spend, delete, or change terms The power to do something quietly becomes the right to do it; a draft path becomes an action path
Learning history Keep results, fixes, overrides, reversals, and retirings, with dates and reasons The current state shifts with no audit trail, and the same tuition gets paid twice

The app is what the owner runs, and it supplies the harness — Claude Code, Codex, and Cursor each bring their own. The adapter is the thin layer that fits one app to the master stores.

Two records need to be kept apart whenever they differ in any of seven ways:

  1. Power: who may decide or change it
  2. Paper trail: what note, source, model, or person made it
  3. Update rhythm: per run, per customer event, now and then, or only by a deliberate choice
  4. Keeping: short-lived, current state, add-only, or retired but still there
  5. Load timing: always, task-only, judging-only, or never straight into a model
  6. Permission: who or what may read, write, override, or run it.
  7. What happens on failure: revise, hold, escalate, reverse, or retire.

Those seven ways are the design's governing test. A new filename is not a split if the same actor can quietly rewrite both objects. Two objects can share one folder and still be apart, as long as their power and life stages are stated and enforced.

Why the splits are plausible, in science and in practice

The evidence supports the parts on their own, and only those. The design as a whole is still just a proposal.

Trained models and outside business records fail in different ways. Research on retrieval-augmented generation splits memory in two. Parametric memory is what the trained model holds in its weights. Non-parametric memory is records fetched at run time. Lewis and colleagues gave two reasons for that split: model knowledge is hard to update, and it offers no paper trail. Their tests found gains over a weights-only baseline on the knowledge-heavy tasks they studied. The narrow design lesson is to keep changing business facts outside model weights. The study does not bless any one vault layout. Lewis et al., NeurIPS (2020)

A longer context window does not remove the picking problem. Controlled long-context tests found that where the relevant fact sits can change how well a model uses it. Facts buried in the middle often fared worse. Addressed retrieval and ranking follow as a direction. No fixed cache size, page limit, or retrieval method follows from the result. Liu et al., TACL (2024)

Fetched sources raise a power problem as well as a fetching problem. A peer-reviewed USENIX study tested five prompt-injection attacks and ten defenses across ten models and seven tasks. A later peer-reviewed KDD study traced indirect prompt injection to one cause: models struggle to tell context they should read from instructions they should follow. Peer-reviewed ACL 2026 tests found another awkward result. Several tuned defenses learned surface cues. One trigger token raised false refusals by up to 50%. Accuracy outside the training data fell by up to 40%. The source-versus-instruction line in this paper answers that failure class. Source material stays untrusted data. It gets few permissions. It is tested with both attack cases and harmless ones. No delimiter, label, or trained defense makes any source safe. Liu et al., USENIX Security (2024); Yi et al., KDD (2025); Li et al., ACL (2026)

Writing text does not carry reliable process state. StateFlow splits states and moves from the model actions taken inside a state. On the InterCode SQL and ALFWorld benchmarks it reported more success at lower cost than ReAct. Those results back stated states and moves as a direction worth testing. They do not show that every business workflow needs a formal state machine with a fixed set of states and allowed moves. Wu et al., COLM (2024)

A writer also makes a poor judge of its own work. In an ICLR 2024 study, self-correction with no outside feedback failed to improve reasoning reliably. NeurIPS 2024 tests found that an LLM asked to judge its own writing favored it. ICML 2025 work found that model judges favored models like themselves, and that errors grew more alike as models grew more capable. Model review is still useful. But as the stakes rise, so should the demand for stronger proof. That means tests locked in advance, outside evidence, checks you can run, and a judge with a different paper trail. Huang et al., ICLR (2024); Panickssery et al., NeurIPS (2024); Goel et al., ICML (2025)

Paper-trail standards give vocabulary and no effect size. W3C PROV splits entities, activities, and agents. It can express use, derivation, credit, revision, and who was responsible. That vocabulary keeps the source, the change, the result, and the responsible actor apart, and lets you trace each one. It does not prescribe a file system for a solo owner. W3C PROV-O (2013)

The Curio Chat Academy split-store design

The Curio Chat Academy design treats its store lines as what sets it apart. How much text gets remembered matters less. It is typed, addressed business memory with a paper trail. Its stores may disagree in only two ways: through a stated order of precedence, or a stated decision.

The nine objects above do not mean nine databases. The model and the harness sit outside the owned stores, and the owner does not build them. They still have to be chosen and recorded, because what a harness cannot enforce, the stores cannot make true. Several other objects can share one place and still be distinct record types with distinct moves. The current proposal maps them into six logical stores:

Logical store Holds Power and update rule Loaded when
Source shelf Raw interviews, exports, research, sent material, and other notes Kept whole while held; approved redaction or deletion follows policy and leaves a paper trail; quiet rewriting is not allowed Only when a task needs the source; always treated as data, never as instructions
Business vault Curated current facts: customer, offer, voice, evidence, money, decisions, and identity lines Owned by the owner; high-stakes entries keep their source and date Addressed pages picked for the task; the whole archive stays unloaded
Craft base General standards for doing a class of work well Kept apart from the owner's facts; refreshed as the general standard moves On demand, for the skill in use
Correction store Lasting fixes about how work should be done, with scope and reason Added or changed by a learning decision; reversible and able to be retired When planning or checking the affected workflow
Expertise store What repeated use of a skill has taught for this owner, apart from general craft and business facts Grows from watched use and correction; the learning survives tool upgrades Only for the skill and owner it applies to
Run and judging ledger Workflow state, chosen input, model and tool paper trail, locked test rules, verdicts, approvals, overrides, and results for one run Add-only for the run; test rules lock before the draft is written or seen by the judge; high-stakes changes name an actor and a reason During work, judging, restart, and later audit

The first five stores describe the current split-memory design at a public level. The sixth is the main improvement this review proposes. Without a run and judging ledger, the system can know the business and still fail one question. Which knowledge, rules, test rules, model, and permissions governed one draft?

No scientific consensus says a solo owner's AI system needs exactly these six stores. The evidence supports narrower splitting parts instead. This store design is the smallest current layout that gives each record a home matched to its power and its life.

Every store declares how it goes stale

Splitting decides where a record lives. It does not decide when the record stops being true. A store with no staleness policy slowly turns into the stale wiki this design exists to replace. The seven ways records differ already make each store declare what it keeps and how often it updates. That statement is not finished until it also says how staleness gets spotted, and what spotting it does.

Each store therefore declares two more things, plus one named case of the failure response it already declares:

  1. Staleness signal: the visible sign that a record no longer matches reality. Usually that is age, disuse, being overtaken by a later decision, or a clash with current evidence.
  2. Detection method: what actually makes that signal. Here you must tell a periodic reread by the owner apart from a check the system runs.

The named case is the retiring path. A positive signal triggers one of the failure responses the store already declares. That is a proposed change, a hold on dependent work, an escalation, or a retiring record. The proposal waits in a review queue until the owner rules on it.

The detection method decides whether the rest works. Say detection is only the owner rereading the vault on a schedule. Then it fails the way capture-by-memory fails. It works right up to the week it is needed most. Some methods do not depend on the owner noticing. You can re-link and merge recent entries on a schedule. You can lower the retrieval weight of records that are both old and unused. You can compare what a stored record claims against current evidence in the source shelf, the run ledger, or a connected tool. Then flag the clash for review. The third one catches an old price or a dead policy before it reaches a draft, not after a customer has read it.

Detection is not power. A staleness signal proposes. It does not retire. Retiring stays a decision someone answers for, and it leaves a record. A system allowed to quietly delete what it judges out of date can lose the business's memory of why it once decided otherwise. Both directions of this failure are real. A store that never prunes piles up confident wrong answers. A store that prunes on its own destroys the evidence that would have corrected it.

Nobody has tested whether automatic staleness detection repays its upkeep for one person. It belongs with the other open claims in Section 14. The narrower claim stands on its own. A store whose only staleness check is the owner's memory has not declared a life. Its decay will be found by a customer-facing draft, not by the system.

Every high-stakes record carries two dates

A record has two times that often differ. Collapsing them costs the very thing the run ledger exists to give you. One is the date the record's content starts to apply to the business. The other is the date the system came to hold it. A price that changed on the first of the month and was fixed in the vault on the ninth has both, eight days apart.

Usually only the first gets recorded. That is enough to answer what is true now. It is not enough to answer what governed one piece of work. A run can be judged fairly only against what the system believed while it ran. With the first date alone, a backdated fix rewrites the past. Every draft made before the ninth now reads as having broken a rule that did not yet exist. The ledger reports holds that never happened, and misses ones that did. The false-hold measure in Section 9 stops working at all. You can no longer tell a gate that failed from a standard that arrived later. The same collapse quietly flatters the system, since a backdated fix makes the history look tidier than the work was.

The requirement is narrow. High-stakes records — facts, decisions, rules, test rules, and the ledger's own entries — carry both dates. One is the date the claim applies from. The other is the date the system wrote it down. Both survive being overtaken and being retired, rather than getting overwritten by the change that ended them. Rebuilding the past then reads the store as it stood then, not as it stands now. That is what lets someone who was not there review a run months later. Nothing here demands a particular storage tool. The property is that the second date exists, is written when the record is, and is never revised. A plain file satisfies it with one line. A database loses it with one careless update.

Moving and running many models are conformance properties

Readable files are needed to move a system, but they are not enough. A folder can move while its links break. A skill can parse while its safeguards vanish. Two agents can each behave well and still quietly overwrite each other. Moving must therefore be proven at three separate layers:

  1. Machine move: the master records can move to another supported machine or path without changing their meaning, ownership, or history.
  2. App move: the same stores and master skills work through different AI apps. No separate copy of the business is made for each one.
  3. Model move: a workflow can change model vendor or model version without treating the old model's chat memory as the source of truth.

These layers should not be blurred. Claude Code, Codex, Gemini CLI, Cursor, and GitHub Copilot are AI apps, or product surfaces. Gemini Flash is a model you pick inside an app. A conforming run records the app and its harness, the model vendor, the model name, the adapter version, and the tool setup. Otherwise "works with Gemini" is too vague to test.

Context control is the property these layers protect

On May 29, 2025, Slack limited the two Web API methods that read message history. New unlisted apps, and new installs of existing unlisted apps, were held to fifteen messages per request at one request per minute. Internal customer-built apps kept 1,000 messages per request at fifty or more requests per minute. A terms update at the same time changed how third-party apps may store, use, and share data taken through those APIs. The changelog names bulk downloads of chat data, for indexing and querying, as the risk being addressed. Existing installs of already-distributed apps were left alone. The narrow reading is the only one this source supports. One platform tightened one path for stated security reasons. Nothing here forecasts what any other platform will do. The design lesson does not depend on this case. Some designs need standing permission to bulk-read another company's platform. That rests a lasting asset on a business policy they do not control, and which can change without their consent. Slack API changelog (2025)

Being able to move is the answer to that exposure. The property it protects is context control. The owner holds the fullest assembled version of the business — not a platform, and not a model vendor. The owner decides who may process which part of it. Context control is the design half of the owner control named in Section 2. It is the same property, stated as a rule about where records live rather than as a mark of success.

Any system that reads across every source builds up that assembled version somewhere. Whoever holds it sits in the strongest lock-in spot in the stack. The pressure comes from two directions that are easy to confuse. App platforms hold slices of the owner's context and can restrict access to them. Model vendors hold the layer that runs the work, and can make the business's working memory a feature of one vendor's account. The two threats are not the same, and they do not substitute for each other. A design that guards against one and ignores the other has covered half the exposure.

Context control is therefore a design requirement, not a preference. It is what the master data plane below exists to deliver. The owner's assembled truth lives in records the owner controls. Every platform and model sits beside it as a part you can swap out, not above it as a landlord.

One master data plane, thin app adapters

There must be one master owner-held store root. A Claude vault, a Codex vault, a Gemini vault, a Cursor vault, and a Copilot vault that drift apart would defeat the design. The lasting records live in that shared root. An app adapter may tell its agent how to find the root and load an addressed context packet. It may call a skill, translate tool and permission names, and emit a paper trail. It may not fork or redefine the business truth.

One shared skill source is practical, because the named apps increasingly support file-based Agent Skills. Claude Code describes its SKILL.md format as an open standard. OpenAI documents the same open Agent Skills base for ChatGPT and Codex. Current Gemini CLI docs support user and workspace skills, including the shared .agents/skills/ location. Cursor describes Agent Skills as portable and finds .agents/skills/. GitHub Copilot accepts Agent Skills from shared skill folders. They still differ in how they find instructions, in optional frontmatter, in permissions, in hooks, in tool names, and in session behavior. The master skill should therefore stay inside the shared standard where it can. A thin, tested adapter owns every app-specific twist. Claude Code skills; OpenAI Codex skills; Gemini CLI skills; Cursor Agent Skills; GitHub Copilot agent skills

The first v1 conformance target is one owner using all five surfaces on the same machine:

Target surface Shared core What the app adapter must handle
Claude Code Master stores and Agent Skills Instruction discovery, paths, permissions, hooks, and paper trail for the Claude Code setting
ChatGPT/Codex The same stores and Agent Skills AGENTS.md or a like instruction projection, Codex permissions and tools, and paper trail for the live Codex surface
Gemini CLI with Flash picked The same stores and Agent Skills Gemini context, policy, tool, and model-choice mapping, with no Gemini-only copy of the records
Cursor The same stores and Agent Skills Cursor rule, permission, tool, and life-stage mapping, with no change to what the master skill means
GitHub Copilot The same stores and Agent Skills Copilot instruction, permission, tool, and life-stage mapping for the supported Copilot surface

Official product docs show that these extension surfaces exist. Whether one skill behaves the same across them is a separate question the docs do not answer. Each row is a release target. It must pass the same typical workflow, and the same failure fixtures, before Curio Chat Academy advertises it as supported. A failure fixture is a prepared input that carries a known fault. An app that lacks a fixed gate, a file ability, or a permission line is unsupported for that workflow. It stays that way until an outside control supplies what is missing. A prompt that says "be careful" is not an adapter for a missing control.

Many apps on one machine

One owner must be able to use several supported apps against the same master records on one machine. Reading at the same time is fine. Lasting writes must pass through one checked write path. Each write lands whole or not at all. It names the record version it expects to replace. It holds a lock, or something like one, while it runs. A stale session must re-read before it commits. A clash must hold the write for sorting out. Last-writer-wins — where the most recent save quietly replaces every earlier one — is not an acceptable learning policy.

Prose cannot enforce that rule while every app keeps unrestricted write access to the master root. A supported writable adapter must make a direct bypass either impossible or visible. It can use file permissions, a sandbox policy, a write broker, integrity checks, or a similar fixed control. Where an app cannot hold that line, it may run read-only, or stay unsupported for lasting writes.

Every run records the app, the vendor, the harness, and the model, with the version of each. It records the adapter and skill versions, the context manifest, and the files read. It records the proposed writes, the verdict, and the approving owner. The stable owner identity belongs to the master system. Vendor accounts are run-time identities noted in the paper trail. They hold no ownership and no approval power. App-specific chat memory never rules. A fix committed through one app becomes available to every other app from the master store, with no syncing between private memories.

Running many apps also creates a disclosure line. An adapter gets only the task's approved context packet. It does not get automatic access to the whole vault. Credentials and secrets stay outside model-readable stores. The owner can let one vendor process a record without letting every installed vendor process it. Gates on high-stakes action sit outside the model, and apply whichever app proposes the action.

Disclosure has an inbound half

That line governs what leaves for a vendor. It says nothing about what may enter, and the two directions are not the same. An outbound mistake exposes a record the owner already holds. An inbound one puts material into the business's memory. The owner may have had no right to store it, and no stated basis to keep it. They may have no way to find it again when that basis runs out.

The places this arrives from are the ones a working solo owner reaches for first.

  • A client's call recording.
  • A support inbox.
  • A payment processor's transaction history.
  • An analytics export.
  • A connected calendar.
  • A screen-watching tool.

All are likely sources of exactly the evidence the vault wants. Each also carries a duty the file itself does not state. Whose material is it? What was it collected for? How long may it be held? And what happens to it when the relationship ends?

The answer is to make entry depend on a record, not on a setting. A source class enters only while a current grant covers it. The grant has to answer five questions before it is worth anything.

  • Why may this material be held?
  • What may be done with it?
  • How sensitive is it?
  • How long does it survive?
  • Between which dates is the grant itself good?

The grant is checked where the write happens, not where the owner set it up. That is the same reasoning the paper applies to a rule in an instruction file. A limit the writer enforces is a control. A limit recorded next to the writer is documentation. Expiry and withdrawal are checked at write time. They are not advisory notes.

One asymmetry keeps the control from taxing ordinary work. The owner's own judgment, said directly while working, is the one source that needs no grant. Everything the system watches, rather than is told, needs one. The same line decides what withdrawal does. It closes the grant and stops further entry. It does not quietly erase evidence that past work relied on. You would be destroying the record of what governed a decision, to suit a wish you have today. The retiring rule above already refuses that. Deletion stays a separate act, clearly scoped and honestly reported.

Moving between machines

A conforming store can move to a different supported machine, user account, or root path. It can then be rebuilt from its master manifest without editing its real records. That requires:

  • lasting links that are relative or ID-based
  • machine-local paths and credentials kept in adapter config you can replace
  • stated text encoding, schema version, and upgrade steps
  • tested backup and restore.
  • offline edits that diverge showing up as clashes before anything is overwritten.

Copying files is not enough. The move test must show six things about the new machine. It finds the same records. It checks they are whole. It discovers the supported skills. It runs a typical governed workflow. It keeps the owner's power line. And it adds a valid run record. Removing one app adapter must leave the shared stores, the other adapters, and the built-up learning whole.

Restoring is admitting, not copying

Export and import are not one job run in two directions. An export proves the records can leave. An import decides what becomes true inside the store. That makes it the lasting write path with the widest reach in the system. A restore, a move, a second machine's first load, an upgrade, and a file a model was asked to make all arrive through it. Each carries records that assert their own standing: this fact was owner-approved, that pattern was ratified, this note was permitted.

Integrity checks cannot settle those assertions. A hash shows that the document has not changed since it was written. It shows nothing about whether the power the document claims was ever granted. A store that checks the bytes and then accepts the claims has built a path for laundering power. The attacker it needs is ordinary: anyone who can hand the owner a file. That includes a stale backup, a well-meaning peer, and the owner's own model output.

Admission therefore re-runs the checks. It does not read each record's account of having passed them. Every admitted record is re-checked against the current contracts. Its permissions are re-checked against grants the receiving store itself holds. Its claimed power is capped at what that store can prove on its own. Anything that cannot be settled fails admission. It does not fall back to a weaker check, because a lowered check on a bulk path is the same as no check at all. The corollary is the part most likely to get skipped, so it is worth saying plainly. The business's own export becomes untrusted material the moment it leaves the store. It comes back on the same terms as any other document.

This is the line the paper already draws around fetched sources, applied to the one path that looks like housekeeping rather than input.

This is a target design. It makes no claim about the current packs. The release evidence must name a target surface and pass it. Until then, treat the Vault Pack and Foundation Pack as native only to the settings their release notes name.

What the review backs, and what stays open

Line Evidence status What follows for the design
Model ↔︎ outside business knowledge Directly backed as a useful technical split in RAG research; also fits ISO and NIST system concepts Keep owner-held knowledge outside model weights and model-specific memory
Model ↔︎ harness A careful design guess. The prompt-injection and tool-permission evidence shows that what a run may reach is set outside the model. No located study measures naming the harness as its own object Record the harness and its version in the paper trail; treat a control it cannot enforce as absent, not as stated
Raw source ↔︎ curated business fact Backed by paper-trail needs and prompt-injection evidence; the exact curation method is untested here Keep the source shelf and business vault; harden the rule that source content is data, never power
Business vault ↔︎ craft base A careful design guess from different ownership, update rhythm, and scope; we found no study that tests this split Keep for now; test whether it stops owner facts and general standards from overwriting each other
Business fact ↔︎ working fix A careful design guess from the difference between describing and requiring; not tested directly in solo-owner systems Keep for now; measure misrouting, contradiction, rule reach, and the cost of retiring
Craft base ↔︎ owner-specific expertise A product design guess meant to hold situated learning across skill upgrades; no direct effect size Keep for now; test moving, reuse, and cross-contamination across owners and skills
Stored context ↔︎ workflow and judging history A gap exposed by paper-trail, state-grounding, and judging evidence Add a run and judging ledger; memory stores should not pose as process state
Writer ↔︎ test rules and judge Backed by self-correction, self-preference, and correlated-error research, plus NIST's practice of splitting build/use from checking actors Lock test rules before the draft is written or seen; raise judge and evidence independence as stakes rise
Hot view (a small, frequently refreshed cache) ↔︎ deeper addressed pages Long-context evidence backs selective retrieval as a direction, not this cache design or its limits Treat cache contents, page size, and load limits as product claims to measure
Checked transport ↔︎ approved record A careful design guess from the gap between proving a document is unchanged and proving its claimed power was granted; the prompt-injection evidence backs the underlying source/power line, not this use of it Re-check on admission, and cap claimed power at what the receiving store can prove on its own
Date of effect ↔︎ date of record A careful design guess from the fact that a run can be judged only against the beliefs it ran under; we found no study that estimates the cost or use of the second date in a one-person business Carry both dates on high-stakes records, and rebuild a past run as of its own date

The design is open to change under a strict test. A line stays only while it improves context picking, the paper trail, enforcement, restart, or owner review burden enough to be worth its upkeep. Science can add a line, merge one, or change how it is built.

Rules are not test rules

Four records look alike and do different jobs:

  • A rule lasts across all eligible work: "Never invent a customer result."
  • A test rule makes that rule checkable on one piece of work: "Every result claim on this page must map to a dated evidence record; an unmapped result claim holds the draft."
  • A verdict records what happened: pass, revise, hold, or escalate, with the failed test rule and the evidence.
  • A permission decides what may happen next: even a passing draft may still need the owner to approve publishing.

Keeping these apart stops a style preference from becoming hard policy. It stops a rule from claiming enforcement just because it was loaded. And it stops a passing check from granting power it never had.

The system in one picture

                            HUMAN POWER
            purpose · identity · risk · final approval

                                │
                                ▼
 SOURCES → BUSINESS VAULT + CRAFT + CORRECTIONS + EXPERTISE
    │                           │
    │                           ▼
    │             WORKFLOW STATE + LOCKED TEST RULES
    │                           │
    │                           ▼
    └──────→ APP ADAPTER → HARNESS ⇄ MODEL ⇄ TOOLS
                                │
                                ▼
                        CANDIDATE DRAFT
                                │
                                ▼
              JUDGING → VERDICT → HOLD / ESCALATE / GATE
                                │          │          │
                                │          │          ▼
                                │          │    approved action
                                │          ▼
                                │    simplify or reset
                                ▼
              RUN LEDGER + CUSTOMER/MONEY RESULT
                                │
                                ▼
          update the right store · reverse or retire a rule

The business loop chooses where work should go. The work loop turns one business need into a bounded draft. The learning loop changes the right store only after evidence, or after a decision someone answers for. The human line owns purpose, identity, risk appetite, and high-stakes action.

This makes the Curio Chat Academy principles visible:

Curio Chat Academy principle Working mechanism Sign that it is present
Pattern before prompts Start from real workflows, customer language, and repeated fixes The context packet cites what was seen; no imagined ideal process stands in for it
Design before automation Split stores, states, test rules, permissions, and moves before you schedule anything The assisted manual path works and leaves a run record you can restart from
Loops before output Ship to a customer, measure the result, and send the learning to the right store A result changes a fact, rule, test rule, method, or constraint, with a paper trail
Ownership before dependence Keep portable owner-held records and human power over high-stakes action The owner can explain what governed a run, and can change models without losing the business
Compounding before tuning Keep proven learning and retire stale rules Repeat errors fall without an ever-growing blob of prompt text
Subtraction before expansion Name the source of capacity, the swap, or the retiring before adding recurring work A new commitment cannot enter the live path without a stated trade-off

A business that saves everything is not always learning. Learning needs picking, typed power, evidence, revision, and forgetting on purpose.

5. The seven-stage workflow

Stage 1. Orient: load the smallest useful truth

Start with the slice of business context that matters. The whole archive stays out of the run.

Useful context can include:

  • the current constraint record
  • the target customer's language, situation, and wanted change
  • the current offer and its evidence
  • voice and identity lines
  • recent decisions and unsettled clashes
  • source material for this task.
  • the meaning of done.

The design treats memory as addressed, in five parts:

  • a short, well-tended hot view for facts you need often
  • a business vault for curated, sourced business knowledge
  • a separate source shelf for raw material, which stays data and can instruct nothing
  • craft, correction, and expertise records that fit the task.
  • a new run record naming what was actually picked.

The hot view is a cache. It points to the deeper material, which stays the source of truth. Raw sources keep their origin, trust status, access line, and how long they are held. Material comes from an interview, a web page, an email, or an uploaded file. None of it gains the right to rewrite rules or call tools just by entering the context window. The context manifest should tell what the workflow asked for apart from what the system can prove it loaded.

Output: a context packet. It holds the truth that matters, the sources cited for this task, the rules that apply, and a manifest of what was loaded.

Stage 2. Choose: name the constraint and the decision

The owner answers three questions:

  1. What belief about the business are we trying to change or test?
  2. What decision will the result inform?
  3. What is the smallest piece of customer evidence worth getting?

Examples of the reframing:

Tool-first request Constraint-first request
"Write ten posts." "Test whether this customer sees the problem in this language."
"Build an email agent." "Learn whether a five-message welcome series lifts activation."
"Research competitors." "Find which promise customers already pay for, and where they stay unhappy."

Output: one decision, one success signal, and one stop signal.

Stage 3. Specify: turn intent into a brief you can run

AI work gets better when intent is spelled out. The brief names:

  • the result and the audience
  • which inputs rule, and in what order
  • what must appear and what must not
  • the evidence standard and the output shape
  • known ways it can go wrong
  • the current capacity tier, defined in Section 13, and the fallback path.
  • the line where the human approves.

The test rules are a separate record. They turn lasting rules and task intent into conditions this piece of work can pass or fail. Lock them before the draft is written or seen by the judge. Otherwise the standard can drift toward whatever the model produced. Say review turns up a fair test rule you missed. Version the test rules, and say plainly that it was added after seeing the draft. Then apply it to a fresh draft, or to a held-out case the test rules were not written against. The brief is where the owner does the authoring and makes the test fair. That is why it cannot be treated as paperwork.

For uncertain work, ask AI to surface its assumptions and any clashes before you ask for the final piece. For high-risk claims, require an evidence ledger with the source, date, scope, and the exact point it supports.

Output: a brief you can run, locked test rules, and an approval line.

Stage 4. Produce: hand over bounded, reversible work

AI suits jobs such as:

  • sorting raw customer language
  • comparing research sources
  • making options from a strategy the owner set
  • drafting and reshaping content
  • applying a known checklist or template
  • playing out objections or edge cases
  • pulling out structured facts
  • preparing test variants.
  • writing down a process that already works.

The default line for the work is:

  • read the least you need
  • use the fewest tools you need
  • make a draft or a reversible change
  • keep the source material
  • record uncertainty
  • do not quietly widen the task or create a new commitment.
  • stop before any public, money, contract, destructive, or privacy-sensitive act.

The run record also names the harness and its version, and the model family and version, where known. It names the settings that matter, the tools made available, the context picked, and the version of the resulting draft. Model identity does not explain every error. It is recorded because you cannot repeat or compare a result when you do not know what produced it.

Tests show that task fit matters. In a study of 453 professionals doing writing tasks, access to ChatGPT cut time by 40% and raised judged quality by 18%. In a field rollout with 5,172 customer-support agents, AI help raised issues resolved per hour by 15% on average. The biggest gains went to less experienced and lower-skilled workers. These are real results, tied to particular tools, tasks, and settings. They say nothing about business work in general. Noy and Zhang, Science (2023); Brynjolfsson, Li, and Raymond, Quarterly Journal of Economics (2025)

Two loops run here, and they are easy to confuse. The outer loop is the seven stages of this section. It turns one business need into one checked piece of work, and the owner sees every turn of it. The inner loop is the harness working: the model proposes, a tool runs, the result comes back, and the model goes again. That loop can turn many times inside Stage 4 without the owner seeing any of it.

The limits in this paper attach to different loops. A revision limit counts turns of the outer loop, after a verdict. A retry limit counts turns of the inner one, before there is a verdict at all. An inner loop with no limit is the usual way a run burns an afternoon and produces nothing to judge. An outer loop with no limit is the usual way a draft gets polished past the point where anyone should still be looking at it.

Output: a candidate draft, a change log, and stated uncertainties.

Stage 5. Verify: match the check to the stakes

What you check depends on what could go wrong. There is no single universal fact-check step.

Output Least checking needed
Private idea work Check strategic fit and freshness; label the assumptions
Internal summary Compare the claims the summary rests on against the source material
Marketing copy Check stated and implied claims; check voice and offer fidelity
Customer advice Check that it is right, that it applies, what it excludes, and when to escalate
Code or automation Run it on typical and hostile cases; look at side effects
Price or money model Recompute the inputs and formulas yourself
Public research Trace every high-stakes claim to a fit source
Send, publish, spend, delete, grant access, or change terms Human approval right before the act

Checking should include running the thing, when running it is the real test. An automation that reads well can still fail on held-out input. Copy that reads well can still mean something else to the target customer, and only asking them shows it. A citation that exists can still fail to support the claim as written.

A checkpoint is not real just because it appears in a procedure or a folder of instructions. It must judge against the locked test rules, be able to return a fail, run on the path that matters, and keep a verdict. For high-stakes work, prefer evidence the writer did not author. That means source records, checks you can run, or held-out cases. It also means a judge with a different source, or human review someone answers for. A second pass by the same model can still help. It is not independent just because it has a different prompt or a different agent name.

The need for this stage rests on evidence. A pre-registered study followed 758 knowledge workers. AI help improved speed, volume, and quality on tasks judged to sit inside what the model does well. Then came one hard task outside that edge. The control group got it right 84.5% of the time, against 60.0% and 70.6% in the two AI-assisted groups. That is an average drop of 19 percentage points. Dell'Acqua et al., Organization Science (2026), §4.2 and Table 7

Output: a verdict of pass, revise, hold, or escalate. It carries the failed test rules, the evidence, and the judge's paper trail. The allowed next step is recorded apart from it.

Stage 6. Ship and measure: let reality answer

Do not confuse an approved draft with business progress. Put the smallest responsible version in front of a customer or a market.

Depending on the constraint, that could be:

  • five customer interviews
  • a sales talk using a revised promise
  • one landing-page variant
  • a paid pilot
  • a small welcome-flow change
  • one channel test, or
  • a revised delivery step tried with current customers.

Write the guess down before you see the result, and pick a window to watch. Tell early signals (clicks, replies, calls booked) apart from customer and money results (sales, activation, completion, repeat business, refunds, referrals, margin).

Output: what you saw, tied to a guess made in advance, including the negative and unclear results.

Stage 7. Capture and recalibrate: make the next cycle better

Close the loop by sending each result to the store whose power it changes:

  1. Business vault: What sourced fact or owned decision changed about the customer, offer, channel, or experience?
  2. Corrections: What should the workflow do differently next time, where, and why?
  3. Expertise: What repeatable, skill-specific pattern now has evidence beyond one good output?
  4. Craft base: Did general evidence change a standard that should apply to everyone, not just here?
  5. Run ledger: What happened, what failed, what was overridden, what shipped, and what followed?

This stage is the most likely to fail, and it fails for a structural reason, not a motivation one. It is the only part of the loop whose running depends on the owner choosing to do it after the work is already done. A step you must remember is the step that stops first under load. The run it skips is rarely a typical one. Capture breaks down exactly in the weeks that produced the fixes worth keeping, which quietly tilts the learning history toward days that went well.

The design answer is to make the run write its own record. Do not ask the owner for more discipline. Running a governed workflow should write its own record as a by-product, the way a delivery leaves an invoice. That record holds the run ledger, the locked test rules, the verdict, any override, and the result. Capture that survives is capture that rides along with work already being done.

What is left for the owner is the call, not the typing: confirming, correcting, redirecting, or rejecting entries the run has already proposed. That split keeps both properties. Automatic capture removes the owner's authorship of the record. It must not remove their power over it. Nothing enters a lasting store on the model's own say-so.

Do not encode a process just because it happened twice. Encode work once you understand its inputs, decisions, quality checks, exceptions, and owner. Every working rule should keep the situation and reason that made it, so it can later be challenged, changed, or retired. Automating confusion makes it faster and harder to see. Keeping dead rules makes the system refuse the wrong things with confidence.

Then go back to the scoreboard. The old constraint may have moved, or a new one may now rule. The system compounds only when each cycle changes where the next one starts.

Output: a closed run record, a stated routing call for each learning, any change or retiring that is justified, and the next constraint decision.

6. What must stay human

AI can take part in nearly every stage. It should not own every decision.

The owner stays responsible for:

Purpose and identity

  • Which customers the business chooses to serve.
  • Which problems it refuses to exploit.
  • What the business promises, and what it will not promise.
  • The values, taste, and point of view that make its work recognizable.

Business judgment

  • Which constraint matters now.
  • Which evidence is enough to act on.
  • Which trade-offs are acceptable.
  • Whether advice fits the real customer and the real setting.

High-stakes actions

  • Publishing in public, and talking straight to customers.
  • Spending, pricing, refunds, and money commitments.
  • Contract, legal, medical, safety, or regulated claims.
  • Granting access, making destructive changes, and using sensitive data.

What becomes lastingly true

  • Which proposed fact, fix, or retiring enters a lasting store.
  • Which of two clashing records the business will treat as current.
  • Which learning is general enough to change a rule, rather than describe one run.

The four high-stakes categories above all point outward. An outward gate is the obvious place for a human, because the damage is immediate and visible. Writes that point inward deserve the same line, for the opposite reason. A wrong entry in the vault sends nothing to a customer on the day it is written, so it trips no high-stakes gate. It waits. Then every later run that touches the topic pulls it up as settled truth. Its cost is paid later, by a draft that is confidently wrong for a reason nobody can see. That is the mechanism behind Failure 9.

Deciding what is true enough to keep needs three judgments. Which of two accounts describes the situation better? Is a one-off exception a new rule or a bad week? Is a fix general, or local to one customer? Those are judgments about the business, not about the text, and the model cannot reach the evidence that settles them. That gap between a person and a model is what makes the gate necessary rather than ceremonial. A model asked to rule anyway will produce a confident merge. A confident merge of a live disagreement is a decision taken with no owner.

This does not mean the owner must write the records. Stage 7 exists so they do not have to. It means a proposal stays a proposal until a person rules on it, and the resulting record names the actor and the reason. The loop pays off because the owner reads candidate entries instead of writing them. It does not pay off because the system decides on its own what the business believes.

Keeping your own skills

  • Customer conversations.
  • Judgment about the offer and its position.
  • Reading primary evidence.
  • Judging quality in the owner's core craft.
  • Working now and then without AI, on skills the owner cannot afford to lose.

This line fits heavy automation. It exists because of a reliance problem. If a person stops forming beliefs, weighing options, or practicing their core craft, short-run output can sit right alongside long-run fragility. A 2025 survey of 319 knowledge workers, covering 936 reported AI uses, found a pattern. More trust in AI went with less reported effort at critical thinking. More trust in oneself went with more. The study watches and self-reports, so it does not prove cause. It does name a risk worth watching. Lee et al., CHI (2025)

7. Use AI as a menu of bounded skills

Many solo-owner systems bring in a whole staff, all built as agents:

  • a chief of staff
  • a marketer
  • a researcher
  • a salesperson
  • an ops manager
  • a board of advisers The image is appealing and often costly to run.

Every standing role adds context, routing, hand-offs, clashes, upkeep, and checking. A one-person business should start with one AI workspace and a menu of bounded skills.

Useful skill modes include:

Skill Job Input it needs Human gate
Research scout Find and compare evidence that fits Question, source policy, date limit Accept the evidence and the reading of it
Knowledge interviewer Interview the owner to draw out unwritten reasoning behind a decision, standard, or exception Topic, records to avoid re-asking, session limit Confirm the transcript, then decide what becomes a source
Customer-language analyst Group pains, hopes, objections, and alternatives Raw interviews or reviews, with permission Decide what is typical
Offer critic Stress-test promise, proof, friction, and risk Offer, customer evidence, money Decide the position and the terms
Content producer Draft from an approved thesis and evidence set Brief, voice, claims, channel Approve publishing
Conversion analyst Diagnose a funnel or a sales talk Event data and what people said Choose the change
Experience analyst Find repeating friction in delivery Support, completion, and feedback data Choose the customer-facing change
Operations steward Apply checklists, keep logs, prepare reviews Approved process and permissions Approve high-stakes actions

Every other skill on that list eats records the business has already made. A one-person business has no colleague to ask, and no internal message archive worth mining. So a structured talk with its owner is likely one of its biggest inputs. Producing that talk is why the interviewer is on the menu. What that talk is for is the reasoning a document leaves out:

  • why a price moved
  • which client type is quietly turned away, and on what sign
  • which objection is fatal and which is noise
  • what the last failed launch really taught

Little of it can be fetched by a connector, because most of it was never written down.

Two lines keep the skill honest. Its output is a source, not a fact. An interview transcript enters the source shelf with its date and its origin as an owner interview. It becomes business truth only through the same curation any other source passes. The interview is also generative in a way that makes it risky. A model that asks a leading question, and gets an agreeable answer, has made a belief rather than found one. A year later that belief will look just like evidence. Prefer questions about specific past cases over questions about general policy. Record what the owner actually said, not a tidied summary. Treat a confident answer to a question the owner has never considered as a guess, never as a fact you retrieved.

Promote a skill into a lasting workflow only when:

  • the work repeats
  • the inputs exist and are governed
  • the output you want can be tested
  • the ways it fails are known
  • a human owner is named.
  • the upkeep costs less than the value it wins back.

Use several agents only when the work truly splits, or gains from independent criticism. Work splits when two parts need no result from each other, so they could run in either order. Work that only looks split still passes something between its parts, and then the second part waits on the first whatever you call it. More agents do not make more truth. A pre-registered meta-analysis of 106 studies found that human-AI pairs beat humans alone on average. They did worse, though, than the better of the human or the AI alone. Results varied a lot by task. Losses were clearest on decision tasks, while creative tasks looked more promising. Vaccaro, Almaatouq, and Malone, Nature Human Behaviour (2024)

8. The work rhythm

A workflow survives only if it works on an ordinary Tuesday, and after a week that fell apart.

Daily: orient, focus, close

Orient

  • Read the live constraint and decision.
  • Clear what arrived overnight: proposals from scheduled runs, and anything yesterday's close escalated rather than settled. The queue should hold candidate entries the system wrote, not work waiting to be written.
  • Load only the context today's work needs.
  • Pick one thing you can ship.

Focus

  • Write the brief before you open the production loop.
  • Use AI for bounded work.
  • Keep customer contact and high-judgment work in the owner's own diary.

Close

  • Confirm that today's runs left records. Add only what the run could not see: why a judgment went the way it did, and what is still open.
  • Correct a proposed entry now, while its cause is still visible. Left in a queue, it will read as agreed fact tomorrow.
  • Leave a re-entry note that makes tomorrow easy to start.

Weekly: reality and reliability review

Look at four views:

  1. Business reality: What did customers do? What did the money do?
  2. Constraint: Is the named constraint still the constraint?
  3. Reliability: Where did work stall, restart, or need redoing?
  4. Learning: What fact, decision, fix, or method should last?

Then choose:

  • one business constraint
  • one test a customer sees
  • at most one working improvement.
  • one thing to stop.

Monthly: calibration and upkeep

  • Review the offer, position, channel, experience, and money together.
  • Work the staleness queue. Retire or change the records flagged by age, disuse, being overtaken, or clashing. Decide which flags were wrong. A month with no flags means detection has stopped running, not that nothing has aged.
  • Settle clashes between stores that precedence did not resolve.
  • Review permissions, sensitive data, connections, and scheduled automations.
  • Test whether the owner can still do and judge core tasks without AI.
  • Retire workflows whose upkeep or checking costs more than they return.

The worst-day version

When capacity collapses — a sick child, a migraine, a client emergency that eats the week — the system shrinks to:

one inbox
one live constraint
one useful customer-facing action
one checking gate
one re-entry note

Getting back on your feet is part of productivity. A system that works only at full energy is not reliable.

9. How to test the design

The split-store design is a product claim built from backed parts. It should be tested as one. A strong theory of why the stores should differ is not evidence that this build improves a solo owner's work.

The main proof of build quality: matching the paper

The strongest sign that The One-Person AI Operating System is well built is a repeatable demo that a named release matches the design it claims. "We used great rigor" is a story about the builders. A versioned conformance table is evidence about the build.

Every release that claims a match should name a fixed paper version, or a content hash. It should then map each need below to four things:

  • something you can inspect
  • a test or demo anyone can repeat
  • the settings tested
  • a dated result

A skipped need is written down as a stated gap and earns no credit. Changing the paper does not change the target a past release was measured against.

ID A matching release must show Least evidence needed
C-01 Master ownership Business truth, sources, fixes, owner-specific expertise, run history, and results stay in plain files the owner controls, outside model weights and private chat memory File and schema check; ownership and export test
C-02 Store power Source, fact, craft, fix, expertise, workflow state, test rule, verdict, permission, and result keep their stated power and life lines Schema checks, plus seeded misrouting and precedence fixtures
C-03 Addressed context A run loads the smallest approved context packet, and records exactly which versions were supplied Context-manifest test with records that are relevant, irrelevant, stale, and forbidden
C-04 Fixes that bite An earlier fix that applies can fail a later draft; locked test rules, evidence, and the call stay recoverable Held-failure, false-hold, retiring, and test-lock fixtures
C-05 Human power A passing verdict cannot publish, send, spend, delete, disclose, or change terms without the separate permission required Negative permission tests at each high-stakes action line
C-06 Learning you can restart Interrupted work can resume from the ledger. Results update, reverse, or retire the right record without erasing history. High-stakes records carry both the date the claim applies from and the date the system wrote it, and both survive being overtaken, so a past run can be rebuilt against what the system believed while it ran Crash/restart, rollback, overtaking, and replay tests; a backdated-fix fixture proving a past run still gets judged against the beliefs of its own date, not today's
C-07 Master skills and adapters One shared master skill keeps one meaning, while app-specific discovery, permission, tool, and life differences stay inside versioned adapters Common-subset check, adapter diff, and skill-version paper trail
C-08 Model swap A typical workflow can change vendor or model version without moving or copying the owner's master records Same-fixture model-switch test, with the model paper trail recorded
C-09 App swap Each advertised app can run the same typical workflow against the same master store, with no app-specific fork of business truth Cross-app matrix covering Claude Code, Codex, Gemini CLI with Flash, Cursor, and GitHub Copilot for v1
C-10 Machine move A clean supported machine, or a new root path, can restore the stores, resolve links, find skills, and finish a governed run Export/restore and changed-path acceptance test, with integrity checks
C-11 Many-writer safety Several installed apps can read at once, while clashing or bypassing lasting writes cannot quietly clobber, interleave, or lose an accepted version Parallel-session lock, stale-write, all-or-nothing, direct-bypass, integrity, and clash-restart tests
C-12 Disclosure and paper trail Each run limits vendor access to approved records, and records the stable owner, app account, vendor, harness, model, adapter, skill, context, writes, verdict, and approval Allow/deny disclosure fixtures, and a full paper-trail check
C-13 Independent life stages Installing, upgrading, disabling, or removing one model or app adapter leaves master records and other supported adapters usable No-clobber, upgrade, uninstall, orphan, and cross-adapter regression tests
C-14 Honest degrading Missing abilities fail closed, or are declared unsupported. A prompt-only imitation is not reported as equal to a fixed control. The response follows the direction: a lasting write that cannot finish fails closed and claims no capture, while a read that cannot finish fails open and emits a visible degraded receipt naming what was not loaded Ability-probe tests, and a versioned supported/partial/unsupported matrix; a forced write-failure fixture proving no capture is claimed, and a forced read-failure fixture proving the run continues, is marked degraded, and can be told apart from a successful load of nothing
C-15 Capture with no separate act Running a governed workflow makes its run record, chosen context, locked test rules, verdict, override, and result with no separate owner act to write them. Interrupted and failed runs leave records too Run-completion test with no manual writing step; abandoned-run, failed-verdict, and crash fixtures showing a record still lands
C-16 Declared staleness Each store declares its staleness signal, its detection method, and its retiring path. A signal that fires makes a proposal, and cannot retire or rewrite a record on its own Seeded stale, unused, overtaken, and clashing fixtures; a detection-to-proposal trace; a negative test proving nothing retires by itself
C-17 Write power No lasting write proposed by an agent enters a store without an owner's call. The proposal is kept whatever the call, and the resulting record names the actor, the reason, and what it replaces Proposal-queue fixtures for accept, correct, reject, and clashing-proposal cases; a negative test proving direct model writes are unavailable or visible
C-18 Admission power Records arriving by restore, move, upgrade, or import get re-checked against current contracts. Their permissions are re-checked against grants the receiving store itself holds. Their claimed power is capped at what that store can prove alone. A check that cannot finish fails admission rather than dropping to a weaker one Import fixtures carrying ratified, owner-approved, and grant-bearing claims the receiving store cannot back up; a round-trip test proving the business's own export re-enters on the same terms as any other document; a negative test proving no bulk path lowers a check the single-record path enforces
C-19 Source admission Any source class other than the owner's own directly stated judgment enters only while a current grant covers it. That grant fixes why the material may be held, what may be done with it, how sensitive it is, how long it survives, and the dates it runs between. The grant is checked where the write happens, and expiry and withdrawal take effect there, not in config Allow and deny admission fixtures per source class; expired-grant and withdrawn-grant write attempts; a test proving withdrawal stops further entry without erasing evidence that past work relied on

This suite proves the build follows the stated design. On its own it proves no more. It does not show the design helps a business. It does not show each supported model writes as well. And it does not show buying beats a capable DIY build. Those are value claims, tested apart, below. The proof order is therefore:

  1. Match: was the system built as specified?
  2. Part performance: do fetching, rule reach, gates, restart, moving, and many-writer safety work on typical cases?
  3. Value against the rest: does Curio Chat Academy cut burden or lift results against the buyer's setup, and against a capable DIY build?
  4. Business result: does the governed workflow lift customer value, quality, capacity, or money in the real world?

The unit of testing

Test one repeating workflow that matters, in one real business. Pick work with source material, repeat fixes, a visible before-state, and a result you can see. Do not start with the easiest demo, or with a one-off creative task that has no stable test rules.

Before you change the setup, save a baseline set of typical work. Include ordinary cases, known repeat failures, exceptions, and cases that should pass. Record the current model and tools. Record the context supplied, the owner review effort, the escaped defects, the repeat fixes, and the restart friction. Record the customer or business result where you can see it.

The comparison

Compare the current setup with the split design. Hold the model, the task, and the source material as steady as you can. State the test rules and the decision rule in advance. Use held-out cases where you can, so the design is not judged only on the examples used to write its rules. Swap which side runs first. Learning or tiredness could otherwise favor whichever runs second.

Testing the purchase claim needs a second comparison. The buyer's old setup against Curio Chat Academy tests practical value. Curio Chat Academy against a capable do-it-yourself build of the same public needs asks a different question: does buying the base layer beat building it? Record setup, upkeep, support, moving, and exit effort on both sides. Without that second comparison, the evidence may show the design beats the old setup. It cannot show that buying beats building the same design.

One blended "AI quality" score would hide which part failed. So the test measures things one at a time:

Measure Question What it exposes
Context picking Did the run load the current facts that mattered, without pulling in unrelated or wrong-company material? Retrieval or addressing failure
Source tracing Can each high-stakes claim be traced to the source and the steps used? A summary mistaken for evidence
Rule reach Did an older fix that applied govern new work, without being taught again? A rule stored but not applied
Gate sensitivity Did the checkpoint hold both seeded and naturally occurring failures? A check that can only say yes
False holds Did the checkpoint reject work that met the test rules? Rules too broad or too old
Judge separation What evidence, model, runnable check, or person judged the draft, and how independent was its paper trail? Writer and judge sharing one error
Permission integrity Did publish, send, spend, delete, or access changes stop at the right human line? Passing work quietly gaining power
Escalation integrity Did a hold, a revision limit, or a crossed line route to the response stated in advance? A failure recorded but not contained
Capture completeness Did every governed run leave a record, including runs that failed, were abandoned, or fell in the worst week? A learning history that only records good days
Staleness detection Which overtaken, clashing, and long-unused records did the system surface, and how many were instead found by a draft, a customer, or nobody? Decay found downstream, not at rest
Write power Did every lasting store change pass an owner's call, and was the original proposal kept? Memory gaining the ability to write itself
Restart Can another session rebuild what happened, what remains, and what governs the next step? Workflow state trapped in chat history
Historical fidelity Can a past run be re-read against what the system believed on its own date, after later fixes have landed? Backdated edits flattering the record and hiding which gate really failed
Capacity fit Did the declared tier match real capacity, and did the downgrade or restart triggers work as designed? Idealized plans creating abandonment or hidden overwork
Owner review burden Which material still needed line-by-line owner review, and why? Checking cost eating the writing gains
Line upkeep What time and error did routing, curation, reconciling, and ledger upkeep add? Design overhead beating the rework it prevents
Cross-app fidelity Did every advertised app load the same records, call the intended skill contract, keep the same gates, and emit a full paper trail? A skill that parses everywhere but loses its meaning
Machine move Did the restored system find the same records and finish the same governed workflow from a different supported machine or root path? Portable files tied to one machine's paths or hidden state
Many-writer safety Did overlapping sessions keep every accepted version and surface clashes before commit? Quiet clobbering, interleaving, or last-writer-wins learning
Vendor disclosure Did each app get only the records approved for that vendor and task? Installing several agents becoming blanket permission to disclose the vault
Admission power Did records arriving by restore, import, or move have their claimed power re-established rather than accepted, and did any bulk path admit what the single-record path refuses? Power laundered through transport
Source admission Did each watched source class enter only under a current grant, and did expiry and withdrawal take effect at the write rather than in config? Material piling up with no stated basis to hold it
Degrading visibility When a load failed, did the run continue and say so? When a write failed, did anything claim a capture that never happened? A failed read reported as an empty truth
Business result Did the workflow change customer value, quality, capacity, or money? A tidy system with no business value

No single threshold fits everyone. The workflow, the stakes, and the baseline decide what counts as a real gain. Report both halves of each fraction: how many were caught, out of how many happened. Report false holds as well as caught failures. Report what got better, what stayed flat, what got worse, and what you could not measure.

What would change the design

The design should change if evidence shows that:

  • a store line adds routing and upkeep without improving retrieval, the paper trail, control, or restart
  • the hot view or addressed loading does worse than a simpler retrieval method
  • business facts and working fixes can be merged with no ambiguity or downstream failure
  • a separate judge adds cost without catching distinct errors at that level of stakes
  • the run ledger cannot be kept up for less than the rework and restart friction it prevents
  • an app adapter cannot hold the required context, gates, permissions, paper trail, or clash handling on its advertised surface
  • one shared writable store creates more clash and restart cost than a safer single-writer design can justify
  • automatic staleness detection raises more false flags than real catches, or costs more attention than the periodic review it replaces
  • the proposal queue becomes a second abandoned inbox. That would mean the owner's bottleneck was never the writing
  • admission re-checking refuses enough good records, or costs enough at restore, that owners route around it
  • the second date gets written on every record and never read. That would mean rebuilding the past answered a need nobody had
  • source grants are issued once at setup and never revisited. That makes them a note of intent, not a live control
  • by-product capture lowers the quality of what is kept enough to offset the runs it stops losing, or
  • another design produces better customer and working results under a fair comparison.

The current cache size, page size, load limit, review rhythm, and file layout are build choices awaiting measurement. None is a scientific finding. The claim that sets this apart is narrower. The system splits objects that have different power and different lives. It then keeps the evidence needed to test whether those splits earn their cost.

10. Worked example: a customer-facing offer test

Take a solo consultant with a service that works, and one repeating problem. AI can draft a landing page fast. But the owner keeps taking out promises with no backing. They keep putting back the words customers really use. They keep checking that the page still describes the real offer. The same three edits, on every draft.

The business constraint is that qualified prospects do not quickly see the value of the offer. "We need more content" names it wrongly. The next useful move is a small message test. A content system can wait.

This example only shows how the workflow runs. It makes no case-study or outcome claim.

Stage What you do
Orient Load the curated current offer and voice from the business vault. Pull customer language and claim support from the source shelf. Add the craft and fix records that apply. Record exactly what was loaded.
Choose State one guess: a clearer description of the customer result will help qualified prospects understand the offer. Decide in advance what evidence would support it, weaken it, or leave it open.
Specify Write a brief for one landing-page variant. Declare the standard capacity tier and the survival fallback. In a separate test-rule record, require every result claim to map to dated evidence, require the real offer terms, and set the voice checks. Lock the hold conditions and the revision limit before anything is written.
Produce Ask AI for options and one candidate draft. Keep the source page, record assumptions, and log the model, tools, context manifest, and draft version. Stop before publishing.
Verify Pull out the draft's stated and implied claims. Test them against the locked test rules and the source evidence. Record pass, revise, hold, or escalate. A hold names the failed test rule and routes the draft to the stated response. A pass does not by itself grant permission to publish.
Ship and measure After a separate owner approval, show the smallest responsible version to a fitting audience. Link the action and the draft version to what prospects understood, asked, and did during the watch window.
Capture and recalibrate Send useful customer notes to the source shelf and the curated vault. Send repeat process failures to corrections. Send skill-specific learning to expertise. Send the whole path to the run ledger. Retire rules that no longer apply, and decide whether messaging is still the constraint.

Four distinctions matter. A better brief improves the draft without replacing the checkpoint. The checkpoint protects the test, and customer evidence is still needed to settle it. Customer evidence updates the system while authorship stays with the owner. The four records also keep their separate jobs. The rule states the standing judgment. The test rules check this piece of work. The verdict records the result, and the permission controls the act.

The practical difference between a business built to produce and one built to learn is this. The second writes down what would prove its plan wrong, before the plan meets the market.

11. The scoreboard

Use a few measures that can change what you do. The exact measures depend on the business model. The groups do not.

Customer and business results

  • Qualified demand, or talks booked.
  • Conversion to sale, or activation.
  • Customers finishing, their result, repeat business, refunds, and referrals.
  • Cash collected, and margin after the cost to serve and to win the customer.

Workflow results

  • Time from a decision to a checked test a customer sees.
  • Redone work caused by thin context, a weak brief, or AI error.
  • Restart time after an interruption.
  • Open loops with no owner, no decision, and no next step.

Quality and risk results

  • Claims that matter, backed by enough evidence.
  • Defects that got past review.
  • Rules that applied and were caught by the checkpoint, and rules that were missed.
  • True holds and false holds, reported with their denominators.
  • Times a human overrode the system, and why.
  • Stale or clashing context that turned up.
  • Sensitive data loaded where a plain summary, with names removed, would have done the job. Note any approved exception.
  • Acts that matter, and whether they reached the right approval gate.

Design results

  • Records that mattered: loaded, missed, or loaded for no reason.
  • Claims that matter, traced from the work back to source and steps.
  • Repeat fixes that came back after you made them binding.
  • Runs you can restart without rebuilding the work from chat history.
  • Errors caught by outside evidence, by checks you can run, by another judge, or by the owner.
  • Time spent holding each line, against the redone work and review time it saved.

Owner results

  • Time spent talking to, or watching, customers.
  • Time spent on high-judgment and core-craft work.
  • Tasks the owner can no longer do or judge with confidence, without AI.
  • Patterns in energy and time that keep predicting failure, or a bounce back.
  • Drops to a lower tier, restart triggers, and time to get useful again.
  • New standing commitments that never named what they replaced, or where the time came from.

Do not make these your main success measures: tokens used, prompts written, agents deployed, automations scheduled, or pieces of content made. They are costs, or just activity, until they link to a result.

12. Failure modes and fixes

Each failure below names the signal that exposes it, and the fix that gives the business a way back.

Failure 1: AI activity outruns business learning

Signal: output rises while demand, conversion, experience, and money stay flat or unclear. A month of daily posts and no new conversations is one common shape. Fix: stop producing, name the constraint again, and run the smallest test a customer sees.

Failure 2: The system has no business truth to trust

Signal: AI keeps inventing, contradicting, or averaging the offer, audience, proof, or voice. Fix: write a short brief you can trust, a sourced vault, and a rule for settling clashes that the human owns.

Failure 3: Source material is allowed to instruct the system

Signal: a web page, email, transcript, or uploaded file changes something it should not. It can change rules, reveal unrelated data, or trigger a tool action, just because the model fetched it. Fix: treat sources as untrusted data. Keep them apart from rules and permissions. Give tools the least power you can. Test both hostile instructions and harmless material that looks like them. Labels help people and systems route content. They are not a complete security control.

Failure 4: Strategy gets smuggled into a prompt

Signal: the owner asks for campaigns, products, or funnels before deciding the result they want and the evidence they need. Fix: keep diagnosis, decision, brief, and production apart.

Failure 5: Advice gets mistaken for judgment

Signal: believable suggestions get acted on because they are specific, confident, or easy. Fix: require options, assumptions, evidence against the idea, a small test, and a human decision.

Behavioral evidence shows that people can lean too hard on advice from an algorithm, even when the context in front of them points elsewhere. How strong that pull is depends on the task and the design. Confident delivery should never stand in for checking. Klingbeil et al., Computers in Human Behavior (2024)

Failure 6: Permission to draft becomes permission to act

Signal: a tool publishes, sends, spends, deletes, grants access, or changes terms with no fresh approval. Fix: split preparing from doing, and put the human gate right before the high-stakes act.

Failure 7: The voice drifts toward the category average

Signal: content is polished and interchangeable with a rival's. Fix: start from your own observations, customer language, and a thesis the owner wrote. Use AI to challenge and express the idea, while the worldview stays the owner's.

In tests on short-story writing, AI raised judged novelty and usefulness, especially for less creative writers. It also made the set of stories more alike. The task is narrow, but the trade-off matters directly in a crowded online market. Doshi and Hauser, Science Advances (2024)

Failure 8: Automation freezes an immature process

Signal: exceptions, workarounds, and manual clean-up rise after you automate. Fix: go back to a manual assisted run. Define inputs, outputs, decisions, exceptions, tests, and ownership before encoding it again.

Failure 9: Memory compounds error

Signal: old assumptions quietly come back as current facts. A price that changed in spring and turns up in an autumn proposal is the usual case. Fix: attach a source and a date to high-stakes claims. Tell what was seen apart from what was concluded. Keep an add-only decision history. Give each store a staleness signal the system can spot. Do not rely only on a review the owner must remember. Then an overtaken record gets flagged at rest, instead of being found inside a draft that has already gone out.

Failure 10: The owner becomes the checking bottleneck

Signal: AI makes more drafts than the owner can responsibly review, and the backlog gets read late at night or never. Fix: produce less, write better briefs, add mechanical checks, raise the bar for what gets made, and ship fewer, higher-value tests.

Failure 11: The system eats the business

Signal: more time goes to tools, prompts, agents, and knowledge upkeep than to customers and delivery. Fix: remove roles and connections. Keep one inbox, one brief, one constraint, one workflow, and one review until the value shows.

Failure 12: Capture depends on the owner remembering

Signal: the learning loop looks whole in good weeks and is quietly missing in bad ones. The run ledger has gaps exactly where the hard work was. The recorded fixes are the ones the owner had time to write up. The review queue holds items from three weeks ago. Fix: move capture into the run, so that running a workflow writes its own record. Cut the owner's remaining task to ruling on what the run proposed. Some part of the loop may still need a separate act of remembering. Treat it as a part that will eventually stop running, and measure it that way.

Failure 13: The brain writes itself

Signal: entries appear in a lasting store that no person decided to keep. The vault gains a confident merged answer to a question the owner still considers open. An exception becomes a rule. A fix gets quietly rewritten to agree with a newer draft. Fix: keep the proposal apart from the record. Require a named call before anything becomes lasting. Look for an automatic write path added for ease during a busy stretch. Check too that the proposal queue has not just moved the review burden of Failure 10. This failure is the mirror image of Failure 12.

Failure 14: A failed read looks exactly like an empty store

Signal: work runs normally through a session in which the memory layer never loaded. Nothing errors. The draft reads well. The fix that applied is simply not there — the same shape as a business that never recorded one. You can only spot it afterwards, usually through a customer. Fix: make the response depend on the direction. A lasting write that cannot finish fails closed. A fix dropped in silence is the loss the design exists to prevent. A system that claims a capture it never made is worse than one that admits it stopped. A read that cannot finish fails open, so a memory outage does not stop the business working. It also emits a visible receipt naming what could not be loaded. An empty context nobody announces is the one failure that looks exactly like success. A blanket rule either way is wrong. Fail every path closed and the memory layer blocks all work. That is the pressure that gets a governance system uninstalled. Fail every path open and fixes evaporate on exactly the days the system is under strain.

Failure 15: Power arrives with the file

Signal: records hold a trust level nobody granted them. Trace it back and you find one of four things. A restore, a machine move, a re-import of the business's own export, or a file a model was asked to make. Integrity checks passed, because the bytes were intact, and whether the file had changed was never the question. Fix: treat admission as a write path, not as transport. Re-check every admitted record against current contracts. Re-check its permissions against grants the receiving store itself holds. Cap its claimed power at what that store can prove alone. Fail admission rather than lower a check to let a batch through. Then look for the convenient bulk path that was added for one move and left in place.

13. The hand-off contract and escalation path

A governed workflow needs a stated contract between human power, machine work, and outside evidence. In its strictest form: the owner decides, AI assists, and evidence updates the decision.

Inside an approved task, the system may do six things. It may summarize notes, sort source evidence, compare options, draft reversible work, apply stated checks, and turn open loops into next steps you can run. It may not choose the business's priority. It may not turn a passing check into permission to act. It may not widen scope or commitments without saying so. It may not hide real doubt, or push past the owner's stated capacity and fixed duties.

The contract turns the human line from a principle into a workflow rule:

Power Who holds it Record required
Name the constraint, the result, and the acceptable trade-offs Owner Constraint record and brief
Write or reshape inside the approved scope AI skill Run record, draft, and uncertainties
Judge the draft against locked test rules Named judge or runnable check Test-rule results and verdict
Publish, send, spend, delete, grant access, or change terms Owner, right before the act Approval tied to the exact draft version and act
Decide what customer or working evidence changes Owner, informed by what was seen Result, routing call, revision, or retiring

Subtraction before addition

New work eats a finite owner. The system may want to add a recurring task, workflow, connection, or automation. First the plan must name one of three things:

  • what it replaces
  • which existing commitment will stop or move
  • where the extra capacity has been shown

The point is to stop believable AI suggestions from growing the business faster than the owner can judge, check, or carry. Deleting one thing for every thing added, as a ritual, misses the point.

The same rule applies inside the design. A new rule should name the failure it prevents and the condition for retiring it. A new store line should name the confusion or control failure it settles. It should also say how its upkeep cost will be measured. Subtraction protects power. Stale rules, dead commitments, duplicate records, and workflows that no longer pay should leave the live path. Their history stays.

Constraint-first work

The input brief should start from the owner's real capacity. Record the time you have and the energy you have. Record the fixed duties that bound the work cycle: customer delivery, caring for others, rest. An idealized schedule records none of these. For planning, the workflow can expose three tiers:

Tier Meant for Shape
Survival Keeping going during disruption or low capacity A three-to-ten-minute manual path: keep one observation, take one immediate action, or leave one re-entry note
Standard Ordinary conditions A twenty-to-sixty-minute governed cycle that moves the current constraint through a brief, bounded work, and a check
Stretch Optional upside after a finished standard cycle Extra work that runs only when the standard result is done and the owner's fixed obligations stay protected

These times are defaults for an owner to calibrate from real capacity. No scientific finding sits behind them. What matters is that the tier is declared before production, and that the failure branch exists before it is needed. For example, two missed planned sessions can be a stated trigger to load the survival path and a small restart script. Another business may need a different trigger.

The escalation protocol

The checkpoint must do more than label a failure. It routes the run into a set response:

Level Trigger Required response
1. Hold and revise The draft fails one or more locked test rules Keep the failed draft and the verdict; attach the failed test rules and evidence; send only the affected part back to production; require a new draft version
2. Simplify and downgrade The stated revision limit is hit, the same failure repeats, or the workflow costs too much to check Stop extra agents, tools, and optional writing; move to the survival tier; ask one constraint-first question: What is the smallest responsible manual path around this failure?
3. Manual reset The writing layer is unavailable, unreliable, or crosses a high-stakes permission line Stop the automated path and escalate to the owner; record the observation in an approved record, choose one small manual test or recovery step, and close with a re-entry note

Two failed revision cycles are a useful default for Level 2. No universal threshold exists. A suspected breach of privacy, security, or a high-stakes permission does not wait for a revision count. It stops the path and escalates at once. The core business loop must keep working when the model, connection, or orchestration layer fails. A tool going obsolete should cost the business convenience and nothing else. Its memory, and the owner's ability to act, must outlive it.

Privacy and the least context needed

Pattern-first work usually needs the shape of a constraint, and none of the private detail inside it. When the name or the exact figures do not change the decision, prefer a short summary of the shape. "A deadline clash on Thursday" carries the decision without a person's name and private circumstances.

Passwords, access tokens, private keys, and other secrets do not belong in task context a model can see. Supply client, health, money, family, or other sensitive data only when three things hold. The task needs it. It is approved for that setting and use. And a stated policy covers access, keeping, redaction, and deletion. Taking out names lowers the risk. It does not by itself make data safe or nameless.

The least-context rule works with the line between source and instruction. Fetched material stays untrusted data. Sensitive sources get narrower access. Neither a source nor a summary gains permission to call tools or reveal unrelated records. NIST's Generative AI Profile names privacy risks and recommends actions. It says to strip personal data from paper-trail records. It says to measure access controls on paper-trail methods. It says to watch output for exposed sensitive data. And it says to set policies for what you collect and how long you keep it. It does not certify any particular local or cloud setup as safe. NIST AI 600-1 (2024), §2.4; MS-2.2-002; MS-2.7-005; MP-4.1-001 and MP-4.1-005

14. What the evidence supports

Findings and parts that are directly backed

Inside the settings they studied, the cited work supports these narrower claims:

  • AI systems hold more than the model. ISO, NIST, and W3C standards treat these as separate concerns: app context, data and input, model, task and output, life-stage activity, and human responsibility.
  • Outside non-parametric knowledge can improve some knowledge-heavy tasks. It also gives an update path and a paper trail that model weights alone do not.
  • A long input window does not promise reliable use of every relevant fact. Relevance and position can both change how well it works.
  • Content you fetch, or a user hands you, can carry hostile instructions. Today's defenses can trade blocked attacks for false refusals. They can also fail outside their training data.
  • Stated states and moves can improve control and results on some multi-step benchmarks.
  • A model correcting itself is not a reliable stand-in for outside feedback, on the reasoning tasks studied.
  • An LLM judge can favor its own writing, or that of models like it. Model errors can also run together.
  • AI can improve speed and judged quality on some bounded writing and knowledge tasks. Effects vary across tasks, users, and settings.
  • AI help can make work worse past the edge of what the model does well. A human-plus-AI pair does not reliably beat the stronger of the two alone.
  • Strict testing of a guess can sharpen business decisions. In the field trial cited here, AI business advice alone produced no average gain you could detect.
  • In a large panel, young U.S. factory plants gained more than others from rented IT early on. Owned IT offered a different, slower path to firm-specific strength.
  • Make-or-buy evidence supports matching the setup to the deal. It does not support a blanket preference for buying or for owning.
  • Industry studies give fair evidence that software reuse lifts quality and output. Its costs, and its effects in small firms, are far less settled.
  • In the very small firms studied, how the owner spent time either kept or damaged the firm's ability to adapt.
  • Official docs for Claude Code, Codex, Gemini CLI, Cursor, and GitHub Copilot show support for file-based Agent Skills, or for finding skills the same way. They do not show that the products run them alike.
  • A platform can narrow bulk-read access by policy. In May 2025, Slack's changelog cut two chat-history methods to one request per minute for non-Marketplace apps. New terms also limited how others may store and reuse the fetched data. This shows the access exists at a platform's choice. It forecasts nothing about other platforms, and says nothing about model vendors.

None of those findings validates the full store design or workflow in this paper.

Careful design guesses

These design choices join backed parts with stated control reasoning:

  • keep model behavior apart from owner-held business knowledge
  • keep the harness apart from the model, and record which one ran
  • a harness sets what a run can reach and what a gate can stop
  • keep raw sources apart from curated facts and owned decisions
  • do not give source content the power of rules, system instructions, or tool permissions
  • keep lasting rules apart from task test rules, from recorded verdicts, and from permissions to act
  • keep workflow state and run paper trail outside chat history
  • lock test rules before the draft is written or seen
  • raise judge and evidence independence as the stakes rise
  • require fresh human approval right before a high-stakes outside act
  • require an owner's call before a proposed record becomes lastingly true, and keep the proposal
  • make run records a by-product of running, not a separate act the owner must remember
  • give every store a declared staleness signal, detection method, and retiring path, and keep retiring a decision someone answers for
  • treat the owner's unwritten reasoning as input you must draw out. Treat what the interview produces as a dated source, not a settled fact
  • buy the shared plumbing that sets no one apart. Build your own truth and edge in portable files you hold
  • keep one master owner-held data plane across model vendors and AI apps
  • treat context control as a design requirement, because a business memory that rests on another company's access policy is not truly owned
  • keep app-specific behavior inside thin, versioned adapters, leaving business truth unforked
  • serialize or clash-check lasting writes from overlapping agent sessions
  • record two dates on a high-stakes claim: the date it applies from, and the date the system wrote it. A past run can then be judged against the beliefs it ran under
  • treat restore, move, and import as write paths that must re-establish power. They are not transport that keeps it
  • admit a watched source only under a current grant, checked where the write happens. The owner's own stated judgment is the one source that needs no grant.
  • let the failure response follow the direction: fail a lasting write closed, and a read open with a visible degraded receipt.

These are controls you can defend. None is a measured effect size for a one-person business.

Product claims still open to revision

These claims must be tested in use:

  • six logical stores beat a simpler or a finer-grained design
  • a hot view plus addressed pages improves reliability enough to justify the upkeep
  • business facts and working fixes need separate stores, and typed records in one store would not do
  • owner-specific expertise should stay apart from the general craft base
  • a run and judging ledger cuts restart, rework, and owner review by more than it costs to keep
  • a separate judge catches distinct high-stakes errors often enough to justify its cost
  • picking one binding constraint beats tuning several business functions at once
  • the daily orient-focus-close rhythm and the weekly reality review improve restart and continuity
  • stated survival, standard, and stretch tiers keep you going without lowering customer or quality results
  • the rule of subtract-before-you-add cuts commitments you cannot keep, without smothering useful tests
  • a stated escalation path cuts recovery time and permission failures
  • a process should be run by hand before it is scheduled or made autonomous
  • capture made while running beats a capture stage the owner does afterwards. It must do so without lowering the quality of what is kept
  • automatic staleness detection finds enough truly decayed records, with few enough false flags, to repay its upkeep for one person
  • a proposal queue costs the owner less attention than writing records by hand. It must not become a second inbox that is itself abandoned
  • structured interviews recover unwritten owner reasoning well enough to store. They must not make beliefs the owner did not hold before being asked
  • one master Agent Skill, plus thin adapters, keeps the workflow meaning across all five apps: Claude Code, Codex, Gemini CLI with Flash, Cursor, and GitHub Copilot
  • one owner can use several supported apps against one writable store, without too much clash, disclosure, or upkeep
  • a second date on every high-stakes record repays its cost by making past runs reviewable. It must not be a field written and never read
  • source grants stay a live control for one person, where the person granting and the person gaining are the same. They must not decay into a form filled in once at setup.
  • a bought Curio Chat Academy base layer beats the buyer's own do-it-yourself setup. It would have to cut one of these enough to matter: repeat fixes, escaped defects, restart friction, owner upkeep, or time to the first governed workflow.

The design should change when a fair test proves one of these wrong. Adding stores is not progress by itself. Fewer lines are better, so long as they keep the same power, paper trail, control, and restart.

Claims this paper rejects

  • Some one prompt, model, agent design, or time split guarantees business success.
  • Research proves most solo owners should buy Curio Chat Academy, or that buying always beats building.
  • RAG research scientifically validates this vault or store design.
  • "Lost in the middle" sets a universal page count, context limit, or hot-cache size.
  • A source label or delimiter makes fetched content safe to follow as instruction.
  • A second prompt or agent name makes a judge independent of the writer.
  • A rule is enforced because it appears in an instruction file.
  • A learning loop is working because its records look complete.
  • A model can decide which of two clashing accounts the business should treat as true.
  • An interview transcript is proof of what the owner knows, rather than a dated record of what they said when asked.
  • Plain Markdown alone makes a workflow portable across models, apps, or machines.
  • A skill that parses keeps the same gates, permissions, and failure behavior in each AI app.
  • AI can work out the right business strategy on its own.
  • More automation, context, agents, or output always makes more worth.
  • A fixed human/AI work split is scientifically settled.
  • The tier times, or the two-revision default, are settled by science.
  • Marketing claims written by AI are safe because a human glanced at them.
  • Reviews, or sources produced by AI, are evidence on their own.

For businesses serving U.S. consumers, the Federal Trade Commission's baseline is plain. Ads must be truthful and not deceptive. A claim of fact needs enough evidence before you publish. Stated and implied claims both count. Rules vary by place and category, so this is a working baseline and not legal advice. FTC, Advertising FAQs: A Guide for Small Business

NIST's Generative AI Profile treats risk as work across a whole life: governing, mapping, measuring, and managing. One output check cannot cover that. It is voluntary U.S. guidance, and nothing like a trial of a solo-owner workflow. It backs stated lines, testing, docs, privacy care, and ongoing watching. NIST AI 600-1 (2024)

15. Conclusion

A solo owner has one built-in edge: a short distance from seeing, to deciding, to doing, to hearing back from customers, to changing.

AI can shorten parts of that distance. It can search faster, reshape more material, make more options, and run steady patterns. It can also add believable error and tempting advice. And it can add overhead, sameness, and dependence.

The useful system is human-led and closed by evidence. That is a different thing from "human only" or "AI first." It remembers what is true. It picks what constrains the business. It briefs the work and hands it over inside bounds. It checks in step with the stakes, ships to reality, and keeps what reality teaches.

That is how a one-person business can grow more capable without becoming less itself.

The buying rule follows the same logic. Buy the shared plumbing that sets no one apart. Build your own truth and edge in portable files you hold. Curio Chat Academy exists to make that second asset compound, without trapping it inside the first.

In the v1 control test, one owner uses five apps on one machine against the same master stores and skills. The five are Claude Code, ChatGPT/Codex, Gemini CLI with Flash, Cursor, and GitHub Copilot. It must do that without copying the business. It must not weaken the gates, lose accepted writes, or give every vendor access to every record. A later model, app, or machine should need a checked adapter, or a checked move. Rebuilding the business's memory from scratch is what the design exists to prevent.

For the smallest place to start, The Three Fixes applies the rule-hold-retire loop to one correction. It does not pretend to install the full design.

Appendix A: Public design record list

This list shows the shape of the design's decisions. It does not publish build schemas, file layouts, automation logic, test-rule libraries, or filled-in working records. A record may be a plain file, a database row, or part of a larger document. Its format matters less than this: can you recover its power, its paper trail, and its life stage?

Each high-stakes record below carries two dates: the date its claim applies from, and the date the system wrote it. Both outlive the overtaking or retiring that ends the record. Where a row names an effective date, a recorded date is implied beside it.

Record Public purpose It must establish Deliberately not shown here
Source record Keep an observation or received file without turning it into business truth Origin, date received, scope, integrity or version, trust status, and any limit on use Intake rules, metadata schema, redaction pipeline, and storage layout
Business fact or decision State what the owner currently treats as true or decided The claim, owner, source or basis, effective date, confidence where it matters, and whether it has been overtaken Vault page design, precedence algorithm, and reconciling procedure
Rule record Keep a fix that eligible work can break Requirement or ban, scope, reason, owner, live state, exception path, and retiring condition Rule grammar, rule library, and enforcement build
Workflow run Make process state recoverable outside chat history Business decision served, current stage, capacity tier, fallback path, chosen context manifest, harness and version, model and tools, draft versions, moves, and next step Run schema, orchestration code, and model-specific adapters
Test-rule set Turn intent and applicable rules into a fair test for one piece of work Scope, test rules, evidence method, severity, hold condition, revision limit, and proof the set locked before writing or review Test-rule templates, scoring logic, and private check libraries
Verdict Record whether the draft passed, needs revision, is held, or needs escalation Draft version, test-rule results, evidence, judge paper trail, uncertainty, the call, and any required escalation state Judge prompts, hidden test cases, and build thresholds
Approval and action Keep a successful check apart from the power to act Named approver, permitted act, exact draft version, time, scope, and what happened Identity-control build and connection credentials
Result and learning Link the work to customer, quality, capacity, or money evidence Prior guess, watch window, denominator, result including flat/worse/unmeasurable, decision, and destination store Product analytics build and internal evidence registers
Proposed write Hold a candidate record between writing and durability, so making it and accepting it stay separate events Proposed content, target store, producing run and model, evidence offered, what it would replace, the call, the deciding actor, and the date Proposal queue build, ranking, and decision interface
Staleness flag Surface a record that may no longer match reality, without acting on it Record affected, signal type, detecting mechanism, evidence of the mismatch, date raised, and the call, including dismissal as a false flag Detection thresholds, decay functions, and drift-comparison build
Retiring or overtaking Let the system forget on purpose without erasing history Record affected, reason, evidence or owned decision, replacement if any, actor, and effective date Upkeep automation and internal review rhythm
Admission record Show what entered a store through a bulk path, and on what basis each record was accepted rather than trusted Origin file and its integrity result, counts admitted and refused, which contracts were re-run, which permissions were re-checked against locally held grants, the power level each record was capped at, and the deciding actor Import build, validator internals, and the hostile fixtures used to test it
Admission grant Let a source class other than the owner's own stated judgment enter any store, for a stated reason and period Which source class; why it may be held; which operations it permits; how sensitive the material is; how long it is kept; the dates the grant runs between; who granted it; and whether it has since been withdrawn, with the date Grant storage, enforcement hooks, and connector credentials
Conformance result Show whether a named release implements this paper on an advertised app and machine Fixed paper version, requirement ID, product and adapter versions, setting, test or demo, result, evidence location, date, and any stated gap Private build code, hidden hostile fixtures, and licensed skill content

The list is deliberately not enough to recreate the build. Its purpose is to make the paper's claims inspectable:

  • source is not fact
  • fact is not rule
  • rule is not test rule
  • test rule is not verdict
  • verdict is not permission
  • permission is not result
  • a proposal is not a record
  • a staleness flag is not a retiring
  • a checked file is not an approved record
  • a date of effect is not a date of record

Appendix B: Evidence register

Accessed August 22–23, 2026. E35 and E36 were accessed August 28, 2026.

ID Source and status Population or scope Finding used here Limit
E1 Noy & Zhang, Science (2023), peer-reviewed randomized trial 453 college-educated professionals on mid-level writing tasks ChatGPT cut time by 40% and raised judged quality by 18% Specific writing tasks, participants, and 2023 tool; not business performance
E2 Brynjolfsson, Li & Raymond, QJE (2025), peer-reviewed field study Staggered rollout to 5,172 customer-support agents 15% average productivity gain, with uneven effects; biggest gains among less experienced and lower-skilled workers One firm, one job, one specialized assistant
E3 Dell'Acqua et al., Organization Science (2026), peer-reviewed pre-registered trial 758 knowledge workers on consulting tasks Gains on 18 within-edge tasks; on one outside-edge task, control correctness was 84.5% against 60.0% and 70.6% in the AI groups, a 19-percentage-point average drop One outside-edge task built with one model generation; the edge moves over time
E4 Vaccaro, Almaatouq & Malone, Nature Human Behaviour (2024), pre-registered systematic review and meta-analysis 106 studies, 370 effect sizes, from 2020 to mid-2023 Human-AI pairs beat humans alone on average, but not the better of human or AI alone; strong variation by task Possible publication bias; varied designs; much evidence predates current AI
E5 Otis et al., Management Science (2026), peer-reviewed field trial 640 Kenyan business owners randomized to a GPT-4 adviser No detectable average revenue or profit effect; weaker performers did worse, stronger performers may have gained One country, one intervention; subgroup readings need care
E6 Camuffo et al., Management Science (2020), randomized controlled trial 116 Italian startups over about one year Scientific-style prediction and hypothesis testing sharpened business decisions Early-stage startups; not AI-specific, and not a trial of a work rhythm
E7 Liu et al., TACL (2024), peer-reviewed controlled model evaluation Multi-document QA and key-value retrieval across several language models Where relevant information sits can materially change long-context performance Tested models and tasks differ from current production systems
E8 Doshi & Hauser, Science Advances (2024), peer-reviewed experiment Short-story ideas and writing AI raised individual judged creativity while cutting the variety of the collection Narrow creative task; not brand strategy or long-run market difference
E9 Lee et al., CHI (2025), peer-reviewed survey study 319 knowledge workers, 936 self-reported examples Trust in AI and trust in oneself were linked to different reported critical-thinking effort Observational, self-reported, not causal
E10 Klingbeil et al., Computers in Human Behavior (2024), peer-reviewed incentivized experiment Behavioral trust and reliance tasks People can lean too hard on AI advice despite relevant context in front of them Behavioral task setting; the size should not be generalized
E11 NIST AI 600-1 (2024), official voluntary framework Cross-sector generative-AI risk management Supports life-stage governance, testing, docs, privacy-risk watching, stripping personal data from paper-trail records, measuring access control on paper-trail methods, and collection and keeping policies Guidance, not causal evidence, not a legal safe harbor, and not certification of any build
E12 FTC Advertising FAQs, official U.S. business guidance Advertising to U.S. consumers Stated and implied advertising claims must be truthful, not deceptive, and adequately backed before use U.S. baseline; category and place can add rules; not legal advice
E13 ISO/IEC 22989:2022, international terminology standard AI concepts and terms Tells a built AI system apart from a model used inside a system to stand for data, knowledge, or process Terminology standard; does not prescribe this design or establish effects
E14 NIST AI RMF 1.0 (2023), official voluntary framework Cross-sector AI life stages and risk management Splits app context, data and input, AI model, and task and output; recommends life-stage testing and best-practice separation of build/use from checking actors Guidance; no solo-owner trial; version 1.0 was under revision at the access date
E15 Lewis et al., NeurIPS (2020), peer-reviewed empirical study Knowledge-heavy NLP tasks using parametric and fetched non-parametric memory RAG beat the weights-only baselines studied, and created a clear path for update and paper trail Wikipedia-based research design; does not validate a business vault or file-store design
E16 Wu et al., COLM (2024), peer-reviewed benchmark study InterCode SQL and ALFWorld StateFlow split process grounding, through states and moves, from actions inside states; reported more success at lower cost than ReAct Two benchmark families; human-designed state machines; not a business-workflow trial
E17 Liu et al., USENIX Security (2024), peer-reviewed security study Five prompt-injection attacks, ten defenses, ten LLMs, seven tasks Formalized prompt injection and found real attack/defense variation across the systems tested Adversarial benchmark; does not prove one universal defense, or that source labels are safe
E18 Yi et al., KDD (2025), peer-reviewed security study Indirect prompt injection across several app scenarios Introduced the BIPIA benchmark and tested black-box and white-box defenses, backing a clear line between outside content and trusted instructions The scenarios and models tested; defense performance does not establish general safety
E19 Li et al., ACL (2026), peer-reviewed security study Two base models and several prompt-attack defense pipelines Found position, token-trigger, and topic-generalization shortcuts that caused false refusals and out-of-distribution accuracy losses Specific supervised defenses and diagnostic datasets; not every defense or deployment
E20 Huang et al., ICLR (2024), peer-reviewed empirical study Self-correction on the reasoning datasets and models studied Self-correction without outside feedback did not reliably improve reasoning, and sometimes made it worse Reasoning tasks only; the paper and its reviews warn against a universal reading
E21 Panickssery et al., NeurIPS (2024), peer-reviewed empirical study LLM evaluation and self-recognition experiments Found that LLM judges can recognize and favor their own writing over equally rated alternatives The models and evaluation designs studied; does not make all model review useless
E22 Goel et al., ICML (2025), peer-reviewed empirical study Model similarity, LLM-as-judge evaluation, and weak-to-strong supervision Preference for similar models went beyond exact self-evaluation; model mistakes grew more alike as capability rose Oversight evaluations, not this workflow; no solo-owner effect size
E23 W3C PROV-O (2013), W3C Recommendation Interoperable representation of provenance Splits entities, activities, and agents, and represents use, derivation, credit, revision, and responsibility A provenance standard; does not prescribe a file system, a store count, or an evaluation method
E24 Jin & McElheran, Management Science (2025), peer-reviewed administrative-data study About 240,000 U.S. factory plant-year records from 2006–2014, including about 41,300 from plants five years old or younger Young plants gained more than others from renting modern IT; renting supported early output and survival, while owned IT offered firm-specific strength over time Factory plants, not solo owners or AI systems; sourcing modes and technology differ from this product
E25 Geyskens, Steenkamp & Kumar, Academy of Management Journal (2006), peer-reviewed meta-analysis Empirical make, buy, and ally studies published or obtained through 2003 Found strong support for matching the setup to the deal; no universal make-or-buy winner Cross-industry boundary research; does not name the best AI product or a solo-owner threshold
E26 Daneshgar, Low & Worasinchai, Information and Software Technology (2013), peer-reviewed mixed-method study Semi-structured interviews and factor ranking in eight Thai SMEs Fit to needs, cost, size and complexity, time, in-house skill, support, and daily factors shaped software buying Small qualitative sample in one country; the factors do not establish a majority preference for buying
E27 Chen, Usman & Badampudi, Information and Software Technology (2024), peer-reviewed systematic literature review 30 industrial primary studies of software-reuse costs and benefits Evidence strength was high for better quality and higher productivity; systematic and verbatim reuse reported more benefit Benefits were studied more than costs; most primary studies were low or moderate quality, and only three concerned small companies
E28 Kevill et al., International Small Business Journal (2021), peer-reviewed qualitative study Three in-depth micro-enterprise case studies How the owner-manager spent time was a core micro-foundation of dynamic managerial capability, and could make capability fragile Qualitative mechanism evidence; does not price the opportunity cost of building software, or compare buy with build
E29 Costa Silva, Rose & Calinescu, IEEE CloudCom (2013), peer-reviewed systematic review 721 cloud-lock-in primary studies screened; 78 selected for detailed analysis Gaps in interfaces, technologies, and semantics hinder portability and interoperability, and can create vendor lock-in Cloud infrastructure literature; many proposed remedies lacked strong empirical validation, and do not directly test local file-based products
E30 OpenAI Codex skills and AGENTS.md instructions, official product documentation ChatGPT and Codex skill packaging; Codex project-instruction discovery Documents Agent Skills as an open-standard workflow format, and layered AGENTS.md discovery for Codex Product docs can change, and do not prove semantic equivalence with another app or with any Curio Chat Academy release
E31 Claude Code skills, official product documentation Claude Code skill packaging and discovery Documents SKILL.md-based skills, and describes Agent Skills as an open standard used across several tools Claude-specific extensions and behavior still need compatibility and conformance testing
E32 Gemini CLI Agent Skills, official Google Gemini CLI repository documentation Gemini CLI user and workspace skill discovery Documents Agent Skills and shared .agents/skills/ aliases, alongside Gemini-specific locations Repository docs and Gemini CLI behavior are version-sensitive; picking Flash does not by itself establish workflow conformance
E33 Cursor Agent Skills, official product documentation Cursor skill discovery, packaging, and compatibility paths Describes Agent Skills as a portable open standard, and documents discovery from shared and app-specific skill folders A supported file format does not establish equal permissions, hooks, tools, or results
E34 GitHub Copilot agent skills, official product documentation GitHub Copilot agent-skill packaging, discovery, scripts, and permissions Documents SKILL.md skills and support for project skill folders, including .agents/skills/ Support varies by Copilot surface and can change; each advertised surface needs its own conformance result
E35 Slite, The Ontology of the Company Brain (2026), 41-page PDF edition retrieved August 28, 2026; access requires the email form at slite.com/ebooks/company-brain, vendor-published survey and practitioner interviews, not peer-reviewed 149 survey respondents (p. 3), plus interviews with builders of shared-context systems (p. 34); sampling frame not described; questions not pre-registered 17% reported a working system in production (p. 29); about half of those who tried had custom-built on Claude or Codex (p. 29); upkeep rather than construction was the usual reason working systems were abandoned (p. 34); practitioners reported that capture depending on human discipline decays, that human review comes before writes to memory in every system described, and that most operational knowledge stays unwritten The publisher sells in the category and has a direct commercial interest in the buy-versus-build answer. The enterprise-pilot figure describes that vendor's pipeline, not the market, and neither the number of pilots behind it nor the number of respondents who had tried to build a system is disclosed. The full text is gated behind an email form, and the publisher's landing page describes the same population as "149 teams" while the ebook text says "149 people"; the neutral "survey respondents" is used here because the source contradicts itself. Used only for direction on upkeep risk, capture decay, write gating, and the unwritten-knowledge gap; no number from this source is relied on
E36 Slack API changelog (2025), official platform documentation conversations.history and conversations.replies access for non-Marketplace apps, effective May 29, 2025 New unlisted non-Marketplace apps, and new installs of existing unlisted apps, limited to fifteen messages per request at one request per minute, against 1,000 messages per request at fifty or more requests per minute for internal customer-built apps; accompanying terms restrict third-party storage, use, and sharing of data obtained via these APIs; bulk downloads for indexing and querying named as the risk addressed One platform, two methods, one date; existing installs of distributed apps unchanged. Shows that platform context access can be narrowed by commercial policy. Forecasts nothing about other platforms, and is not evidence about model vendors.

Appendix C: Method and claim limits

This paper draws on four kinds of material: peer-reviewed research, official guidance, documented working frameworks, and practical build records. They carry different kinds of weight. Research backs narrow findings in the settings studied. Official guidance backs care with risk and with claims. The full workflow is an open synthesis. No single study has tested it as one thing.

Scientific limit

Nobody has tested the full design as one thing. That design covers the business map, the store lines, and the seven-stage workflow. It also covers the rhythm, the capacity tiers, the escalation path, the build order, the moving profile, and the scorecard. No paper in the evidence register tests the Curio Chat Academy business vault or the six-store proposal as a package. No paper tests whether buying Curio Chat Academy beats building an equal system. Official vendor docs show which surfaces existed on the date we looked. They do not show equal behavior, safe shared use, or a Curio Chat Academy match across those surfaces. "Evidence-informed" means a part or a control fits bounded evidence and standards. It does not mean science has validated the design, or the choice to buy.

The design review used three warrant levels:

  1. Direct support: the cited source studies or defines the part used in the claim, within a stated scope.
  2. Careful guess: the paper proposes a control because records differ in power, paper trail, life, or failure cost. The source backs the part, but not the whole control or its effect size.
  3. Product claim: the exact store line, build choice, rhythm, or threshold must be tested in real use.

Named evidence gaps

  1. We found no study that tests a full file-based AI operating system for skilled solo owners, against their current prompt, memory, or second-brain setup.
  2. We found no study comparing a typed multi-store vault with one single knowledge base, while holding model, task, and source material steady.
  3. We found no evidence that sets any of these for everyone: hot-cache size, page size, context-load limit, review rhythm, tier length, or revision limit.
  4. We found no study that measures the specific benefit of splitting owner business facts from working fixes.
  5. We found no study on whether splitting owner-specific expertise from the general craft base helps. That would mean easier moving, better upgrade survival, or better work.
  6. We found no study on what a run and judging ledger saves in a one-person business. That would mean less rework, restart time, escaped defects, or owner review.
  7. We found no study that gives a portable effect size for judge-independence on customer-facing solo-owner work.
  8. We found no long-run study showing that the seven-stage workflow, or the daily and weekly rhythm, improves customer value, business money, or owner control.
  9. Prompt-injection studies justify a line between source and instruction. They do not give a complete defense for any file-connected desktop AI system.
  10. The field evidence on AI for business owners cited here comes from one intervention and one population. It does not establish who will benefit from this design.
  11. We found no study that estimates what share of solo owners should buy such a system, and what share should build one. None compares the current Curio Chat Academy packs with an equal do-it-yourself build. The only on-topic evidence we found is a practitioner survey published by a vendor. It does not describe its sampling frame, and its publisher has a direct stake in the answer. It points to where the risk of abandonment sits, and supplies no usable size.
  12. We found no study showing that a shared Agent Skill behaves the same across Claude Code, Codex, Gemini CLI, Cursor, and GitHub Copilot. That covers retrieval, following instructions, permissions, tools, and failure.
  13. We found no study on several AI apps writing one owner's master file stores at the same time — neither its safety nor its value. This is a question for engineering tests and product trials.
  14. We found no study on whether capture made while running beats a separate capture stage, or on what it costs in record quality. Practitioners report that capture based on discipline decays. They supply no rate, no denominator, and no comparison.
  15. We found no study that sets a staleness-detection method, threshold, decay function, or acceptable false-flag rate for a one-person business. None shows that automatic detection repays its upkeep against a periodic human review.
  16. We found no study on whether structured interviews recover unwritten owner reasoning well enough to store. None counts how often an interview makes a belief rather than finds one. Practitioners report this as unsolved at scale.
  17. We found no evidence establishing how much of a solo owner's working knowledge is unwritten. The often-repeated figure of roughly 80% traces to practitioner assertion, not measurement, and is not relied on here.
  18. We found no study on whether recording both dates on business records repays its upkeep for one person. None counts how often people rebuild the past once they can.
  19. We found no study on whether source grants stay meaningful in a one-person business. There, the person granting and the person relying on the grant are the same. The usual outside check on a consent record is therefore missing.
  20. We found no study comparing failure responses for a personal memory layer. So why should a failed write fail closed, and a failed read fail open with a visible receipt? Because the two silent outcomes cost very different amounts. That is a control argument, not a measured result.

Considered and not used

  • "AI memory makes fixes permanent": rejected. Lasting, fetching, applying, judging, and enforcing are five different things.
  • "RAG research validates the Curio Chat Academy design": rejected. The RAG studies test retrieval designs and tasks, not these business stores.
  • "Lost in the middle proves shorter prompts are always better": rejected. The study shows that position matters in the tasks and models tested. It sets no universal best context size.
  • "A second model is independent": rejected. Models from one family, fed shared data, shared prompts, and shared sources, tend to make the same mistakes.
  • "A human in the loop makes the workflow safe": rejected. Four things decide whether it is a real control: where the human sits, what they know, what they can do, and what power they hold.
  • "A rule in a file cannot be forgotten": rejected. A stored rule can fail to load, fail to apply, be overridden, be contradicted, or be judged badly.
  • "Plain files make the system portable": rejected. After the files move, much can still fail: finding paths, matching schemas, finding skills, permissions, tools, writing at the same time, and life-stage behavior.
  • "Agent Skills support makes apps equal": rejected. A shared packaging standard does not even out model behavior, which instruction wins, permissions, hooks, tools, or failure handling.
  • "More stores make more reliability": rejected. Each line adds routing and upkeep cost, and must earn it in practice.
  • "A successful draft validates the system": rejected. One output shows almost nothing. It cannot show how well the system fetches, how well the gate catches, or how often it holds good work. It cannot show whether permissions hold, whether it restarts, or what it is worth.
  • "A checked backup is a trustworthy source": rejected. An integrity check shows a document has not changed. It never shows that the power the document claims was granted.
  • "Consent is a setting": rejected. A permission recorded beside the writer, and never checked at the write, is documentation. That is the same objection the paper makes to a rule in an instruction file.
  • "A system that never errors is working": rejected. A read that fails in silence looks just like a store that holds nothing. Only one of them is a fault.

Design-change rule

The public design is built so it can be proved wrong. Merge, split, or redesign a line when a fair test shows a better layout. Better means it keeps power, paper trail, control, and restart at lower cost, or gets better results. New science can change the design before any product evidence exists. It might reveal a failure the current lines cannot hold.

Any design comparison should state six things in advance:

  • its workflow
  • its population
  • its test rules
  • its hold conditions
  • how it measures owner review
  • its decision rule It should keep false holds and negative results. It should name the model, the tool, the context, and the judge's paper trail. Otherwise the design can "improve" merely because the test moved.

Upkeep limit

Review the evidence base at least once a year. Review it again whenever a big claim, a rule, a model's abilities, a prompt-attack result, or a test method changes. A number reused here must keep its population, task, tool, date, and denominator. Changes should update both the public claim and its evidence-register limit. A new model feature alone is not evidence that a design claim has passed.

The worldview this paper turns into an operating mechanism is the CurioChat manifesto.

← Back to white papers  ·  Read the proposition summary