We build software with coding agents, AI that writes and tests code on its own for long stretches while a person reviews the result. When these agents got good enough to work for hours without help, software teams learned something the hard way. The model (the AI a lab such as Anthropic or Google trains) was rarely the problem, in our experience. An agent with vague instructions and broad access writes a great deal of confident code in the wrong direction, and sometimes it breaks things on the way.
Anthropic's guide to building agents names three needs (clear success criteria, guardrails, and pausing for a person at checkpoints), and that matches what we see in our own work. We call them the three Cs: clarity, constraints, and checkpoints. Every AI role a manufacturer adds to its org chart needs the same three. When an AI project stalls, one of them is usually missing.
1. Software teams wrote the three Cs first
A coding agent such as Claude Code starts each session by reading a project file its team wrote for it: what the project is, how the code is organized, which commands to run, and what good work looks like. Claude Code calls this a CLAUDE.md file. That file is clarity.
The agent also runs under a permission list. The team decides which actions it may take on its own, which ones need a person to approve, and which ones it may never take. In Claude Code, these permission rules are checked in a fixed order: deny first, then ask, then allow. That list is constraints.
Then the agent's work goes through checkpoints before it reaches anyone. Automated tests have to pass, a person reviews the change before it is merged into the product, and the tools keep a history so a bad change can be rolled back. Claude Code even names its rollback points checkpoints, and Anthropic's own guide to building agents says agents "can then pause for human feedback at checkpoints." Google says 75% of its new code is now AI-generated and approved by engineers.
None of this is new management thinking. What changed is that software teams had to write it down, because an agent cannot pick up the unwritten rules by watching. Your AI roles cannot either.
2. Clarity: write the job down before the AI starts
Clarity works at two levels, and both need to be on paper.
The first is the company's direction: what you want AI to attack first and why. A fabricator that loses jobs because quotes take four days should start with quoting. A repair shop drowning in cert requests should start there. Picking one target, in order, keeps the effort from spreading across ten half-finished experiments.
The second is the role itself. Write its job in one sentence, then say what a good result looks like with real examples. For a quoting role, clarity sounds like this: "Drafts quotes for repeat parts and simple machined work, priced from our cost tables and past jobs, ready for an estimator to review within two hours of the RFQ arriving." Attach ten past quotes your best estimator considers good. Those examples teach more than any paragraph of instructions.
If your team cannot write that sentence, the role is not ready. We wrote about this in Software Got Cheap, Clarity Didn't: the expensive part of an AI project is getting the job out of people's heads and onto paper.
3. Constraints: enforce limits with a login, not an instruction
Constraints are the things you will not allow, written as a short "never" list. For the quoting role:
- It never sends anything to a buyer.
- It never prices below the margin floor set by the sales manager.
- It never changes a record in the ERP. It only drafts.
- It never sees payroll or HR files.
- It never runs on an account that lets the vendor train on your data. Anthropic's business products do not train on your inputs by default.
Software teams learned that a written instruction is a request. Anthropic's own documentation says its agent treats the project file "as context, not enforced configuration," and tells teams to use a hard rule in the software when an action must be blocked every time. So you do not ask an agent to stay out of the live customer database. You give it a login that cannot reach it.
The incidents that taught this were public. In July 2025, Replit's AI coding agent deleted a customer's live database during a declared code freeze. Replit's CEO called it "unacceptable" and began separating test and live data so it could not happen again. The same month, Amazon reported that an over-broad access token (a digital key with more access than the job needed) let an attacker slip malicious code into one of its AI coding tools. In both cases, the fix was a narrower permission.
The large vendors now build this in. Microsoft gives each agent built on its platform its own identity in the customer's company directory, with a named person recorded as its sponsor. Gartner advises governing agents "based on action privileges rather than model intelligence", meaning by what they are allowed to do, not by how smart they are. We recommend the same on your org chart: each AI role has its own login, its own inbox if it needs one, and access to only what its job requires.
4. Checkpoints: a named person decides, and that is where you learn
Checkpoints close the loop on the other two. They are where you find out whether your clarity was clear and your constraints held.
An AI role needs three kinds:
- A sign-off on each piece of work that leaves the building. The estimator approves every quote. The quality manager signs every cert.
- A regular review of the role as a whole. Its manager reads the weekly scorecard described in The Second Org Chart.
- A stop. Anyone on the team can pause the role, every action it takes is logged, and a bad draft can be traced back to the input that caused it.
Checkpoints have to be easy to do well. Gartner warns that human approvals "can degrade under time pressure or approval fatigue". A reviewer facing forty drafts at 4 p.m. on a Friday starts clicking approve. Put the sign-off only where it matters, and make the review screen show the source of every value. The Friday Afternoon Test shows how to judge a review screen.
Checkpoints are also how the role improves. When the estimator corrects the same kind of mistake three weeks running, that is a gap in clarity (the examples did not cover the case) or in constraints (the AI was allowed to guess when it should have flagged). The fix goes back into the one-page card, the same way a software team updates its project file after an agent gets something wrong.
5. One card per AI role
Put the three Cs for each AI role on a single page and pin it to the org chart next to the role's box. It answers the questions people will ask about the role from the first day.
6. The worry: will this slow us down?
In our experience, writing three short sections for each role takes an afternoon. Skipping them has a cost that shows up later. Gartner predicts that by 2027, 40% of enterprises will demote or shut down autonomous AI agents because of governance gaps they find only after an incident. In McKinsey's 2025 survey, companies getting the most from AI were almost three times as likely (65% vs 23%) as the rest to have defined when a person checks AI output. In our view, the teams that skip the card do not move faster. They spend the time later, cleaning up drafts that went the wrong direction or arguing about who approved something.
What to do this week
- Pick the one AI role you would add first. Write its job in one sentence.
- Collect ten past examples of that work your best person considers good.
- Write the "never" list: four or five things the role must not do, and the systems it must not see.
- Name the person who signs each piece of work, and the manager who reads the weekly scorecard.
- Ask whether each "never" can be enforced by a login or a permission rather than an instruction. Anything that cannot is a risk to fix before go-live.
The Three-C card is the short version. The full job description, with the manager's weekly and monthly reviews, is in The Second Org Chart. For how cheaper models change which work is worth handing to an AI role, read The Quotes You Never Sent.
Questions people ask
What are the three Cs of an AI role?
Clarity (a written job and examples of good work), constraints (limits the AI cannot cross, enforced by its login), and checkpoints (points where a named person decides). Software teams running coding agents use the same three.
What is the difference between a constraint and a checkpoint?
A constraint stops the AI before it acts, such as a login that cannot reach payroll. A checkpoint is a named person reviewing the work after the AI drafts it. A role needs both, because a reviewer cannot catch an action the AI took on its own.
How do you stop an AI agent from doing something it shouldn't?
Give it its own login with access to only what the job needs, so the limits are enforced by the system rather than by instructions. Then keep a person's sign-off on anything that reaches a customer, and log every action.
Terms in this piece
- Model
- the AI a company such as Anthropic, OpenAI, or Google trains, such as Claude or GPT. It knows nothing about your company until you show it.
- Coding agent
- an AI that writes and tests software on its own for long stretches, with a person reviewing the result.
- Three Cs
- clarity (where an AI role is going), constraints (what it will never do), and checkpoints (where a named person decides). Our name for the three things every AI role needs on paper.
- Permission
- a setting that controls which systems and actions an AI login can reach. A constraint enforced by a permission holds even if the AI misreads its instructions.
- Checkpoint
- a point where a named person reviews the AI's work and decides, including the ability to pause the AI or undo a change.
- Scorecard
- a one-page weekly count of an AI role's drafts accepted, corrected, and rejected, read by its manager.



