Specification Sovereignty
Why You Can Outsource Execution But Never the Definition of Success
One of the great attractions of generative AI is that execution keeps getting cheaper. Work that once required scarce professional time can increasingly be delegated to a model, an agent, a vendor, or some combination of all three. That can be a very good thing; I have no attachment to having expensive people perform routine work simply because expensive people have always performed it.
But there is a question hiding underneath all that delegation: what must you still know how to do yourself?
I do not mean whether someone in your organization must continue manually performing every task the AI performs. That would defeat the whole point. I mean something far more fundamental. If you delegate execution, do you still possess enough internal knowledge and judgment to specify what the system is supposed to accomplish, recognize when it has subtly failed, define the exceptions that matter, and evaluate whether something else would do the job better?
If you can outsource execution, can you also safely outsource the definition of success at the same time? This question is surprisingly difficult to answer.
Specification vs. Execution
Execution and specification are fundamentally different capabilities. We notice execution because it produces visible work in the form of memos, briefs, filings, and processed claims. Specification is quieter. It is the institutional judgment that tells you which result you actually want, which compromises are acceptable, and which technically successful output still counts as a failure for your organization. As execution gets easier to delegate, specification becomes the core capability you must protect. Even before AI, good playbooks were difficult to produce and maintain.
Organizations have relied on third-party vendors, consultants, and managed services for decades. Good outsourcing has never required pretending that the buyer knows how to perform every task better than the seller. Specialized technical expertise is often exactly what you are buying.
There is a critical difference, however, between relying on someone else’s execution expertise and allowing someone else to define your objective. Suppose an AI vendor provides an automated legal or financial workflow and also supplies the metrics by which that workflow will be judged such as speed, cost per transaction, percentage of matters handled automatically, or a vendor-defined accuracy score. Those may all be useful metrics, but do they describe what your organization really means by success?
What happens if the workflow gets faster while creating more difficult exceptions downstream? What if a high aggregate accuracy rate hides catastrophic failures concentrated in the few edge cases that matter most? What if users love the system because it gives quick answers, while experienced professionals notice those answers quietly omit a distinction your business considers vital?
At that point, the vendor is doing more than supplying software. It is setting your business objectives. If someone else defines what “good” looks like, they are no longer merely your supplier. They are governing your strategy. If they provide metrics and benchmarks for you and measure their success to their own benchmarks, what have you lost control over?
The Failure That Looks Like Success
This risk is subtle because the most dangerous AI failures rarely look like failures. If an automated system crashes, refuses to produce an answer, or outputs obvious nonsense, you have an obvious operational problem. As frontier systems become more capable, however, the real threat is the output that looks competent. It is polished, responsive, and broadly correct. It may even score perfectly against the vendor’s benchmark suite.
It just isn’t quite doing what your organization needs it to do. That “quite” can contain an enormous amount of risk. The model might miss a category of exceptions that experienced practitioners know matter disproportionately. It might optimize for speed when the real objective is reducing downstream rework. Or it might follow the formal rule correctly while missing the institutional judgment that dictates when the formal rule should not control.
The deepest risk is not simply that AI will make mistakes or hallucinates. People make mistakes and misunderstand instructions every day. The deeper risk is that you gradually lose the internal experts who can tell the difference between a plausible-looking output and a genuinely successful one.
That loss happens quietly. As execution is delegated, fewer people do the underlying work. As fewer people do the work, fewer people encounter the edge cases and operational consequences through which practical judgment develops. Eventually, the organization possesses plenty of people capable of operating the system, but almost no one capable of challenging its definition of success.
The system works. The dashboards are green. And nobody quite remembers what the dashboards are failing to measure.
Specification Lock-In
The question is not which activities must remain internal, but which evaluative capabilities must remain in-house when manual labor disappears. Execution capability can be hired. Specification capability must be retained.
How much specification capability you need varies. A commodity workflow requires very little. A workflow tied directly to competitive advantage, professional judgment, client obligations, or high-consequence decisions requires a great deal. Specification is an argument for knowing which evaluative judgments you cannot afford to lose rather than a simplistic argument for keeping everything in house.
Test your organization with a straightforward diagnostic question:
If execution were completely delegated tomorrow, what specific knowledge and evaluative capability must you retain in-house to independently specify, audit, challenge, and replace the system if it fails?
The word replace is decisive. If you cannot specify the system independently, you cannot evaluate it independently. If you cannot evaluate it independently, you cannot challenge it. And if you cannot challenge it, replacing it simply means moving from one externally defined objective to another.
This is the ultimate form of vendor dependence: specification lock-in. If a vendor defines the objective, the metrics, and the acceptance criteria, switching vendors does not restore your independence. You no longer possess enough internal judgment to tell the next vendor what you need or to evaluate whether its proposed system is any better.
The Sovereignty Boundary
This is where Prudent AI draws the line differently from a conventional build-versus-buy debate. The question is never which operational tasks to keep inside, but which judgments must remain sovereign.
You can let an AI write the first draft while retaining the internal capability to state what a good draft must accomplish. You can automate legal or financial analysis while retaining the knowledge needed to identify the underlying assumptions that matter. You can let a vendor execute a complex workflow while retaining the authority and capacity to decide whether its acceptance criteria match your own.
That is specification sovereignty. It does not require doing everything yourself. It requires remaining capable of defining for yourself what success means.
Prudent AI governs consequential operational commitments under uncertainty. As AI systems become more deeply embedded, delegate execution to gain leverage, but preserve the sovereign internal judgment required to specify and evaluate what your organization needs.
Accelerate reversible learning. Pace irreversible commitment.
Dennis Kennedy – CC BY 4.0 license
[Originally posted on DennisKennedy.Blog (https://www.denniskennedy.com/blog/)]
DennisKennedy.com is the home of the Kennedy Idea Propulsion Laboratory
DennisKennedy.Blog is part of the LexBlog network.