A successful demo proves that an agent can complete a selected task under favorable conditions. It does not prove that the organization can operate it. Before production, eight boundaries must be explicit: outcome, audience, knowledge, authority, controls, evaluation, lifecycle and ownership. If one remains implicit, the agent may still be useful, but it is not yet an accountable production service.
A demonstration tests capability. Production tests responsibility.
Prototype conversations usually focus on answer quality, orchestration and tool use. Production introduces different questions. Who is affected when the answer is wrong? Which identity performs an action? What happens when a source is unavailable or contradictory? Who receives an alert, approves a change, handles an incident or disables the service?
Microsoft’s agent design guidance covers goals, triggers, data, tools, governance and evaluation. Its operating checklist adds environment separation, data policies, monitoring, application lifecycle management and named ownership. The durable lesson is broader than Microsoft: an agent becomes production-ready only when its behavior is surrounded by an operating system.
This distinction matters because a prototype often hides favorable assumptions. The maker knows the intended wording. Test data is clean. Permissions are broad. The happy path receives most attention. Production removes those protections. Users ask ambiguous questions, source systems fail, ownership changes and a seemingly minor update alters behavior.
Record eight boundaries before release
The production record should fit on one page. Its purpose is not to document everything. It is to make the decisions that determine exposure, evidence and accountability visible.
| Boundary | Decision to record | Evidence before production |
|---|---|---|
| Outcome | Which business result should change? | Baseline, target measure and review cadence |
| Audience | Who may use it and who may be affected? | Access groups, excluded cases and affected-party analysis |
| Knowledge | Which sources are permitted and how fresh must they be? | Source register, permission tests and stale-content behavior |
| Authority | What may the agent read, recommend, create, change or send? | Tool permissions, acting identity and prohibited actions |
| Controls | Where are validation, approval, escalation and stop mechanisms required? | Control tests and named escalation paths |
| Evaluation | Which normal, ambiguous and adversarial cases must pass? | Test set, acceptance thresholds and known limitations |
| Lifecycle | How are versions built, tested, released and reversed? | Environment strategy, dependencies and rollback procedure |
| Ownership | Who owns value, operation, risk and support? | Named roles, review dates and incident duties |
An unresolved cell is not automatically a reason to stop. It is a reason to classify the current state honestly. An experiment may tolerate unresolved production questions if its audience, data and consequences are tightly bounded. A production service may not.
Authority creates the largest change in risk
An agent that summarizes documents exposes different risk from an agent that sends messages, changes records or operates a browser. Authority should be described as verbs, not as a vague statement that the agent has access.
For example, “access to the service-management system” is incomplete. The useful record says that the agent may read tickets assigned to the current user, draft an internal work note and propose a category. It may not close a ticket, change priority, contact the requester or read another team’s queue without a person approving the action.
This level of precision helps engineering and governance at the same time. Engineers know which identities and tools to configure. Security can assess least privilege. Operations can define alerts. The business owner can decide whether the remaining manual step is acceptable.
Put controls where uncertainty meets consequence
Not every model output needs human approval. Not every agent action should be automatic. Use deterministic validation for facts and rules the model should not guess, such as required fields, identifiers, totals, permitted recipients or policy thresholds.
Human review belongs where interpretation remains necessary and the consequence of an error is material. Irreversible, legally significant or externally visible actions deserve a stronger boundary than a reversible internal draft. Some actions should remain prohibited because the available controls cannot reduce the exposure sufficiently.
The design goal is not maximum automation. It is the smallest amount of human intervention that preserves accountable judgment.
Evaluate the system, not only the conversation
Answer-quality testing is necessary but incomplete. A production evaluation should include permission boundaries, source failures, unavailable tools, malicious or irrelevant instructions, duplicate actions, long-running tasks and recovery after interruption.
The expected result is not always a correct answer. It may be a refusal, a clarifying question, an escalation or a safe stop. These behaviors need explicit test cases because users experience them as part of the service.
Evaluation also continues after release. Monitor tool calls, errors, retries, latency, user feedback, overrides and business outcomes. A passing pre-release test set does not prove that sources, usage patterns and dependencies will remain stable.
A worked release decision
Consider an agent that triages incoming support requests. In the demo, it reads an email, identifies a category and drafts a response. The team proposes automatic sending.
The production record exposes three missing boundaries. The customer database contains similarly named accounts, so identity resolution is uncertain. Some categories carry contractual response commitments. The knowledge source contains outdated instructions. Automatic sending would combine uncertainty with external consequence.
A defensible pilot allows the agent to classify the request, retrieve approved guidance and draft a response. Deterministic checks validate the account identifier and required fields. A support employee reviews the response before sending. The pilot measures classification quality, edit distance, handling time and escalation reasons. Automatic sending is reconsidered only after the evidence shows where it is safe.
This is not a failed automation. It is a deliberately bounded service that produces evidence for the next decision.
Use three states instead of a binary launch decision
Experiment: optimize for learning. Use synthetic or low-risk data, a small expert group, no consequential autonomous action and a fixed end date.
Bounded pilot: use real work within a defined audience and authority. Assign owners, monitor behavior and state the evidence required for expansion.
Production service: operate with approved controls, lifecycle, support, monitoring, incident response and periodic review.
The state should be visible to users and reviewers. Calling a pilot “production” creates false assurance. Calling every controlled test an “experiment” can hide real operational exposure.
Common failure modes
The most common failure is allowing the demonstration architecture to become the production architecture by inertia. Other warning signs include shared maker credentials, broad connectors used for convenience, no baseline for value, knowledge sources without owners, manual changes made directly in production and monitoring that counts conversations but not failures or outcomes.
Another failure is governance arriving only after the build. Late review often produces redesign because identity, environment or data choices are already embedded. The production record should begin before implementation and mature with the agent.
The final release question
An agent is ready for production when the team can explain its permitted behavior, demonstrate representative tests, identify accountable owners and show how changes, incidents and capacity will be handled. The answer may still be “not yet.” That is a useful result if the missing evidence and next bounded step are clear.
Amplified Pi uses the eight-boundary production record to connect AI strategy with delivery. We help organizations select the right pattern, build the system, establish evidence and operate it under proportionate control. The next step is a production-boundary review of one real agent, not another generic readiness survey.