The phrase agentic AI is used for almost everything: a chat assistant, a search experience, a workflow with an AI step, a coding tool and a system that plans and acts across several services. This ambiguity is not merely linguistic. It lets procurement, architecture and risk conversations proceed without agreement about what the system can actually do.
Classify observable operating properties first. Product categories can then inform implementation, but they should not replace the description of delegation.
Apply the six-part operational test
1. Goal
The system accepts an outcome that persists beyond producing one response. It can determine that several steps are needed and can assess whether the goal has been reached. “Answer this question” can remain ordinary conversational AI; “resolve this supplier exception” may imply ongoing work.
2. Tools
The system can select or sequence capabilities such as retrieval, APIs, workflows, code, messages or other agents. A fixed integration called at the same point every time is automation. Agentic behavior increases when the system chooses which tool to use and in what order.
3. State
The system retains task-relevant information across steps or events: what has been tried, which evidence is missing, what decision is pending and which commitments already exist. State is not synonymous with unrestricted long-term memory. It should be purposeful, bounded and governed.
4. Authority
The system has permission to affect data, systems or people. Authority can range from drafting a change to executing it without prior review. This property determines much of the operational risk. Tool availability alone does not mean every action should be authorized.
5. Feedback
The system observes tool results, errors, changed conditions or human responses and adapts the next step. Retrying the same call mechanically is different from choosing an alternative route or escalating when confidence falls.
6. Accountability
A named owner defines the outcome, boundaries, evaluation, monitoring, incident response and stop decision. Accountability never transfers to the model. Without it, the organization has created delegation without governance.
A system does not need maximum intensity on every property to be useful. The test explains what kind of system it is and what evidence it needs.
Use a four-level delegation scale
Level 0, Assist: the system generates or retrieves content, and the user directs every step. Examples include summarization and grounded question answering.
Level 1, Recommend: the system assembles evidence and proposes a decision or action, but a human executes it. This can be valuable for complex judgment without delegating authority.
Level 2, Act within a bounded process: the system chooses steps and invokes tools within explicit limits. Humans approve defined consequences or exceptions.
Level 3, Operate toward an outcome: the system initiates or continues work from events, adapts across several steps and completes permitted actions with limited intervention. Monitoring, containment and recovery become service requirements.
The scale is not a maturity ladder. Moving upward is justified only when additional delegation creates enough value to cover increased variability and control cost.
A worked classification: supplier invoice exception
Consider four proposals for invoices that do not match a purchase order.
A chat experience explains the policy when a clerk asks. It is Level 0. A recommendation service reads the invoice and order, highlights discrepancies and suggests the next action. It is Level 1.
A bounded agent chooses which source system to query, requests missing evidence, prepares an adjustment and routes material differences for approval. It is Level 2 because it selects tools and steps but cannot post the adjustment without a defined approval.
A Level 3 system monitors incoming exceptions, contacts permitted internal roles, retries temporary integration failures, posts low-value corrections within a delegated threshold and escalates anomalies. It requires durable state, explicit identity, audit of every tool call, cost controls, exception queues and tested stop behavior.
Calling all four solutions “invoice agents” would hide the material differences in authority and operation.
Match controls to the property that creates risk
Goal persistence requires termination conditions and safeguards against endless work. Tool choice requires allowlists, dependency ownership and validation of tool outputs. State requires retention, access and correction rules. Authority requires least privilege, consequence-based approval and transaction controls. Feedback requires retry limits, escalation and detection of loops. Accountability requires an owner, telemetry, evaluation and incident response.
Microsoft’s agent design framework reflects these operating concerns by separating goals, triggers, tools, knowledge, orchestration, instructions, governance and evaluation. It also recommends deterministic flows or topics for compliance-driven, high-impact or irreversible steps. Agent architecture can therefore contain deterministic components where predictability matters.
Know when not to build an agent
Choose a prompt or search experience when the user mainly needs information and remains in control. Choose a deterministic workflow when inputs, rules and sequence are stable and exceptions can be explicitly routed. Choose an application when users need persistent records, structured interaction and visible state.
An agent becomes appropriate when cases vary, the next step depends on observed results, several tools may be needed and hard-coding every path would be brittle. Even then, delegate only the decisions whose variability the organization can evaluate and contain.
Avoid equating autonomy with intelligence or business value. A tightly controlled recommendation can outperform an autonomous process if the latter creates more exception handling than it removes.
Turn the label into a design record
For every proposal, document the six properties, delegation level, human boundaries, failure consequences and evidence required for release. If a team cannot explain these elements, the architecture is not ready for procurement or implementation.
Amplified Pi uses this record to move discussions from ambition to an implementable pattern. The result may be an agent, a workflow, an application or a combination. The right answer is the least complex system that delivers the required outcome with a control model the organization can operate.