Distribution no longer ends at installation
Imagine buying software without opening a product page. You tell an assistant to compare three suppliers, read the renewal clauses in their contracts, check the team's calendar, and prepare a recommendation. The assistant selects several capabilities, asks for access to the relevant documents, calls the tools, and returns with a decision packet. Only the final email still needs your approval.3, 4
That interaction is not a better App Store search. It changes the unit being distributed. The App Store packages software as an application that a person discovers, installs, and operates. Software as a service moves the application to infrastructure managed by a provider, but the user still enters a product, learns its interface, and drives the work. An agent market distributes something less self-contained and more consequential: a capability that may plan, call tools, preserve state, and act under delegated authority.1, 2, 3
Apple's account of the App Store's first decade describes a launch with 500 apps and a single place for discovery, distribution, review, and payment. That arrangement solved a mobile-software problem: how to put a trustworthy package on a device. The next distribution problem is different. It is how to let a system find and combine capabilities without granting it an invisible, permanent key to the user's work.1, 8, 4
| Decision | App Store | SaaS | Agent market |
|---|---|---|---|
| What is distributed | An installable application | Access to a hosted application | An agent, tool, skill, connector, or workflow |
| Typical instruction | Install this | Open this workspace | Complete this task |
| Who operates the interface | The user | The user | The user delegates; the system may operate tools |
| Main permission boundary | Device and operating-system access | Account, role, and tenant access | Task, tool, data, write, and transaction authority |
| Common billing unit | Purchase, in-app purchase, subscription | Seat, plan, or usage | Model, call, action, task, or outcome |
| What failure looks like | The app crashes or misbehaves | The service fails or corrupts data | The system may take the wrong external action |
| Strongest lock-in | Platform purchases and device services | Data, workflow, and contract | Memory, permission graph, execution history, and integrations |
The useful name for this layer is still unsettled. It appears as GPTs, apps inside a conversation, agents in an enterprise directory, marketplace listings, skills, and connectors. Those differences matter. A consumer directory, an administrator-approved tenant catalog, and a cloud marketplace do not offer the same rights or review. “Agent Store” is therefore best treated as a category in formation, not one finished commercial format.4, 7, 8, 9
Capabilities cross application boundaries
An application can be complex, but its submitted binary and declared services give a reviewer a relatively stable object. An agent is assembled at runtime. Its behavior depends on a model, instructions, context, memory, tool connections, identity, external content, and the state of every service it calls. Change one model version, retrieved page, connector response, or organization policy and the same initial request can take a different route.3, 5, 13
This is why two new protocols deserve attention. Anthropic introduced the Model Context Protocol as an open way for AI applications to connect to data sources and tools. Google's Agent2Agent protocol addresses a different relationship: how agents can describe capabilities, discover one another, exchange task information, and coordinate across vendors and frameworks. A short working distinction is useful: MCP helps an agent use a tool; A2A helps one agent work with another.5, 6
Open protocols reduce the cost of wiring systems together. They do not decide whether the connection should exist. A travel task can cross an assistant, a corporate identity provider, email, calendar, expense policy, airline inventory, and payment. Each hop introduces credentials, logs, retention rules, and a new place where external content can alter the task. Interoperability expands the reachable surface faster than it settles responsibility.5, 6, 13

The durable architecture is layered. Apps remain the boundary for device sensors, local files, and deep visual work. SaaS systems remain the records for customers, contracts, inventory, projects, and payments. Agents become an operating layer that can plan across those systems. The market above them determines which capability is discoverable, approved, billable, and eligible to act.1, 2, 3, 8
| Part | Product question | Failure to rehearse |
|---|---|---|
| Model | Which model performs which step, and what can change? | An upgrade changes behavior without an operating review |
| Instructions | Which policies win when instructions conflict? | A persuasive external page redirects the task |
| Context and memory | What is retained, for how long, and for whom? | Old or cross-user context contaminates a decision |
| Tools | Which calls are read-only, write-capable, or transactional? | A lookup quietly becomes a change in another system |
| Identity | Whose authority is used at each step? | A broad user token becomes permanent agent authority |
| Runtime | Where are state, retries, budgets, and timeouts controlled? | Loops and retries create cost or duplicate actions |
| Evaluation | Which failures block release or trigger shutdown? | A high average score hides one catastrophic route |
Each listing must disclose its permission contract
A conventional store page can show screenshots, ratings, privacy labels, and an install button. An agent listing needs to explain behavior that may not fit on a static page. “Uses calendar” is too vague. Buyers need to know whether the agent reads or writes, which fields leave the tenant, where those fields are sent, what is retained, which model or subcontractor processes them, and which operations require a new confirmation.7, 8, 16
The minimum-permission principle is familiar; the interaction design is not. People will accept a broad access request when it stands between them and a promised result, especially after they have already invested time in the task. A credible delegation flow exposes the plan before the grant, separates read from write, asks again at the irreversible step, and keeps the stop control in the same place as the progress view.7, 13, 14

That five-part chain—plan, permission, action, log, revoke—should be visible in the product, not buried in policy prose. The plan says what the system intends to do. The permission step binds access to this job. The action view distinguishes a proposal from a change already sent to an external service. The log records the route and evidence. Revoke stops future use and, where the action allows it, points to a recovery or compensation path.7, 8, 14
For a low-risk research task, one approval may cover several read-only calls. For sending email, publishing content, deleting records, signing a contract, or paying a supplier, “human in the loop” must become a named control: who approves, at what step, against which preview, within what amount, and with what emergency stop. A generic confirmation at the beginning is not equivalent.13, 14, 15
- Show the proposed route and the tools it will call before the first sensitive grant.
- Bind access to a task, data scope, duration, and maximum spend instead of reusing an unlimited token.
- Separate reading, drafting, writing, publishing, deleting, and paying into distinct authority levels.
- Preserve the source, output, approver, model or agent version, tool response, and final external effect in the log.
- Put revoke, retry, rollback, and incident reporting beside the running task rather than in an administrator maze.
Outcome pricing needs a measurable event
SaaS buyers learned to ask for price per seat, billing cadence, minimum commitment, and the boundary between plans. Agent economics break that tidy model. One task can consume model inference, search, file retrieval, code execution, database calls, third-party APIs, human review, and payment fees. A cheap-looking subscription can still create an unpredictable operating bill; an expensive task can still be economical if it reliably replaces slow, high-value work.11, 12, 10
AWS publishes AgentCore charges across services such as runtime, memory, identity, observability, gateway, and policy. OpenAI lists model and tool charges separately. Salesforce presents Agentforce through editions, usage, and credit-based units. The names and meters differ, but they point to the same procurement problem: “How much is the agent?” is incomplete until the buyer defines a task and its full route.11, 12, 10
| Cost layer | Meter to capture | Question before approval |
|---|---|---|
| Model | Input, output, cached context, or session | Which steps need the most capable model? |
| Retrieval and tools | Search, storage, code, connector, or API call | Can repeated calls be bounded and inspected? |
| Runtime | Time, memory, retries, and state | What stops a loop or duplicate action? |
| Marketplace | Listing, distribution, or revenue share | Is paid placement distinguished from task fit? |
| Transaction | Payment or business outcome | Who pays for reversals, disputes, and failure? |
| Human control | Review, exception handling, support, and audit | Which labor is saved, and which labor is newly required? |
Outcome pricing sounds cleaner because it links payment to value. It also turns the definition of “success” into a contract. A qualified lead, completed claim, saved dollar, or booked trip can be measured in several ways. The party controlling attribution may also control the bill. Buyers should require a reproducible event definition, exclusion rules, a dispute path, and a cap before treating outcome pricing as aligned incentives.10, 16
Durable advantage lives beyond the prompt
Natural-language instructions and ready-made connectors lower the cost of producing a demonstration. They do not lower the cost of operating a dependable product. The gap appears after the first successful task: tool errors, prompt injection, expired credentials, changing model behavior, ambiguous user intent, duplicated actions, support requests, audit evidence, and regional policy all arrive at once.3, 5, 13
AgentDojo makes the security problem concrete. Its benchmark combines realistic tool-use tasks with indirect prompt-injection attacks. The important lesson is not one leaderboard position. It is that useful task completion and resistance to malicious instructions must be evaluated together. A system that refuses every tool call is safe but useless; one that completes the happy path while obeying hostile content is not a product.13
The strongest developer advantages will therefore come from assets that survive model substitution: permissioned vertical data, reliable tool connections, a well-understood industry process, difficult evaluation cases, low exception cost, responsible support, and a record of correcting failures. A prompt wrapper can be copied. A clean permission graph and a year of audited edge cases are harder to reproduce.13, 14, 8
Store review must change with the product. An application review can inspect a declared package at submission and revisit it after updates. An agent's behavior also changes when models, instructions, external content, connectors, and tenant policies change. Approval has to become continuous: signed versions, regression suites, policy checks at runtime, abnormal-call monitoring, trace review, periodic reapproval, and a kill switch.8, 14, 13
Platforms gain power over selection and execution
Today a recommendation system influences what a person watches or buys. In an agent market it may influence what software acts, which data it receives, and which transaction it attempts. That is a larger grant of platform power. A directory should therefore separate organic task matching from paid placement, explain why a capability was selected, and preserve enough route information for the user or administrator to challenge the choice.16, 8, 9
The near-term advantage belongs to platforms that already hold enterprise identity, data, workflows, and billing. Microsoft can place agents inside an organization-controlled directory. Salesforce can connect a marketplace to customer records and business processes. Cloud providers can combine runtime, identity, observability, procurement, and metering. Their advantage is not simply a better model; it is that they already sit at several control points in the task.8, 9, 11
That does not mean apps disappear. Operating systems still govern device permissions and local resources. Rich creation, analysis, and visual control still benefit from dedicated interfaces. SaaS systems remain the databases and operational records behind many tasks. The likely hierarchy is more layered: the app is a trusted interaction boundary, SaaS is the system of record, the agent is the operating layer, and the agent market governs discovery, approval, and trade.1, 2, 4
Regulation will follow the chain rather than the label. The EU AI Act organizes obligations around roles and risk, while the Digital Services Act adds marketplace traceability duties for traders. China's interim rules for public generative-AI services establish another set of provider responsibilities. None of these regimes is a finished “Agent Store law,” but together they make a useful question unavoidable: who developed, deployed, recommended, authorized, and profited from the action?15, 16, 17
NIST's AI Agent Standards Initiative focuses on areas including interoperability, security, and identity. Standards can reduce ambiguity, but they will not remove commercial incentives to create new lock-in. Even when tool calls are portable, memory formats, permission history, evaluation methods, billing units, and marketplace reputation may not be. The right to export an execution record will matter as much as the right to export a document.14, 5, 6
Controls buyers and builders need now
The category will change quickly, so a static winner list would age badly. The more durable decision is to make delegation inspectable and reversible before allowing it to become routine. Start with one bounded task whose evidence and external effects can be checked. Do not begin with the workflow that can move the most money or expose the most data.14, 13
| Reader | Three actions now |
|---|---|
| User or buyer | Choose agents that expose their plan, sources, permissions, cost, and log; require fresh confirmation for payment, deletion, signing, and publishing; export configuration and execution history on a schedule. |
| Developer | Measure task success together with error cost and unit margin; build least-privilege tools, budgets, timeouts, regression tests, and incident shutdown into the product; compete on workflow evidence rather than a prompt alone. |
| Platform | Disclose ranking and paid placement; issue verifiable agent identities and signed versions; provide task-scoped permission, trace, revoke, portability, and emergency-disable primitives. |
| Regulator | Classify by the consequence of an action; make the model, developer, deployer, marketplace, tool, and authorizer visible in the responsibility chain; support interoperable identity, audit, incident, and evaluation standards. |
The shift from App Store to Agent Store is not a move from icons to chat bubbles. It is a move from selecting a tool to delegating authority. The scarce resource is no longer the number of available apps. It is a combination of trustworthy permission, verifiable result, predictable cost, and a portable record of what happened.3, 14, 16
The platform that controls discovery, identity, permission, execution, and recovery will look less like a storefront and more like an operating system. Buyers should not wait for the market to choose that platform for them. They can start by insisting that every delegated action has an owner, an evidence trail, a cost boundary, and an exit.8, 14, 11

