← Back to blog
Choice

The AI Vendor Scorecard: How to Buy Without Getting Locked In

A buyer's guide to evaluating agent platforms through the only four questions that survive the next release: Context, Control, Cost, and Choice.

15 min readby Team BricksNotes
enterprise AIagentic AIdata professionalsprocurementvendor selectionlock-inbuyer's guideagent platformsgovernanceportability
Share
01

The demo that closed the deal and the runtime that closed the door

Every enterprise AI program has the same story, told with different vendor names. A senior leader sees a demo that answers a question the team has been arguing about for a year. Legal is looped in. Procurement runs a compressed evaluation. A contract is signed. Six months later, the same team is quietly building a second internal project called migration. The demo was real. The runtime is a cage.

This is not a failure of due diligence. It is a failure of the questions the diligence asked. Traditional software procurement was built for a world where the product was mostly the product. You compared features, SLAs, and pricing. You did not need to ask whether the definitions inside the tool would be portable, because the tool did not have a semantic layer of its own. You did not need to ask whether the audit trail would survive an export, because the audit trail was just log lines. You did not need to ask whether the evaluations would run against a different model, because there was only one model, and it was called your codebase.

Agent platforms are different. What you are buying is not a feature set. It is a runtime that will accumulate meaning, memory, evaluations, and policy as your team uses it. The value of the platform to your organization compounds over time, and so does the cost of leaving. The right question at signature is not which one has the best demo. It is which one will still be a good decision the day you want to swap the model underneath.

This essay is the scorecard we wish every data leader had walked into their last agent platform evaluation with. It is organized around the four questions the book runs on — Context, Control, Cost, and Choice — because those are the four questions the market will still be answering three years from now, when the vendor logos on today's slides have merged, pivoted, or disappeared.

Editorial illustration of a vintage four-tray balance scale, with each tray holding a labeled weight: Context, Control, Cost, and Choice, on cream paper with coral accents.
Four weights on the same scale. A vendor evaluation that ignores one of them is a vendor evaluation that will regret it.
02

Why the 4 C's, and not the RFP template

Most enterprise AI evaluations still run on an RFP template that was written for a SaaS purchase in the mid-2010s. It asks about uptime, SOC 2, data residency, price per seat, and integrations. Those questions matter. They are also, in 2026, the ones the vendors have all learned to answer identically. Every serious agent platform has SOC 2. Every serious agent platform can host in your region. Every serious agent platform integrates with your data warehouse and your identity provider. The RFP no longer discriminates between the vendors that will compound your advantage and the vendors that will slowly absorb it.

The four questions that do discriminate are the ones the book calls the 4 C's. Chapter 4 introduces the framework and explains why these four survive the churn. Context — does the platform let you build, own, and export the meaning layer that decides whether the agents give correct answers on your data. Control — does the platform give you the guardrails, approvals, and audit trails that keep an autonomous system inside the lines your policy and your regulators have drawn. Cost — does the platform expose the real unit economics of an agentic workflow, or does it hide the exploration tax inside a friendly-looking seat price. Choice — can you leave. Not in theory. Not with an export button. Actually leave, with your context, your evaluations, and your policies intact, running on top of a different model or a different runtime.

The 4 C's are not a marketing device. They are the four places where lock-in accumulates and where advantage compounds. A vendor can score well on one and poorly on another, and the shape of that scorecard tells you exactly what you are agreeing to when you sign.

03

Context: are you renting definitions, or building them

The first and most consequential question is whether the meaning that makes the agents useful lives inside the vendor's runtime, or inside your organization. Every agent platform has to represent your business somehow. It has to know what revenue means, what a customer is, which region a transaction belongs to, and how churn is calculated on Tuesday versus Wednesday. That representation is the semantic layer, and the platform is going to build one whether you plan for it or not.

The difference between vendors is who owns it. Chapter 16 argues that the semantic layer is the highest-leverage artifact in an enterprise AI program, and Chapter 15 argues that ownership of that layer is not a technical detail — it is the difference between a moat you have built and a moat you have leased. If your metric definitions live inside a vendor-specific catalog with no export path, you have leased them. If your customer entity model is expressed in a proprietary DSL that nobody outside the vendor's engineering team can read, you have leased it. If the retrieval logic is a black-box service call with no way to inspect or replay what it fetched, you have leased that too.

The vendors themselves will not describe this as leasing. They will describe it as convenience. And in the first six months, it will be convenience — the platform will take care of things you did not want to build. In the second year, it will be the reason a competing runtime that scored better in the last evaluation is not a realistic option, because the semantic layer that took your team eighteen months to curate cannot come with you.

The Context questions on the scorecard are therefore not about whether the vendor has a semantic layer. Every serious vendor does. The questions are about who owns it, where it lives, and whether it can leave. Can the definitions be exported in a format that a different platform can consume without a rewrite. Are the entity models expressed in an open standard, or in a proprietary schema. Is the retrieval logic inspectable — can you see, for any answer, what was fetched and why. Are the citations that back an answer stored in a way that survives a vendor change. If the answers to those questions are unclear, you are not evaluating a Context-strong platform. You are evaluating a Context-flavored one.

A hand-drawn ledger with columns labeled Vendor A, B, and C and rows labeled Context, Control, Cost, and Choice, with coral checkmarks and crosses filling the cells, and a serif quill pen resting beside it.
The scorecard is the artifact. Every serious evaluation should leave one behind, in writing, before signature.
04

Control: can the agent do what it is not supposed to do

The second question is whether the platform treats the agent as an actor that must be governed, or as a chatbot that occasionally calls tools. This distinction sounds academic until the first time an agent in your environment takes an action that surprises somebody. That day will come. The only variable is whether the platform gave you the guardrails to prevent it, the audit trail to explain it, and the reversal path to undo it.

Chapter 13 makes the case that action control matters more than answer quality in the next phase of enterprise AI, because the actions are the surface where a mistake becomes a compliance event. Chapter 14 walks through what a control plane actually contains — the list of tools an agent can call, the approval routes for the actions that are not fully autonomous, the reversal paths for actions that turn out to be wrong, and the audit trail that records what was decided and by whom. Chapter 19 closes the loop with the human-in-the-loop patterns that keep an agentic workflow inside a regulator's line of sight without slowing it to a stop.

The Control questions on the scorecard are the ones the vendor demos will not naturally invite. Can you scope tool access per agent, per role, per environment, without a services engagement. Can you route specific action types through human approval, with the reason for approval or denial captured in the record. Is every agent action logged with the inputs, the retrieved context, the decision, and the identity of the human or policy that approved it. Can that log be exported to your SIEM in a format your existing tooling already understands. Can an action be reversed, and if so, by whom and on what SLA. If any of those answers requires a professional services contract to arrange, you are looking at a platform that treats governance as an upsell rather than a foundation.

05

Cost: what is the actual unit economics of a successful task

The third question is the one that catches almost every buyer, because the vendor's answer to the pricing question is almost always technically true and practically misleading. Agent platforms are typically priced on some combination of seats, tokens, tool calls, and platform fees. The number on the invoice is not the number that decides whether the program is economically viable.

The number that decides that is cost per successful task. An earlier essay on this site, The Compute and Latency Budget, argues that agentic workflows are governed by a three-axis budget and that the compounding advantage in cost comes from context density, not from clever pricing negotiation. Chapter 18 frames the accounting: it is not enough to know what a token costs, because a task that requires forty tool calls and three retries to reach the right answer costs a multiple of a task that reaches it on the first try. Chapter 20 explains why the exploration tax — the reads, retries, and speculative queries an underinformed agent has to run before it can converge — is where most of the money quietly goes.

A vendor evaluation that asks only for a price sheet will get a price sheet. A vendor evaluation that asks the right Cost questions will get something more useful: an honest picture of what a real workflow costs to run to completion on real data. Can the platform expose cost per task, not just cost per token. Can you attribute cost to the specific unit of work — an intent, a workflow, an agent — that caused it. Does the pricing model align with successful completions, or does it reward the platform when the agent has to retry. Is there a floor of platform fees that has to be paid regardless of usage, and if so, what workloads have to exist for that floor to be worth paying. What happens to the bill when the exploration tax spikes — do you find out at the end of the month, or in real time.

The most important Cost question is the one about ceilings. Every vendor is happy to talk about how their platform performs at the size of your first pilot. Ask what happens at ten times that size, and at a hundred times, and whether the pricing curve is linear or sub-linear. If the vendor cannot answer, the pricing model has not been designed for the world you are about to enter. It has been designed for the pilot you are running now.

Hand-drawn cross-section of an iceberg. Above the waterline: sticker price, seat and token fees. Below: exploration tax, retry cost, migration cost, eval rebuilds, governance gap.
The invoice is the tip. The bill is what sits underneath it.
06

Choice: can you actually leave

The fourth question is the one the book returns to more than any other, because it is the one that decides whether every other advantage is durable. A platform that scores perfectly on Context, Control, and Cost, but that you cannot leave without giving up all three, has priced itself into an eventual renegotiation you will lose. Choice is not a preference. It is the structural feature that keeps the other three honest over time.

Chapter 22 frames what portability actually requires. It is not a button labeled Export. It is a set of standards, contracts, and internal disciplines that keep every artifact the platform accumulates — semantic layer, retrieval logic, memory, evaluations, policies, audit trails — in a form that a different runtime can pick up. Chapter 12 argues that portable context is the moat, because a context layer that lives only inside one runtime is a moat you are renting.

The Choice questions are therefore not the ones a vendor will love to answer. Which artifacts can be exported, in which formats, on what cadence, and can you export them without a services engagement. Are the tool contracts your agents call — the interfaces that connect them to your systems — expressed in an open standard like MCP or A2A, so a different runtime can call the same tools without a rewrite. Are the evaluations you accumulate model-agnostic, so that when a new frontier model arrives you can benchmark it in place, or are they wired specifically to the current runtime. Is there a documented, tested migration path that the vendor's own customers have used, or is Choice a slide in the sales deck.

The related essay MCP and A2A: The Open Protocol Moment walks through why the current standards moment is the single most important Choice-preserving development of the last two years. A platform that supports MCP for tool calls and A2A for agent-to-agent handoffs is a platform that has quietly told you it does not intend to make leaving expensive. A platform that has invented its own equivalents and has not published a bridge to the open standards is a platform that has quietly told you the opposite.

Hand-drawn illustration of a vintage padlock on a cloud-shaped platform labeled Vendor Runtime, with a chain labeled context, evals, and tool contracts leading to a coral key labeled Export Contract.
The export contract is the key. Sign the platform, and sign the key too.
07

The twenty questions, in the order to ask them

A working scorecard is not a spreadsheet with two hundred rows that nobody reads. It is a short list of questions, each with a defensible answer, that the whole buying committee can rally around. The following twenty are the ones we have found consistently separate the vendors that will still be a good decision in three years from the ones that were a good demo.

On Context: One, who owns the semantic layer, and can it be exported in a format a different platform can consume without a rewrite. Two, are entity and metric definitions expressed in open standards or a proprietary DSL. Three, is the retrieval path inspectable — can you see, for any answer, what was fetched and why. Four, are citations stored with each answer in a durable format. Five, can institutional memory be versioned, reviewed, and rolled back like production code.

On Control: Six, can tool access be scoped per agent, per role, and per environment without a services engagement. Seven, can specific action types route through human approval with the decision captured. Eight, is every agent action logged with inputs, retrieved context, decision, and approver. Nine, does the audit log export to your existing SIEM in a format your tooling already reads. Ten, can an action be reversed, and by whom.

On Cost: Eleven, can the platform report cost per successful task, not only cost per token. Twelve, can cost be attributed to the specific intent, workflow, or agent that caused it. Thirteen, does the pricing model reward successful completions or reward retries. Fourteen, what is the platform floor, and what workloads justify paying it. Fifteen, how does the cost curve behave at ten and one hundred times pilot volume — linear, sub-linear, or worse.

On Choice: Sixteen, which artifacts can be exported, in which formats, on what cadence, and without a services engagement. Seventeen, are tool contracts expressed in MCP, A2A, or equivalent open standards. Eighteen, are evaluations model-agnostic and portable to a different runtime. Nineteen, has any customer publicly executed a migration off the platform, and on what timeline. Twenty, will the vendor commit contractually to a bounded migration cost and a defined data-export SLA at contract termination.

Score each question on a simple three-point scale: yes with evidence, yes with caveats, or no. Total the scores by C. Any vendor with a zero column has already told you the answer. Any vendor with balanced coverage across all four is worth a second conversation. Any vendor that has strong Context, Control, and Cost, and a weak Choice column, is exactly the vendor most enterprises regret signing in year two.

08

The contract clauses that make the scorecard real

A scorecard that never leaves the evaluation phase is a scorecard the vendor's account team will happily let you keep. The scorecard becomes real when it is written into the contract. That is where most enterprise AI programs quietly give up their negotiating leverage, because the buying committee that ran a rigorous evaluation hands the contract to a legal team that has never seen an agent platform before and defaults to the standard SaaS template.

There are five clauses that turn the 4 C's from an evaluation exercise into a durable protection. First, a data and context export clause that specifies the artifacts, formats, cadence, and SLA at termination — including semantic layer, memory, evaluations, and audit logs. Second, an open standards clause that commits the vendor to maintaining MCP-compatible or equivalent tool contracts for the term of the agreement, so your integrations remain portable. Third, an audit and reversal clause that requires every agent action to be logged with a defined schema and made available to your SIEM in a specified format. Fourth, a unit economics reporting clause that requires the vendor to expose cost per task and cost attribution at a defined granularity, not only aggregate token or seat spend. Fifth, a migration cooperation clause that caps the vendor's ability to charge for termination assistance and defines the reasonable-effort standard for cooperation with a successor runtime.

None of those clauses are exotic. All of them exist, in some form, in mature enterprise software contracts. What is exotic in 2026 is remembering to ask for them for an AI platform, because most legal teams still treat these platforms as ordinary SaaS. They are not ordinary SaaS. The value they accumulate for your organization is the value that will be hard to move if you do not plan for it at signature.

09

The three failure patterns to name in the room

Even a well-run evaluation can fail three specific ways, and naming them before they happen is the single cheapest defense against them.

The first is the demo capture. A senior stakeholder sees a demo that answers a real internal question and forms an emotional attachment to the vendor before the scorecard has been filled in. The rest of the evaluation quietly becomes a search for reasons to justify the decision that has already been made. The defense is simple: run the scorecard before the demo, not after. Publish the questions and the scoring rubric to the buying committee before the first vendor conversation. The demo then becomes evidence for the scorecard, not a substitute for it.

The second is the RFP autopilot. The evaluation is delegated to a procurement team that runs the standard SaaS template and never asks the 4 C's questions. The vendors return identical, technically correct answers, and the decision falls to price and brand. The defense is to co-author the RFP with the buying committee and the data leadership, and to require the 4 C's questions to be answered with evidence rather than assertion. A vendor answer that says our platform supports export is not an answer. A vendor answer that names the formats, cadence, and SLA is an answer.

The third is the pilot-to-production trap. A promising pilot on a narrow workload gets extended into production on a broad one, without a re-scoring against the 4 C's at the new scale. The Cost curve that was linear at pilot volume turns sub-linear against you at production volume. The Context that fit inside the vendor's memory layer at pilot scale starts to demand a schema the platform does not support. The Control mechanisms that were adequate for a friendly internal audience are inadequate for a regulated one. The defense is to re-score the platform at every major expansion — not just at signature.

10

A short worked example

Imagine a mid-sized enterprise evaluating three agent platforms for a customer operations workload. Vendor A is a household name with an integrated agent runtime bundled into its cloud data platform. Vendor B is a specialist agent framework with strong open-standards support and a smaller footprint. Vendor C is a scrappy startup with the best demo of the three and a hand-rolled runtime.

On Context, Vendor A owns a proprietary semantic layer that is deeply integrated with its own catalog. Definitions are portable in name but not in shape. Score: yes with caveats. Vendor B expresses entities and metrics in an open format and provides a documented export. Score: yes with evidence. Vendor C has no semantic layer yet and asks the customer to bring their own. Score: no.

On Control, Vendor A has a mature audit and approval framework, integrated with major SIEMs. Score: yes with evidence. Vendor B has the primitives but requires configuration. Score: yes with caveats. Vendor C has logging but no approval routing. Score: no.

On Cost, Vendor A hides token and tool-call cost inside a bundled platform fee, with cost-per-task reporting available only at enterprise tier. Score: yes with caveats. Vendor B exposes cost per task natively. Score: yes with evidence. Vendor C prices at a low seat fee and passes model cost through, but has no unit economics reporting. Score: yes with caveats.

On Choice, Vendor A supports MCP for tool contracts and provides export for definitions but not for evaluations. Score: yes with caveats. Vendor B supports MCP and A2A end to end, with a public migration playbook a customer has already executed. Score: yes with evidence. Vendor C has invented its own tool protocol and has no migration story. Score: no.

The scorecard reveals what the demos would have hidden. Vendor A is a defensible default but concentrates risk in the Cost and Choice columns. Vendor B is the strongest scorecard in aggregate. Vendor C, the demo winner, would have been a two-year regret. A worked example like this one, run inside a buying committee before signature, is worth every hour it takes.

11

What great looks like

A great vendor evaluation ends with three artifacts, not one. The first is the scorecard itself, filled in, dated, and archived. The second is a short memo, one page, that summarizes why the winning vendor won and what the losing vendors would have needed to do to change the outcome. That memo becomes the reference document the next time this evaluation runs — and it will run again, because the AI vendor landscape rotates faster than any category we have priced before. The third is the export contract: the specific clauses that turn the scorecard's Choice column into a legal instrument.

None of this is exotic procurement discipline. It is the same discipline that mature organizations already apply to their largest cloud and data platform contracts. What is different, and what deserves to be named clearly, is the pace. Agent platforms are still an early market. The scorecard you fill in this quarter will need to be re-filled next year. A vendor that scores well today can, and often will, be a different company a year from now — a merger, a pivot, a change in leadership, a shift in pricing strategy. The scorecard is not a one-time gate. It is a practice.

The organizations that will win the next three years in enterprise AI are not the ones that pick the best vendor. They are the ones that build the muscle to re-evaluate the vendor landscape on a regular cadence, with a scorecard that stays consistent even as the vendors change. Context, Control, Cost, and Choice are the four columns of that scorecard. They will still be the four columns in 2029.

12

The Monday move

The reader of this essay does not need a strategy offsite. What most enterprises need this week is a working draft of the scorecard, in a shared document, with the twenty questions filled in for the vendors already in the environment. Not the vendors under evaluation. The ones already signed. The exercise is uncomfortable, because the answers reveal decisions that were made without asking the right questions. It is also cheap, because the artifact it produces is the reference document that will run every future evaluation, and the checklist that will surface the export clauses missing from every current contract.

The next agent platform in the market is already being demoed to somebody in your organization. The scorecard is what turns that demo from a threat to your existing stack into a data point on a durable framework. That framework is the 4 C's. And the only real cost of not writing it down is that the next contract will be signed without it.

"The right question at signature is not which platform has the best demo. It is which one will still be a good decision the day you want to swap the model underneath."
Mini checklist

Try this at work

  • Publish the twenty-question scorecard and the three-point scoring rubric to the buying committee before the first vendor demo, not after.
  • Require every 4 C's answer to arrive with evidence — a format, a cadence, a customer reference — rather than an assertion.
  • Fill in the scorecard for the agent platforms already signed in your environment; the gaps you find are next quarter's contract-amendment agenda.
  • Write the five export clauses — context, open standards, audit, unit economics, migration cooperation — into every new agent platform contract before signature.
  • Re-score the platform at every major expansion, not just at renewal; pilot economics rarely survive contact with production volume.

The Context Advantage is the long-form playbook behind this scorecard — thirty-four chapters on Context, Control, Cost, and Choice, with the role definitions, contracts, and evaluation practices that make the framework operational. Start with the free chapters at [/context-advantage/blog](/context-advantage/blog), or unlock the full book at [/context-advantage/buy](/context-advantage/buy).

Explore the book →
Over to you

If your top agent vendor announced a pivot next month, how many days would it take to move your semantic layer, evaluations, and audit trail to a different runtime — and who owns the answer?

Found this useful? Share it with a teammate.
Share
BricksNotes updates
Liked this? Get the next essay in your inbox.

One thoughtful piece a week on context, control, cost, and choice for data and AI teams. No spam.

By subscribing you agree to receive emails from Team BricksNotes. Unsubscribe anytime.

This is a companion post to The Context Advantage — a living book by Team BricksNotes.