AI Agent Contracts: What to Put in Writing in 2026
Buying or building an AI agent? The contract decides who pays when it gets things wrong. Here is what to define, guarantee and negotiate before you sign.
Most AI agent deals are signed on the strength of a demo. The agent answers three questions perfectly, everyone nods, and the contract that follows is a standard SaaS template with "AI" pasted into the product name. Then the agent goes live, gets something wrong in front of a real customer, and nobody can point to the clause that says whose problem that is.
An AI agent is not ordinary software. It makes decisions, acts on your behalf, touches your data and your customers, and it does not behave identically every time you run it. That last point breaks the assumptions most software contracts are built on. A normal app either works or throws an error. An agent can be confidently wrong, and a contract that never mentions accuracy, escalation or liability leaves you carrying that risk alone.
This is the contract you actually want, whether you are buying an agent from a vendor or hiring someone to build one. If you are still choosing who to work with, read it alongside our guide on how to choose an AI agent development company. This is about what goes in writing once you have.
Define what "working" actually means
The single most expensive gap in AI agent contracts is the absence of a definition of done. "The agent handles customer support" is not a specification. It is a wish.
Pin down scope in concrete terms. Which requests is the agent responsible for, and which are explicitly out of scope? What is the expected resolution rate, measured how, over what sample? A support agent might be contracted to fully resolve 70% of tier-one tickets without a human, with everything else handed off cleanly. Write that number down. Without it, "working" means whatever the vendor says it means on renewal day.
Agree on how accuracy is measured before launch, not after a dispute. That usually means a shared test set of real cases, a scoring method you both accept, and a minimum score the agent has to hit to be considered live. Our guide on how to evaluate AI agents covers how to build that evaluation so it reflects your real traffic rather than a cherry-picked demo.
Service levels for something that thinks
Classic SLAs cover uptime and response time. Those still matter, but an agent needs two more.
The first is an accuracy or quality floor: a measured threshold the agent must stay above once in production, checked on an agreed cadence, with a remedy if it drifts below. Models get swapped, prompts get edited, your data changes, and performance can quietly degrade. The contract should say what happens when it does.
The second is escalation and fallback. When the agent is unsure, or the request is out of scope, or a dependency is down, what does it do? A good contract specifies that the agent hands off to a human or a safe default rather than guessing. Silence here is how you end up with an agent inventing a refund policy at 2am.
Tie the remedies to outcomes, not effort. If accuracy falls below the floor, you want service credits or a fix timeline, not a meeting. This connects directly to how you pay in the first place, which we cover in outcome-based AI pricing.
Data, privacy and who owns the output
Three separate questions hide inside "data", and each needs its own answer.
- Your input data. Can the vendor use your documents, tickets and customer records to train or improve their models, including models other clients use? For most businesses the answer should be a firm no, written down. Get the data residency, retention and deletion terms in the contract, especially if you operate under the GDPR.
- The agent's output. Who owns what the agent produces? If it drafts content, writes code or generates analysis, you want clear ownership of those outputs, the same way you would for custom software. Our piece on who owns the code applies here almost line for line.
- The configuration. The prompts, tools, workflows and fine-tuning that make the agent yours have real value. Decide upfront whether that work is yours to keep or the vendor's to lock away.
Security belongs in writing too. If the agent can take actions, not just answer questions, the contract should set out what it is permitted to touch and how those permissions are controlled. Our guide on securing AI agents is a good checklist for what the security exhibit should cover.
Liability when the agent gets it wrong
This is the clause everyone skips and later wishes they had read. An agent that acts can cause real harm: a wrong price quoted, bad advice given, a transaction approved that should not have been.
Standard software contracts cap liability near the fees paid and disclaim almost everything else. That allocation made sense when software only displayed information and a human made every decision. It makes much less sense when the software is the one deciding. In the EU this is no longer just a negotiating point: the revised Product Liability Directive now treats software as a product, which changes who can be held responsible when it causes damage.
Negotiate the split deliberately. Where the agent operates inside guardrails you approved, responsibility should sit differently than when the vendor's model simply failed. Define what "human in the loop" means for high-stakes actions and who signs off. The point is not to win every clause, it is to make sure a real person owns each failure mode before one happens.
Exit terms, before you need them
The best time to agree how you leave is the day you join. Agents accumulate value that is easy to trap: conversation history, tuned prompts, evaluation sets, integrations.
Make sure the contract gives you your data back in a usable format and lets you take the configuration you paid to develop. The EU Data Act strengthens your right to switch providers and port data, but a specific exit clause saves you the argument. Watch for soft lock-in too: proprietary formats, undocumented workflows, or a model only the vendor can run. If leaving means rebuilding from scratch, you do not really have a choice at renewal, and the vendor knows it.
A short checklist before you sign
- Scope and definition of done, with a measured resolution or accuracy target.
- An evaluation method agreed in advance, on your real cases.
- An accuracy floor and escalation path written into the SLA, with remedies.
- Data terms: no training on your data without consent, clear residency, retention and deletion.
- Ownership of outputs and of the configuration you funded.
- Liability allocated to match who actually makes each decision.
- Exit rights: data export, config portability, no hidden lock-in.
None of this slows a good project down. It mostly forces the conversations that a demo lets you avoid, and those conversations are exactly where you learn whether a vendor understands their own agent well enough to stand behind it. If they resist putting an accuracy floor or a clean exit in writing, that answer is worth more than the demo was.
Written by
Rafael Costa
Software Engineer & Technical Writer
Rafael is a software engineer at Lusivision who writes about web development, cloud architecture and applied AI. He has spent over a decade shipping production software for companies across Europe and enjoys turning hard technical topics into clear, practical guides.
View all articles