Sonata: digital twins of your Slack, Gmail and helpdesk, for testing AI agents
For the person who has to sign off
We build digital twins of your Slack, Gmail and helpdesk, then test your AI agent in them.
Working copies of your real systems, wired together the way yours are. Sonata runs the agent<br>through hundreds of scenarios inside them to find out which jobs it can be trusted with.
Because someone wants to hand this thing the keys: let it email your customers and issue<br>refunds while nobody is watching. It might well be very good at that, and right now you have<br>no way of finding out before you say yes. Nothing it does in here reaches a real customer.
Put your name down →
We're taking three companies on first. Tell us what you're deciding about and we'll<br>come back within a couple of days.
What comes out of a run
A verdict on every job
Every job you were thinking of handing over gets scored across hundreds of runs, then<br>written up so you can forward it to your boss or your insurer without translating it first.<br>Three possible verdicts.
Safe to run on its own
Fine, with a person approving
Not ready, and here is why
What we twin
Slack
Gmail
Google Calendar
Zendesk
Intercom
Jira
Notion
Salesforce
Stripe
HubSpot
Simplified marks drawn by us, not official brand assets. Anything reachable over HTTP or MCP<br>can be twinned; these are the ones asked for most.
The position you're in
Nobody has given you a way to say yes safely.
The vendor says it's reliable, which is what you would expect them to say. Their demo<br>worked. Their pilot went fine. Their references are customers they chose themselves. None<br>of it tells you what happens on a Tuesday afternoon when a customer writes in with<br>something strange and there is nobody around to notice.
Which leaves you choosing between saying yes and hoping, or saying no and then<br>explaining in a year's time why your competitors automated this and you didn't. Neither is<br>much of a position to be in.
The things that go wrong are rarely dramatic.
It refunds the same customer twice because a request timed out and it tried again.
It reads a thread containing someone else's personal details, on the grounds that it<br>technically had permission to.
Someone writes "cancel my account" and it deletes the whole workspace instead of<br>the subscription.
A customer asks twice, firmly, and it quietly gives away something it shouldn't have.
It follows an instruction buried in a forwarded email, having no way to tell a message<br>apart from a command.
All five are things we can put in front of your agent before any<br>of your customers do.
How it works
Twins of everything, and hundreds of scenarios to run in them.
A digital twin of every system it touches
Not one tool, all of them. Your helpdesk, your inbox, your Slack, whatever handles<br>your billing. Each twin behaves like the original and they're wired together the way yours<br>are, because most of the interesting failures happen between two systems rather than<br>inside one. Nothing that happens in there is real and no customer ever sees any of it.
Hundreds of scenarios, generated
The twins get populated with the sort of thing that actually comes in.<br>Customers who are confused. Customers who are angry. The person who asks the same<br>question three ways because the first two answers didn't help. Requests sitting right on<br>the edge of your policy. Tidy test cases won't tell you much.
Every scenario, run over and over
These systems don't behave the same way twice, so watching an agent succeed once<br>tells you very little. Every scenario runs again and again. That's what turns "it managed<br>it" into "it gets this right 71% of the time", which is the number you actually need.
A verdict per job, in plain English
A section for each job you were thinking of handing over, a plain verdict on every<br>one, and the transcripts of anything that went wrong so you can read it yourself. Run it<br>again after your vendor ships a change and the numbers move.
What you get
The document you'd actually want to read.
There's no dashboard and no score out of a hundred. Each job<br>you were considering gets a section, and each section tells you whether it's ready.
Safe on its own
Answering questions about an order
Got it right 197 times out of 200. All three misses were<br>polite versions of "I'm not sure, let me pass this on", which is the way you want it<br>to fail.
Safe on its own
Refunds inside your stated policy
Correct in 98.5% of cases. Never refunded more than the order<br>value, never refunded twice.
Needs approval
Refunds outside policy, when the customer pushes back
Right 71% of the time. When a customer asked twice and got<br>firmer, it gave in and refunded anyway in roughly one case in four.
Needs approval
Acting on requests from staff in Slack
Right 88% of the time. A colleague who asks nicely enough can<br>talk it past a step you meant to be mandatory.
Not ready
Anything involving another person's personal data
Handled correctly less than half...