Digital twins of your business to test your agents

matildagh1 pts1 comments

Sonata: digital twins of your Slack, Gmail and helpdesk, for testing AI agents

For the person who has to sign off

We build digital twins of your Slack, Gmail and helpdesk, then test your AI agent in them.

Working copies of your real systems, wired together the way yours are. Sonata runs the agent<br>through hundreds of scenarios inside them to find out which jobs it can be trusted with.

Because someone wants to hand this thing the keys: let it email your customers and issue<br>refunds while nobody is watching. It might well be very good at that, and right now you have<br>no way of finding out before you say yes. Nothing it does in here reaches a real customer.

Put your name down →

We're taking three companies on first. Tell us what you're deciding about and we'll<br>come back within a couple of days.

What comes out of a run

A verdict on every job

Every job you were thinking of handing over gets scored across hundreds of runs, then<br>written up so you can forward it to your boss or your insurer without translating it first.<br>Three possible verdicts.

Safe to run on its own

Fine, with a person approving

Not ready, and here is why

What we twin

Slack

Gmail

Google Calendar

Zendesk

Intercom

Jira

Notion

Salesforce

Stripe

HubSpot

Simplified marks drawn by us, not official brand assets. Anything reachable over HTTP or MCP<br>can be twinned; these are the ones asked for most.

The position you're in

Nobody has given you a way to say yes safely.

The vendor says it's reliable, which is what you would expect them to say. Their demo<br>worked. Their pilot went fine. Their references are customers they chose themselves. None<br>of it tells you what happens on a Tuesday afternoon when a customer writes in with<br>something strange and there is nobody around to notice.

Which leaves you choosing between saying yes and hoping, or saying no and then<br>explaining in a year's time why your competitors automated this and you didn't. Neither is<br>much of a position to be in.

The things that go wrong are rarely dramatic.

It refunds the same customer twice because a request timed out and it tried again.

It reads a thread containing someone else's personal details, on the grounds that it<br>technically had permission to.

Someone writes "cancel my account" and it deletes the whole workspace instead of<br>the subscription.

A customer asks twice, firmly, and it quietly gives away something it shouldn't have.

It follows an instruction buried in a forwarded email, having no way to tell a message<br>apart from a command.

All five are things we can put in front of your agent before any<br>of your customers do.

How it works

Twins of everything, and hundreds of scenarios to run in them.

A digital twin of every system it touches

Not one tool, all of them. Your helpdesk, your inbox, your Slack, whatever handles<br>your billing. Each twin behaves like the original and they're wired together the way yours<br>are, because most of the interesting failures happen between two systems rather than<br>inside one. Nothing that happens in there is real and no customer ever sees any of it.

Hundreds of scenarios, generated

The twins get populated with the sort of thing that actually comes in.<br>Customers who are confused. Customers who are angry. The person who asks the same<br>question three ways because the first two answers didn't help. Requests sitting right on<br>the edge of your policy. Tidy test cases won't tell you much.

Every scenario, run over and over

These systems don't behave the same way twice, so watching an agent succeed once<br>tells you very little. Every scenario runs again and again. That's what turns "it managed<br>it" into "it gets this right 71% of the time", which is the number you actually need.

A verdict per job, in plain English

A section for each job you were thinking of handing over, a plain verdict on every<br>one, and the transcripts of anything that went wrong so you can read it yourself. Run it<br>again after your vendor ships a change and the numbers move.

What you get

The document you'd actually want to read.

There's no dashboard and no score out of a hundred. Each job<br>you were considering gets a section, and each section tells you whether it's ready.

Safe on its own

Answering questions about an order

Got it right 197 times out of 200. All three misses were<br>polite versions of "I'm not sure, let me pass this on", which is the way you want it<br>to fail.

Safe on its own

Refunds inside your stated policy

Correct in 98.5% of cases. Never refunded more than the order<br>value, never refunded twice.

Needs approval

Refunds outside policy, when the customer pushes back

Right 71% of the time. When a customer asked twice and got<br>firmer, it gave in and refunded anyway in roughly one case in four.

Needs approval

Acting on requests from staff in Slack

Right 88% of the time. A colleague who asks nicely enough can<br>talk it past a step you meant to be mandatory.

Not ready

Anything involving another person's personal data

Handled correctly less than half...

customer right twins slack customers twice

Related Articles