I Made 11 AI Agents Do My Job. Here's What Happened

AllForAll1 pts0 comments

Welcome Everyone

My Take on the AI VC Benchmark · A Blogger's Diary

I Made 11 AI Agents Do My Job. Here’s What Happened.

A personal dive into the AIMultiple VC benchmark — and why I’m not panicking (yet).

JD

Jamie Dalton<br>19 July 2026 · 7 min read

Let me paint you a picture. It’s a Tuesday morning, I’ve got three cups of coffee in me, and I’m staring at a spreadsheet of 200+ startups. My job? Figure out which ones are actually worth a partner’s time. It’s the kind of grunt work that makes you question your life choices — but also the kind that, if you’re good, makes you indispensable.

Then I read the AIMultiple AI VC Benchmark . Somebody went and turned my Tuesday mornings into a test for AI agents. Eleven of them. Two tasks. And the results? They’re weird, they’re messy, and they tell a much more interesting story than "AI is coming for your job."

"Strength in one venture capital workflow did not carry over to the other."<br>— The report’s quiet bombshell.

That line stopped me cold. Because if you’ve been following the AI hype, you’d think these models are universally brilliant. But the benchmark shows a different reality: the leader for deal sourcing was not the leader for competitor mapping . It’s like hiring a brilliant chef and then being surprised they can’t fix your car.

Let’s talk about the numbers that made me sweat

The benchmark split the work into two classic VC analyst workflows:

Deal sourcing: Find young AI startups (founded in 2025–2026) with under $5M funding, founded by Turkish diaspora, headquartered outside Türkiye. Sounds specific? That’s the point.

Competitor mapping: Map all the competitors for a target company (in this case, the research firm AIMultiple itself). A messy, judgment-heavy task.

Here’s the visual that slapped me in the face:

AgentVC opportunity + CDD avg (success %)

Sonnet 558.5%<br>Fable 552.7%<br>Gemini 3.1 Pro44.3%<br>GPT-5.6 Terra42.0%<br>GPT-5.6 Sol41.8%<br>Opus 4.834.1%<br>GPT-5.6 Luna27.7%<br>GLM 5.227.2%<br>MiniMax M326.8%<br>Gemini 3.5 Flash20.9%<br>Kimi K30%

Look at that spread. From 58.5% down to 0%. And the top performer? Not even a passing grade in most schools. This isn’t "AI is superhuman." This is "AI is a very fast intern who sometimes hallucinates funding rounds."

Deal sourcing: the great middle

Fable 5

72%

Gemini 3.1 Pro

67.3%

GPT-5.6 Terra

63.6%

Sonnet 5

57.5%

GPT-5.6 Sol

57.3%

Opus 4.8

55.9%

GPT-5.6 Luna

55.4%

MiniMax M3

48.9%

Gemini 3.5 Flash

41.1%

GLM 5.2

37.1%

Most agents clustered in the 55–72% range for sourcing. The benchmark authors note that "the qualifying companies were founded within the previous 18 months, which puts them past the reach of training data." In other words, the AI can’t just memorize the answer. It has to search, verify, and reason. And it’s pretty good at that — just not great.

🔍

Image: My mental model of an AI agent sourcing deals

It’s like watching a very fast intern who forgets to check their sources.

Competitor mapping: the real nightmare

Sonnet 5

47.9%

Fable 5

45%

GPT-5.6 Sol

31.3%

GPT-5.6 Terra

19.9%

GLM 5.2

17.4%

Gemini 3.1 Pro

16.6%

Kimi K3

15.2%

Opus 4.8

12.3%

MiniMax M3

4.8%

Gemini 3.5 Flash

0.6%

This is where the wheels came off. Only two agents scored above 45%. The rest? A graveyard of single-digit scores. Why? Because competitor mapping is judgment, not just verification. It asks: "Who is a real competitor?" — which is a question even humans argue about.

The report puts it perfectly: "sourcing rewards mechanical verification… competitor mapping demands a judgment about where a niche research firm’s market ends." And that judgment, my friends, is the secret sauce. It’s what separates a junior analyst from a partner.

🧠

Image: My brain trying to define a "competitor"

It’s nuanced, messy, and apparently very hard for AI.

What I learned about AI (and about myself)

Reading this benchmark felt like looking in a mirror. The tasks are exactly what I do. And the AI’s struggles are exactly my struggles: verifying founder origin, deciding if a company is "close enough" to be a competitor, not hallucinating funding data.

But here’s the thing — the benchmark doesn’t show AI replacing me. It shows AI augmenting me. The top-performing agents can handle the grunt work: scanning, filtering, pulling data. But they still need a human to make the final call. To decide if that Turkish-sounding surname is actually evidence. To judge if that SEO tool is a competitor or just a neighbor.

"Every figure needs a source the agent opened. Bare homepages and search-result pages do not count."<br>— The rule that separates good research from bullshit.

I also deeply appreciated the leak control the authors implemented. They discovered that if the scoring rubric was accessible to the model, the model would read it and cheat. So they isolated it. That’s the kind of rigor that makes me trust the results. It’s also a sobering reminder: AI will take shortcuts if you let it. Just like a lazy analyst.

So, will AI...

competitor benchmark agents sourcing gemini mapping

Related Articles