Open Source Tax Engine outperforming GPT sol and Fable 5

asmigulati2 pts0 comments

OpenTax by Invaro | Turn any AI into a precise tax preparer

githubInstall

The deterministic tax engine for AI agents<br>Turn any AI into a<br>precise tax preparer<br>Language models are probabilistic. Tax law is not. OpenTax is a deterministic computation engine: the same facts produce the same return, every time, with every line traced to the statute that produced it. It is the only open-source tax engine on the market, and the highest-scoring engine ever measured against one.<br>Read the sourceInstall in 30 seconds

Deterministic<br>Same facts, same return, every single time. Rules and exact arithmetic, not sampling. Nothing an LLM can drift on.

The only open-source one<br>There is no other open-source tax engine on the market. AGPL-3.0, every rule and every test in public.

The HIGHEST-scoring one<br>96% exact returns on TaxCalcBench. No model, no engine, open or closed, has ever scored higher.

The model isn't the ceiling.<br>The engine is.<br>TaxCalcBench, the industry benchmark for AI tax work, asks a model to prepare 50 complete US tax returns. A return only counts if every line is exact. Claude Sonnet on its own gets 6% of returns right. The same Sonnet, using OpenTax, gets 96%. The highest score ever recorded, above every top model, with or without web search. The engine amplifies whatever model you point at it.

before<br>Claude Sonnet, on its own

6%

+ connect the OpenTax MCP<br>one URL, two minutes — nothing else changes

after<br>the same Sonnet, same prompts — with the engine

96%

correct returns · TaxCalcBench TY25HIGHEST SCORE EVER RECORDED<br>Claude Sonnet + OpenTax MCP96%

GPT-5.6 Sol + web search58%

GPT-5.5 + web search54%

Claude Fable 5 + web search34%

Claude Sonnet, unaided6%

48/50<br>returns exact to the dollar

98.2%<br>of all 820 scored lines exact

16×<br>better than the same model alone

and the two misses?<br>We traced both to reference returns whose worksheets are inconsistent with their own inputs — no published model run has matched those two cases either. We filed the bug upstream. An engine precise enough to audit the exam it's taking.

methodology: TaxCalcBench TY25 by column-tax, 50 returns, strict scoring (every line exact) · scored with the benchmark's own evaluator against its reference returns · leaderboard rows from its published July 2026 results · agent runs cold, pass@1

01<br>Cited, verbatim<br>Every rule in the corpus carries its statute: 26 U.S.C. section, effective window, and a verbatim source excerpt. No number enters the engine on model recall.

02<br>Provable<br>Every computation returns a Merkle-rooted proof: the facts, the rules, the arithmetic. Re-verify any answer offline, byte for byte.

03<br>Documents in, facts out<br>A deterministic compiler turns W-2 and 1099 boxes plus birth dates into engine facts. Ages, dependent classifications, penalties, excess withholding, all with zero judgment calls.

04<br>Refuses loudly<br>Anything outside the corpus refuses with a named reason instead of guessing. A wrong answer that looks right is the one failure mode we do not ship.

05<br>Federal + state<br>Individual, corporate, and fiduciary: Forms 1040, 1120, and 1041, complete line sets in one call. State coverage across 29 states and growing, every line under the same citation discipline.

What your agent does with it<br>OpenTax is the computation layer under whatever AI you already use — Claude, ChatGPT, Cursor, or your own agent. The model does the reading and the writing; the engine does the math. In practice:

tax & accounting firms<br>Review every return with a second engine<br>Connect OpenTax to Claude and drop in the client file: W-2 and 1099 boxes compile deterministically into facts, the full 1040 line set recomputes to the cent, and every line your software disagrees with gets flagged — with the statute attached.<br>→ to your agent, verbatim<br>“Here's the Hendersons' W-2 and the draft 1040 from our software. Recompute every line and flag anything that doesn't match — citation for each flag.”

freelancer & S-corp clients<br>Quarterly estimates that survive lumpy income<br>Safe-harbor prongs (100% / 110% of prior year) and the § 6654 annualized-installment worksheets are encoded rules, not vibes. Windfall in Q3? The engine computes the installment that actually avoids the penalty.<br>→ to your agent, verbatim<br>“Client draws a $60k salary and just booked a $180k Q3 licensing windfall. Compute the annualized Q3 estimate and show the safe-harbor comparison.”

advisory engagements<br>Entity elections, modeled to the cent<br>S-corp election letters stop being hand-waves: run sole-prop vs S-corp at a reasonable salary — SE tax versus payroll taxes, the QBI wage limit doing its thing, entity and personal pass in one call.<br>→ to your agent, verbatim<br>“Sole prop, $220k net. Model an S-corp election at a $110k salary: SE vs payroll tax, QBI under the wage limit, total federal delta — cited.”

financial planners<br>Find the cliffs before your client does<br>The solver sweeps any input across a range and reports every cliff and kink on the way: Social Security taxability...

engine model returns line opentax claude

Related Articles