Dulus is the operating layer for inference34 model providers unified into one runtime • 100+ inference backends through LiteLLM • 5.9B tokens processed through Claude • 98.8% prompt cache hit rate • Lookback compresses 2,000 conversation turns into a 20-turn inference window • Python Console avoided 1.5–2M tokens by keeping large datasets outside the model s context • 2,186+ MCP tools • 100,000+ skills • 400+ commits in 90 days • 58 public releases.None of your inference projects are throwing out these numbers rn except for $P0