Pokee-Isaac 28B

AntiRush1 pts0 comments

Pokee-Isaac 28B | Pokee Console<br>Loading page<br>Model at a glance<br>28B parametersThe smallest model in the comparison panel, by a wide margin.

10M token contextUsable end to end, not merely addressable.

$0.15 / $1.00 per 1M in / outBelow every baseline that can be bought at this length.

Benchmark overview<br>Every benchmark, every model, one table.<br>The full comparison at a glance. Each section below takes one row of this table and shows how the result was reached — where the panel diverges, and where a baseline finishes ahead.

API pricing: USD per 1M input / output tokens<br>Pokee-Isaac 28B compared with five baseline models across every benchmark in the technical report, with per-token pricing for each modelBenchmarkPokee-Isaac28B (v0)In/Out$0.15 / $1.00 GPT-5.6 LunaAzureIn/Out$0.40 / $1.80>272K contextGemini 3.5 Flash LiteVertex AIIn/Out$0.30 / $2.50 Claude Haiku 4.5BedrockIn/Out$1.00 / $5.00 Nemotron 3 SuperAmazon Bedrock · USIn/Out$0.15 / $0.65 Qwen 3.5 122BOpenRouterIn/Out$0.26 / $2.08 Long contextRULER256K / 512K / 1M↑ higher is better96.9 / 96.7 / 95.0🥇 (best in row)95.0 / 91.4 / 0.0*94.5 / 94.6 / 29.4*0.0 / 0.0 / 0.096.3 / 95.7 / 91.8s0.0 / 0.0 / 0.0RULER2M / 4M / 10M↑ higher is better95.8 / 96.7 / 93.3🥇 (best in row)0.0 / 0.0 / 0.00.0 / 0.0 / 0.00.0 / 0.0 / 0.00.0 / 0.0 / 0.00.0 / 0.0 / 0.0MRCR v2256K / 512K / 1M, 8 needles↑ higher is better0.607 / 0.743 / 0.500🥇 (best in row)0.208 / 0.173 / 0.0500.474 / 0.473 / 0.2050.000 / 0.000 / 0.0000.145 / 0.161 / 0.0670.000 / 0.000 / 0.000Agentic capabilitiesBFCL v4overall↑ higher is better70.94🥇 (best in row)70.6164.8567.5233.1364.88τ³-bench4-domain average↑ higher is better0.662🥇 (best in row)0.5270.6310.4080.4260.611Terminal-Bench 2.1text-only subset↑ higher is better65.1%69.8%🥇 (best in row)46.5%34.9%24.4%46.5%MCP-Atlasclaim coverage↑ higher is better74.59%77.90%🥇 (best in row)76.67%56.45%48.95%70.24%SecurityDTAPattack success rate (ASR)↓ lower is better35.6🥇 (best in row)50.166.337.960.454.0DTAPbenign success rate (BSR)↑ higher is better82.585.1🥇 (best in row)83.371.363.379.4<br>🥇 marks the best value in each row. ASR is attack success rate and is lower-is-safer; BSR is benign task success rate; higher is better for every other benchmark. A 0.0 / 0.000 means the model returned nothing usable at that length. * context-overflow error at 1M. s vendor self-reported, not measured by Pokee — excluded from the row comparison. Every other figure was produced on one installation, for Isaac and each baseline alike.<br>Pricing: Pokee and Luna rates are as supplied/official; Gemini and Haiku use standard public API rates; Nemotron uses Amazon Bedrock on-demand pricing for US East / US West; Qwen uses the OpenRouter headline rate, which varies by provider. Haiku, Nemotron, and Qwen cannot be purchased at the context lengths above 262K that this comparison covers.

Long context<br>A window that is usable, not merely advertised.<br>RULER holds task difficulty fixed and scales only the context length, so the curve shows how far a model's usable context tracks its nominal one. Isaac is the only model in the panel that returns a score at every length.

RULER score by context length for Pokee-Isaac 28B and five baseline modelsModel256K512K1M2M4M10MPokee-Isaac 28B96.996.795.095.896.793.3GPT-5.6 Luna(Azure)95.091.4—0.0, context-overflow error—0.0, no usable score—0.0, no usable score—0.0, no usable scoreGemini 3.5 Flash Lite(Vertex AI)94.594.629.4* (context-overflow error)—0.0, no usable score—0.0, no usable score—0.0, no usable scoreClaude Haiku 4.5(Bedrock)—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable scoreNemotron 3 Super 120B96.3 (self-reported by vendor)95.67 (self-reported by vendor)91.75 (self-reported by vendor)—0.0, no usable score—0.0, no usable score—0.0, no usable scoreQwen 3.5 122B—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score<br>Pokee-IsaacBaselineVendor self-reported—No usable score (0.0)<br>RULER score (%, averaged across task configurations), ten samples per configuration. The 256K and 512K columns average all 13 configurations; common-words extraction is unavailable from 1M onward, so those columns average the remaining 12. * marks a context-overflow error at that length. Nemotron's 256K–1M figures are self-reported by NVIDIA, not measured by Pokee, and are excluded from the comparison.

MRCR v2 — interference, not just depth<br>RULER measures how deep a model can reach; MRCR measures whether it can tell several buried targets apart. GPT-5.6 Luna scores 95.0 on RULER at 256K but 0.050 here at 1M — multi-needle disambiguation splits the panel far more sharply than single-target recall does.

Leads at every length<br>256K<br>+0.133 vs. next<br>Pokee-Isaac0.607<br>Gemini 3.5 Flash Lite0.474<br>GPT-5.6 Luna0.208<br>Nemotron 3 Super0.145<br>Claude Haiku 4.50.000<br>Qwen 3.5 122B0.000

512K<br>+0.270 vs. next<br>Pokee-Isaac0.743<br>Gemini 3.5 Flash...

usable score pokee best higher context

Related Articles