Hi HN,Today I m showcasing Trunchbull, a benchmarking platform designed for authoring benchmarks and running them against different models. We have direct support for benchmarks that use the harbor authoring system, custom tool authoring via the vercel ai sdk and configuration limits.We ve also already imported terminalbench 2.0, as a sort of proof of concept that our harbor task orchestrator works, although you currently need a paid account as we are provisioning sandbox environments.I ve made several popular benchmarks publicly available for testing. You dont need an account or your credit card information to access these:[Public benchmark demos](https://trunchbull.dev/sandboxes)[GSM8K math reasoning](https://trunchbull.dev/try/gsm8k)[SkateBench](https://trunchbull.dev/try/skatebench)[ARC-Challenge](https://trunchbull.dev/try/arc-challenge)[TruthfulQA](https://trunchbull.dev/try/truthfulqa-mc1)[Medical AI Failure Atlas](https://trunchbull.dev/try/medical-ai-failure-atlas)These demos let you pick from a preselected list of models, and will systematically test the selected models against their case scenarios.Any and all feedback is welcome, but i m particularly interested in: - knowing what kind of benchmark evidence youd like to expect - any improvements on our benchmark run page, anything that can provide clarity or better understanding of the benchmark u just ran. - what you d prefer to see on the overview page. - improvements on our documentation - better configuration and spend limits.