Skills x2, 108 Runs: Optimising Efficacy and Token Efficiency

darvh1 pts0 comments

Ponytail vs SignalBench has the runs albeit not in an organised manner. Models showed memorization of the answers, which is kinda expected these days.More than the skills, it was fun getting the benchmark running :).Let me know if the preamble on how the bench was conducted or constructed is flawed.Blog Post: darvh.com/posts/when-coding-agents-raced-through-108-bugs/Bench: github.com/darvh/benchSignal: github.com/darvh/signal

bench signal darvh skills runs github

Related Articles