We built a computer use agent once and got it benchmarked at the 2nd position in the biggest benchmark for this field but in the process, we realized one thing, how outdated and inefficient these benchmarks are and how they don t truly capture how these tasks went.So we thought a community led benchmarking arena would be the way to go about. https://coarena.ai/.You just have to put in a task and two side by side windows with hidden models will show up and do the same task and at the end you vote which one did better(or none).We have a leaderboard where we rank all the models based on how they performed on the tasks so far, so please do check that out!Just a little sneak peak: OpenAI seems to dominating in terms of success but for speed it s actually something else.This is still pretty new and we re quite open to ideas and feedback so would appreciate any.