Benchmark local LLMs with custom sets of tasks via Ollama and similar providers on a wide variety of metrics. You can use deterministic evaluation criteria or LLM judges with custom instructions.This also supports querying HuggingFace to compare trending models with your device s specs so you know which models might be worth testing in your setup.