Private Inference for Coding Agents

Otsar_zro1 pts0 comments

Zro

Open-model inference<br>Private inference for coding agents.<br>Fast, EU-based endpoint for open-weight models. Zero data retention, zero training, and optimized for long-context workloads.<br>Connect your coding agentView pricing

CodePrivate<br>ModelsOpen<br>SetupMinutes

Agent integrations

Claude Code<br>Codex<br>Cursor<br>Cline<br>Pi<br>OpenClaw<br>OpenCode<br>Hermes

API<br>AgentsCLI + IDE<br>InfraEU regions

EU infrastructure<br>Private inference on EU infrastructure.

Zro serves coding-agent workloads from EU infrastructure with zero request retention and no training on customer data.<br>Location<br>EU regions

Retention<br>Zero by default

Training<br>Never

Zro is tuned for long-context, multi-turn coding sessions. Under the endpoint, MoonMath applies HyperQuant compression, custom kernels, and hardware-aware deployment.<br>Compression<br>HyperQuant

Kernels<br>Custom attention

Hardware<br>AMD · NVIDIA · TPU

Performance<br>Built for coding‑agent workloads.<br>View technology→

Coding agents<br>$ zro launch claude, codex, opencode, hermes, openclaw, pi$zrolaunchclaudeclaude<br>Install the npm package, log in once, then launch supported coding tools with temporary session config.<br>View integrationsCreate API key

@moonmath-ai/zro npm packageClaude Code · Codex CLI<br>npm install -g @moonmath-ai/zro<br>zro login<br>zro launch claude<br>zro launch codexnpm install -g @moonmath-ai/zro<br>zro login<br>zro launch claude --model minimax-m3<br>zro launch codex --model glm-5.2Supports Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, and Pi.

Coding models<br>MiniMax M3MiniMaxBuilt with MiniMax M3GLM-5.2Z.aiFrom Z.aiDeepSeek V4 ProComing soon<br>Kimi K2.6Coming soon

Open coding models.<br>One endpoint.<br>Open-source models are becoming competitive with closed-source systems for coding tasks. iZro starts with MiniMax M3 and GLM-5.2, with more open coding models coming soon.

Details<br>Built for private inference.

What is Zro?Zro is a private inference endpoint for coding agents. It serves open-weight coding models from EU infrastructure, with zero request retention, no training on customer data, and setup paths for the tools developers already use.

Can I use existing OpenAI-compatible clients?Yes. Zro exposes OpenAI-compatible access for chat completions, so existing clients and agent tools can point at the Zro base URL.

Can I use Anthropic-compatible clients?Yes. Zro also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.

How do I use Zro with coding tools?Use the @moonmath-ai/zro npm package for launch-supported tools: run zro login once, then launch your harness with zro launch. The integrations page also covers manual setup for Cursor and Cline.

Do you retain prompts or completions?No. Prompt and completion bodies are not retained by default after inference is processed.

Is it fast?Yes. Zro is built for responsive, streaming inference, so developer tools, agents, and production apps do not have to trade speed for privacy.

Which models can I use?MiniMax M3 and GLM-5.2 are available now, with more open-model options being added across regions.

Are prompts or completions used to train models?No. Customer prompts and completions are never used for training, fine-tuning, evaluations, analytics, or dataset creation.

Where does inference run?Zro runs on privacy-forward EU infrastructure. Current regions include Finland and France.

How does billing work?Plans start at $20/month for $60 of inference spend. Usage packs are available without a subscription and expire after 90 days. Plan spend resets monthly.

coding inference launch models open tools

Related Articles