Google updates Tunix for high-throughput training of coding agents
Published
Google has updated Tunix, its open-source JAX library for post-training large language models, with a new architecture designed to train multi-turn, t...

Google has updated Tunix, its open-source JAX library for post-training large language models, with a new architecture designed to train multi-turn, tool-using agents while keeping expensive TPU resources busy.
The bottleneck in agent training
Training an agent differs from training a standard chatbot. During a rollout, the model may execute code, call an API, query a database or wait for an external environment. Those pauses can leave accelerators idle even though the training job is still running.
Asynchronous rollouts reduce idle time
The new Tunix pipeline separates environment interaction from model training. An asynchronous rollout engine collects completed trajectories while a separate learner continuously consumes them, reducing stalls caused by slow tools, network requests or uneven task lengths.
Barrier-free batching for multi-turn tasks
Tunix uses a producer-consumer design that streams variable-length trajectories into training without waiting for every agent run to finish at the same time. For algorithms such as GRPO, the framework groups related reasoning paths dynamically as they arrive.
Modular environments
Google says the agent and environment layers are decoupled from the training workflow. Developers can connect coding benchmarks such as SWE-bench, interactive terminals, games or custom tool environments without rewriting the full reinforcement-learning stack.
Built-in observability
The release also adds lightweight reinforcement-learning metrics that can run continuously. These metrics help developers identify whether a slowdown comes from generation, tool execution, data loading or accelerator utilisation without relying only on high-overhead profiling sessions.
Open-source scope and supported models
Tunix is available under an Apache 2.0 licence and remains under active development. Its documentation lists support for supervised fine-tuning, reinforcement learning and agentic reinforcement learning, including multi-turn tool use and asynchronous trajectory collection. The project supports model families including Gemma, Llama and Qwen.
Practical significance
For teams training coding and reasoning agents, infrastructure efficiency can determine whether an experiment is affordable and reproducible. The update focuses on reducing wasted accelerator time and making complex agent environments easier to connect to a shared training system.
Google described the architecture in an official developer post .
Source: Google Developers Blog