Today we’re introducing Valar-27B, a scientific reasoning model trained using Actuator to lead complete investigations. In our paper-reproduction evaluation, it scored above GPT-5.5 and Claude Opus 4.8 and reached 88% of GPT-6 Astra’s score and 89% of Claude Opus 5.5’s. All in a 27-billion-parameter investigator.

We built Valar because the cost of a poor scientific decision can be far larger than the cost of generating it. An experiment can run correctly and still test the wrong hypothesis. A useful investigator needs to recognize when new evidence undermines an earlier assumption and decide whether a failed experiment warrants another attempt or a different approach. As agents carry out more work autonomously, that judgment determines how much of their work is useful.

Valar is trained for this kind of long-horizon reasoning. It directs investigations over multiple exchanges with its tools, interpreting results and choosing subsequent experiments in the context of what came before. Coding agents handle implementation and execution. That division lets us concentrate a smaller model’s capabilities on scientific judgment while benefiting from continuing advances in coding tools.

Scientific reasoning across six domains

Valar was trained on 812 demonstrations spanning biology, chemistry, earth and climate science, materials science, physics and machine learning. Paper reproduction tests its reasoning across those fields: we remove a result figure from a published paper and ask the investigator to reproduce it. The model has to interpret the method, direct an implementation and assess what the executed experiment actually established.

On the 28 held-out problems shared with our frontier comparisons, Valar scored 0.831, about 5% above the 0.790 scored by both GPT-5.5 and Claude Opus 4.8. Valar used our current coding agent; those two models used an earlier one. GPT-6 Astra and Claude Opus 5.5 scored 0.945 and 0.936, respectively.

Scientific reproduction scores on 28 held-out problems: GPT-6 Astra 0.945, Claude Opus 5.5 0.936, Valar-27B 0.831, GPT-5.5 0.790 and Claude Opus 4.8 0.790.
Figure 1 Mean scientific reproduction score on 28 shared held-out problems.

Valar produced 10 reproductions scoring at least 0.9, compared with five for GPT-5.5 and three for Opus 4.8. It also exceeded Opus 4.8 on four of the five graded dimensions, including reproducing the scientific claim and the fidelity of the implementation.

In one groundwater-modeling investigation, Valar requested checks of boundary conditions and limiting cases, read the execution log and asked for a second review of the numerical results. The completed experiment recovered approximation errors close to those reported in the paper. Together with results across other fields, this suggests scientific reasoning that transfers to unfamiliar papers. We expect broad-domain training to support stronger out-of-distribution generalization, including investigations that combine methods from different fields.

A smaller investigator with a choice of tools

We also evaluated Valar with Claude Sonnet 5.5 as its coding agent. Our strongest Valar configuration for that setup scored 0.748, about 12% above Opus 4.8 at maximum effort and 94% of Opus 5.5’s score at maximum effort. It outscored Opus 4.8 on 18 of 21 problems.

Sonnet configuration on 21 development problems: Claude Opus 5.5 0.794, Valar-27B 0.748 and Claude Opus 4.8 0.671.
Figure 2 Mean scientific reproduction score with Sonnet as the coding agent on 21 development problems.

The choice of coder lets teams adapt the system to their research and deployment requirements. Work involving proprietary methods or restricted datasets can use a privately hosted investigator; a fully air-gapped setup would also need local coding and experimental tools.

What Actuator makes possible

Valar was trained using Actuator, our patent-pending system for closed-loop model training. Actuator feeds measurements of model behavior back into the training process, adjusting training pressure when protected capabilities start to drift. The same control paradigm spans fine-tuning, alignment, distillation and compression, allowing a model to specialize while protecting the capabilities its application depends on.

Valar shows the value of that approach in a demanding scientific role: a 27B investigator scoring above named frontier models on complete research tasks. For product teams, specialization offers a way to build models around their own workflows and deployment constraints. We built Actuator to make that process repeatable, so teams can develop capabilities specific to their products while retaining control over how those models are trained and run.

Scientific reasoning researchers can trust

For AI to have a meaningful impact on science, researchers need to trust the rigor and quality of its work. As we discussed in our post on vibe science, that requires continued advances in scientific interpretation, hypothesis generation and long-horizon reasoning while reducing hallucinations. These capabilities still need to improve even in the strongest frontier models.

That commitment to scientific rigor is at the heart of Marvin. Valar brings more of that reasoning capability to a 27B model, making it practical on workstations and private GPU servers and expanding where scientific agents can run. Smaller models also allow more investigations to run in parallel, bringing this capability within reach of more research teams. We’re uploading the weights to Hugging Face so other researchers can build on Valar and help advance agentic science.

We’re also using Kev to convert Valar into a fast Jev-style decision model. Kev turns LLMs into models that return choices and probabilities directly, extending Valar’s scientific judgment to applications that need quick decisions without generating a full written response.

We’re building toward a future in which every research team can work with an AI scientist capable of sustained, rigorous investigation. Our contribution is to advance that scientific judgment and make it practical at more scales of deployment. More researchers should be able to pursue difficult questions and produce findings others can trust and build upon.

Learn more about Actuator · Explore Marvin