Why Ordered Reasoning Budgets Matter for Coding Agents
Qwen3.8-27B-pi explores a practical property for agentic models: higher reasoning settings should buy at least as much work and capability as lower ones. The project offers useful lessons about evaluating cost, correctness, and tool-using behavior together.

Reasoning effort should be evaluated as an end-to-end ordering contract: higher settings must justify their cost with at least as much useful work and capability.
A reasoning-effort control sounds simple: choose a low setting for fast work and a higher setting when a task deserves more deliberation. For a coding agent, however, that promise can break in a subtle way. A nominally lower setting may consume more reasoning than the setting above it while solving fewer tasks. The label then stops being a useful budget control.
Qwen3.8-27B-pi is an experiment in correcting that behavior for the Pi terminal coding harness. Its author reports a two-stage process: supervised fine-tuning on successful, complete agent sessions, followed by reinforcement learning that rewards correct lower-effort attempts for staying within the reasoning used by successful higher-effort attempts on the same task. The reported evaluations show non-decreasing reasoning use and pass rate from low through medium to xhigh across Terminal-Bench 2.1, GPQA Diamond, and SciCode. Those results are project-reported measurements, not an independent replication.
The broader contribution is a useful way to think about controllable inference. Ordering, rather than a particular token ceiling, is the central product property.
Effort is an interface contract
An effort selector is meaningful only if users can predict what moving the control will do. The precise token count does not need to be fixed; repository size, test latency, and task ambiguity make that unrealistic. But the direction should be dependable. A higher setting should not routinely spend less reasoning while also performing worse.
This distinction matters because an agent's cost is wider than its hidden reasoning. It reads files, invokes shell commands, receives observations, revises edits, and may repeat the loop. A short internal thought can still trigger an expensive tool sequence. Conversely, a longer plan may prevent several failed edits. Evaluators therefore need to connect effort levels with outcomes across the whole run: correctness, reasoning tokens, total output, number of turns, tool calls, and elapsed time.
Ordering is best treated as a statistical contract over a representative task set, not a guarantee for every prompt. An easy task may terminate immediately at every setting. A noisy test can also make one run longer than another. What matters is whether the aggregate control remains monotonic enough to guide real decisions.
Why successful trajectories help but do not settle the problem
Training on complete successful sessions teaches more than answer formatting. It exposes the sequence that makes an agent useful: inspect the repository, form a hypothesis, change a file, run a check, and react to feedback. That is especially valuable when the deployment harness has a compact tool vocabulary and a distinctive interaction loop.
Yet imitation preserves quirks as readily as strengths. If collected low-effort sessions happen to deliberate more than medium-effort sessions, supervised learning can reproduce that inversion. Token-level validation loss may still improve because it measures how well the model predicts the training distribution, not whether the deployed agent uses a budget sensibly or completes more tasks.
That leads to a practical checkpoint-selection rule: evaluate checkpoints inside the intended harness. A model that predicts demonstrations slightly better is not automatically the model that edits repositories more reliably. Behavioral panels should include multiple effort settings and should record failure shape, such as repeated commands, excessive turns, or premature stopping, rather than collapsing everything into one average loss.
Correctness must remain the first gate
A naive efficiency reward invites an obvious shortcut: stop early. A concise failure is cheap but useless. The approach described for Pi avoids that incentive by conditioning length pressure on success. Incorrect attempts receive no benefit merely for being short. Successful low and medium attempts are compared with successful higher-effort attempts for the same task, while the highest setting is left without length pressure.
This structure has two attractive properties. First, it uses the task itself as the comparison unit. A one-line configuration fix and a multi-file refactor do not share a sensible universal budget. Second, it preserves a capability ceiling: the top effort level can explore without being pulled toward an arbitrary brevity target.
There are boundaries. A verifier can be incomplete, and optimizing against it may reward changes that pass tests without satisfying the real request. Higher-effort attempts can also fail, leaving no trustworthy budget reference. The comparison data must therefore be complete, structurally valid, and scored by credible task checks. Reward design cannot compensate for weak evaluation contracts.
A practical evaluation matrix
Teams assessing an effort-aware coding model can start with three questions. Does accuracy stay flat or rise as effort increases? Does reasoning spend stay flat or rise? And does total agent cost follow a useful pattern once tool use and repeated context are included?
Run the same task sample at every supported setting with enough repeats to expose variance. Separate successful and failed trajectories: failures often run until a limit and can dominate averages, while successful-only numbers reveal whether the model needed extra work to reach the same outcome. Inspect per-task inversions as well as aggregate means; a clean average can hide a class of repositories where the control behaves backward.
Finally, compare adjacent settings for decision value. If medium is both cheaper and more accurate than low, the low option is not serving users even if the overall benchmark score looks strong. If xhigh costs substantially more for only rare gains, document the task types that justify it. The goal is not uniformly minimal reasoning. It is a control whose extra cost buys a credible chance of extra capability.
What the experiment changes
Qwen3.8-27B-pi frames inference effort as something that can be trained and audited, rather than a decorative runtime parameter. It also illustrates why agent evaluation must join model behavior to harness behavior: checkpoints, tool loops, task verifiers, and effort settings form one system.
The reported benchmark gains and deployment formats make the model available for further testing, but the durable lesson is methodological. When exposing low, medium, and high reasoning modes, define the ordering users should expect, select checkpoints with end-to-end tasks, and refuse efficiency improvements that trade away correctness. A budget selector earns trust when its direction is predictable—even when the exact cost of the next coding task is not.
Source: Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding ↗. How we write


