GitHub Copilot's HydraFusion Comes to VS Code and the App - Picking an Execution Pattern Instead of a Model
On September 30, 2026, GitHub widened the HydraFusion research preview, which combines several models, to VS Code and the Copilot app. Billing is at each used model's standard rate, and the auto model selection discount does not apply. Here is how it works and when to reach for it.
GitHub announced on September 30, 2026 that its HydraFusion research preview now runs in two more places: the Visual Studio Code editor and the GitHub Copilot app. Before this, Copilot CLI was the only way in1. Reaching more surfaces, the company says, topped the list of things early testers asked for.
HydraFusion sits in the model picker without being a model. The documentation is explicit about it: “You select HydraFusion in the model picker, but it isn’t a model. It’s a system that orchestrates models at runtime.”2 Relative to choosing a single model, both the behaviour you get and the invoice you receive work differently.
Auto decides which model; HydraFusion decides which procedure
Automatic model selection, or Auto, has been in Copilot for some time, and it resolves to one model per request. The changelog draws the line this way: Auto makes a per-request model decision, whereas HydraFusion is an experiment in whether Copilot can choose a workflow and drive several models inside one turn1. The documentation states the same contrast more compactly — Auto answers the question of which model should take a request, and HydraFusion answers the question of which procedure should take a task.
Three procedures are in place for now2.
| Execution pattern | What it does |
|---|---|
| Single | One model handles the whole task by itself |
| Cascade | An efficient model produces the first attempt, and a quality gate either accepts that answer or hands the task up to something stronger |
| Critique | One model writes a draft, a model from another family comments on it, and the original model gets exactly one revision pass |
The procedure is picked per prompt. A prompt handled as Single can be followed by one handled as Critique, according to the documentation. That decision is cheap: it barely adds time, and none of the eventual answer comes out of it. Critique is described as close in spirit to the rubber duck agent in Copilot CLI.
The picking itself is framed as an optimization problem. Signals about how each model performs at reasoning, at generating code, at debugging and at using tools feed into it, and the aim is the most efficient procedure that still clears a quality bar1. The pool mixes quick models with models that reason harder, and because it is re-cut whenever new models land or the company revises its evaluations, GitHub publishes no roster of what is in there. You also have no say in which models get used2.
The invoice is whatever the models added up to
Worth knowing before you switch. Every model HydraFusion calls is charged at that model’s own list rate, the documentation says. The orchestrator adds no line item of its own, but the discount that comes with automatic model selection is withheld. Since one task can pull in more than one model, the credits it burns can exceed what a single model would have burned2.
Caching is handled by keeping the main conversation pinned to one model as far as possible, so cached tokens keep working for you. Helper models — a reviewer, say — are passed just enough context to do their part.
Context is unusual as well. There is no context window attached to HydraFusion itself; a given step lives inside the limits of whatever model is executing it. What the picker displays is deliberately pessimistic — it reflects the tightest limit in the pool — and if your current conversation exceeds it, you are asked to compact before switching2.
GitHub’s own advice on the split: keep Auto for day-to-day work, where one model per request plus a paid-plan discount on model costs is the efficient answer. Bring in HydraFusion for bigger, clearly bounded coding jobs — untangling a hard bug, or a change that spans several files — when extra passes over the problem might justify the extra minutes and credits. Anything quick or routine, it says, belongs back on Auto. The review passes inside Cascade and Critique cost time; Single is close to an ordinary single-model request.
Edits from a thrown-away draft stay where they are
Among the documented limits, this is the one that lands directly on your pre-commit habits. When a draft gets discarded, whatever it already wrote into the workspace — edited files, for instance — is left in place rather than rolled back. Check your changes before committing, the documentation says2.
Steps are visible while the work happens, but drafts in flight may be rewritten or dropped, so the only thing rendered is the answer at the end. Combine those two facts and you can have edits from an abandoned draft sitting in your working tree with no indication of it on screen. Our comparison of how Claude Code, Cursor, and Google Antigravity are driven found that each of the three tools puts the place you inspect results somewhere different. Here, the same question — where you look, and how finely — lands inside a single turn.
Other documented limits: HydraFusion operates separately from subagents and does not launch them; and while the preview lasts, both the procedures and the models are subject to change. No service level agreement applies during the preview, and production workloads are explicitly not the intended use2.
An administrator may have to switch it on
You need VS Code 1.140 or later, or VS Code Insiders. Should the picker not show it, switch on the chat.copilot.hydraFusion.enabled setting1. In the Copilot app, take the newest build, look up HydraFusion in Settings, flip it on, and then choose it in the picker.
Where Copilot arrives through an organization or an enterprise, someone with administrative rights may first have to permit preview features at that level. The changelog names Copilot Pro, Pro+, Business, and Enterprise as the audience, and states that administrators have to enable preview features for the Business and Enterprise cases1. The September 4 research preview post read differently: it offered HydraFusion “to users on all GitHub Copilot plans,” reachable through the /experimental switch in Copilot CLI3. The two posts diverge here, and neither explains the gap.
The documentation names no plans whatsoever. What it does say is that only models your plan carries and your organization’s or enterprise’s model policies permit are eligible, and that the picker hides HydraFusion outright if none of them are2. The default policy for new features aimed at Copilot Business and Enterprise states plainly that features still in preview fall outside it. A separate setting governs preview access, so October 22 arriving will not by itself unlock HydraFusion.
On two benchmarks out of three it trails Opus 5 slightly
The September 4 post carries measurements from three agentic coding benchmarks. Claude Opus 5 and GPT-5.6 Sol served as the points of comparison, and the figures published are those of whichever HydraFusion configuration tuned up best3.
| Benchmark | Cost vs. Opus 5 | Quality vs. Opus 5 |
|---|---|---|
| TerminalBench 2.1 | 67% lower | +4.9 points |
| DeepSWE | 36% lower | -1.5 points |
| CheckpointBench | 65% lower | -0.1 points |
On GitHub’s own numbers, the cost figure fell somewhere between 36% and 67% in all three cases, while the quality figure came out ahead in one — TerminalBench 2.13. Quality here means verified task quality, the proportion of tasks judged to have been answered correctly; the cost is an estimated total that sweeps in drafting, critique, revision, escalation, retries and fallbacks alike. GitHub itself calls TerminalBench 2.1 relatively saturated and says that makes wider validation important, which is why DeepSWE and its tougher repository-scale problems sit in the same set of three. CheckpointBench, the internal one, is multi-turn and built out of genuine Copilot sessions.
The caveats attached to those numbers deserve equal weight. GitHub calls them controlled offline results, tied to the particular benchmark revisions, workflow settings, pool of models and pricing assumptions used, and notes that every model was held at the medium reasoning setting3. What happens when real developer workloads meet the same policies is precisely what the preview is supposed to reveal.
The baselines are also worth dating. These measurements went out on September 4, against Claude Opus 5 and GPT-5.6 Sol. Claude Opus 5.5 followed on September 22, GPT-6.1 Sol on September 29, and the pool HydraFusion reaches into is not a fixed thing either. Read the table as a snapshot of the models available at the time.
What to verify before you switch
The gesture resembles choosing one model, while the billing and the behaviour underneath do not. Three things you can settle yourself first.
Spending visibility is the first. GitHub’s instructions for finding out which models ran are to hover the footer of a finished response in VS Code, or the response itself in the Copilot app; in Copilot CLI, /usage reports the credits each model consumed during a session2. The second is whether your pre-commit routine can absorb the question of what an abandoned draft left behind in the workspace. The third, for organizational use, is simply whether preview features have been permitted for you.
GitHub frames HydraFusion as a move away from picking the best available model and toward assembling, task by task and on the fly, the best procedure for solving it — the company’s opening wager on that premise3. On September 4 it conceded that withholding drafts leaves developers waiting with little to look at, named that an honest trade-off, and said progress reporting was the next thing on its list. Three improvements appear in this changelog: what happens at each step is easier to read, updates arrive more often, and a long-running task no longer looks stalled1. They land on exactly that trade-off.
Sources
- HydraFusion in VS Code and the GitHub Copilot app - GitHub Changelog (September 30, 2026)
- Using HydraFusion - GitHub Docs (accessed October 1, 2026)
- Project HydraFusion: Frontier quality via multi-model orchestration - GitHub Blog (September 4, 2026; the research preview announcement and benchmark results)
Was this article helpful?
Thank you!
Received. Thank you!