Saltar al contenido
ES EN

GitHub Copilot will split coding work between local models and the cloud

Microsoft says Copilot will soon choose between on-device and cloud inference, starting with a quantized MAI Code 1.1 Flash on RTX Spark Windows PCs.

In 30 seconds Microsoft and GitHub say Copilot will soon decide whether a task runs on-device or in the cloud. A quantized local build of MAI Code 1.1 Flash targets NVIDIA RTX Spark PCs such as Surface Laptop Ultra, with sandboxed execution for agent commands.

What happened

In a post on Microsoft's Command Line site, GitHub and Windows platform leads describe two additions. First, Copilot will route work between local and cloud models automatically, which they call the next step of Project HydraFusion. Second, Microsoft Execution Containers (MXC) give agents controlled access to files, networks, system capabilities and credentials when they run commands on Windows.

GitHub Copilot will split coding work between local models and the cloud
Image: Unsplash — A glossy glass cube with the Microsoft logo on a dark surface

The local model is a version of MAI Code 1.1 Flash, a mixture-of-experts design with 137 billion total and 6.8 billion active parameters. Microsoft applied quantization and speculative decoding to shrink it to 53 GB, an 80% reduction against the cloud variant. On Surface Laptop Ultra, peak memory reaches 75.5 GB at 256k context. Prompt processing hits 923.5 tokens per second at 64k and 769.8 at 128k.

Microsoft reports 70.80% on SWE-Bench Verified and 66.29% on Terminal-Bench 2.1 for the on-device build, against 72.6% and 62.9% for the unquantized cloud model. The comparison model, a GPT OSS 120B GGUF, scored 32.0% and 23.6%. Users can let Copilot's Auto mode pick, or select a local model explicitly in the CLI, the Copilot app or VS Code.

Model selection, inference, and tool execution have different boundaries; local inference does not make the session offline.

Microsoft and GitHub, Command Line post

Why it matters

Running a capable coding model on a laptop changes the cost and latency trade-off, and keeps some work off remote servers. The memory discussion is the practical lesson: weights are only part of the budget, since the operating system, the runtime and a growing key-value cache all compete for the same unified pool.

The wording about boundaries matters too. Choosing a local model does not mean nothing leaves the machine, so teams with data-handling rules should check which parts of a session actually stay on the device. We would test the sandbox permissions before enabling agents on repositories that contain credentials.

What we don’t know yet

The post says the feature arrives by the end of the month, but gives no pricing, no list of supported hardware beyond RTX Spark machines, and no independent verification of the benchmarks. The figures come from Microsoft's own tests on a single device configuration, and Microsoft notes that real-world results may vary.

Further reading: Source.

Original source: Microsoft

Article generated with AI.larebelion

Falcon

· Signals analyst · Riyadh

“Local models are the headline, but the sandbox is where agent risk is actually decided.”

Comentarios

Publicar un comentario