Urgent.News

What's breaking now, across thousands of outlets.

AI

GitHub Copilot is going local — but Microsoft won’t say what gets sent to the cloud

GitHub Copilot will soon decide whether coding tasks run locally or get sent to cloud models, with automatic routing expected The post GitHub Copilot is going local — but Microsoft won’t say what gets sent to the cloud appeared first on The New Stack .

GitHub Copilot is going local — but Microsoft won’t say what gets sent to the cloud

Microsoft's GitHub Copilot is set to determine whether coding tasks should run locally or be sent to cloud models, with automated routing planned for October's end. This announcement, co-authored by GitHub product manager Patrick Nikoletich and Windows platform partner architect Stuart Schaefer, came alongside the availability of new sandboxing controls for GitHub.

However, the extent of data sent to the cloud varies depending on the tools used by Copilot. While shell commands and local MCP servers are subject to OS-level restrictions, built-in file tools rely on checks within the agent harness. Remote MCP servers remain outside the local process sandbox. Although Microsoft acknowledges that local inference does not make a session completely offline, they have not disclosed how much repository context Auto sends to cloud models, nor whether developers can view routing decisions or restrict inference to local models.

Copilot decides the location of inference based on task context and cache state, and developers can choose between auto-routing or using a local model directly in the Copilot CLI. The options include MAI Code 1.1 Flash through the Windows ML provider and OpenAI-compatible local endpoints. However, Microsoft has not specified how much conversation history or repository context Auto sends to the cloud when it routes a task, nor whether developers can see these decisions or limit inference to local models.

For teams with strict data-handling policies, the extent of data sent to the cloud remains unclear. Selecting a local model ensures inference stays on the device, but it does not prevent the agent from making external network requests through its tools. Developers requiring a fully local session would also need to secure what those tools can access.

MAI Code 1.1 Flash, a mixture-of-experts model with 137 billion total parameters, has been reduced to 53GB through mixed-precision quantization and speculative decoding. This model is targeted at NVIDIA RTX Spark Windows PCs, such as the Surface Laptop Ultra, which offers up to 128GB of unified memory. However, peak memory use of 75.5GB at a 256K-token context makes it challenging for most developer laptops with 16GB or 32GB of RAM.

Benchmarks show that the quantized model scored 70.8% on SWE-Bench Verified, compared to 72.6% for the full-precision version, and outperformed the original on Terminal-Bench 2.1. While these improvements suggest a reduction in model size without significant coding performance loss, they fall short of proving that quantization made the model better.

Microsoft employs Microsoft's open-source Execution Containers (MXC) library to enforce sandbox policies, with BaseContainer, Seatbelt, and bubblewrap used on Windows, macOS, and Linux, respectively. When sandboxing is enabled, Copilot applies OS-enforced restrictions to shell commands and local MCP and language servers. However, GitHub's claim of offline workflows may not be verifiable, as the prompt used in the demo may have retrieved metadata over the network instead of using local repositories and tests.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at thenewstack.io →

More in AI

More from Thursday 8 October →