Urgent.News

What's breaking now, across thousands of outlets.

AI

Microsoft Brings MAI Code 1.1 Flash Local Inference to GitHub Copilot

Microsoft plans to bring MAI Code 1.1 Flash to local coding workflows in GitHub Copilot, allowing supported Copilot experiences to run the coding model on a device instead of relying solely on cloud inference. The rollout is planned to begin with a limited audience by the end of October 2026, covering Copilot CLI, the Copilot app, and IDE integrations such as Visual Studio Code. The development…

Microsoft is set to introduce local inference capabilities for its AI coding assistant, MAI Code 1.1 Flash, within GitHub Copilot. This development allows supported Copilot experiences to execute the coding model directly on a user's device rather than relying solely on cloud-based inference. The rollout is scheduled to commence with a select group of users by the end of October 2026, encompassing Copilot CLI, the Copilot app, and various IDE integrations such as Visual Studio Code.

The introduction of local model execution offers developers an alternative pathway for AI-assisted coding, providing an option for on-device processing alongside cloud-hosted models in scenarios where local execution is more appropriate. The decision to implement local inference hinges on factors such as the availability of compatible hardware, the efficiency of routing between local and cloud models, and the final pricing structures.

Microsoft has not disclosed the pricing for local inference in its official announcement. According to the company's official documentation, MAI Code 1.1 Flash is designed as a local coding model with a 256,000-token reference context window. It contains 137 billion total parameters, including 6.8 billion active parameters, and employs quantization and speculative decoding techniques to minimize its footprint on devices.

Microsoft outlines two approaches to utilizing local inference within Copilot. The first is auto orchestration, where Copilot determines whether a request should be processed locally or in the cloud based on Microsoft's HydraFusion orchestration approach. This dynamic routing aims to coordinate models across edge and cloud environments. The second option involves explicit local-model selection, which allows users to choose MAI Code 1.1 Flash via the Windows ML provider or by connecting the Copilot to local endpoints.

The way local inference operates differs depending on the selected approach. In auto orchestration, Copilot dynamically routes requests between local and cloud inference, ensuring optimal use of the available resources for each request. In explicit local selection, users can manually opt for MAI Code 1.1 Flash by leveraging Windows ML or by establishing a connection to local endpoints, granting them direct control over where a coding request is processed.

This distinction is crucial for developers who prefer to have explicit authority over the processing environment for their coding tasks.

In addition to local model execution, Microsoft is integrating sandboxed tool use through its Microsoft Execution Containers (MXC). These containers are designed to isolate tool execution, which is particularly relevant for agentic coding workflows where a model may need to interact with external tools as part of completing a development task. Local inference complements this capability, enabling the execution of complex coding tasks that go beyond mere text generation.

The hardware requirements for running MAI Code 1.1 Flash locally are significant. Microsoft's reference measurements were conducted on a Surface Laptop Ultra equipped with an NVIDIA RTX Spark GPU. Under these conditions, the model consumes approximately 75.5GB of peak memory at a 256,000-token context length. The decoding throughput for this setup is reported at 923.5 tokens per second when operating on a 64,000-token context, and 769.8 tokens per second at a 128,000-token context.

These figures highlight the importance of hardware specifications in determining the feasibility and performance of local inference.

The rollout of local inference is contingent on hardware capabilities, and businesses should refrain from assuming that all developer machines will meet the necessary requirements or perform optimally at the reported context sizes and performance levels. Microsoft has not provided a general minimum specification or a comprehensive list of supported devices, emphasizing that the adoption of local inference will largely depend on individual hardware configurations.

For developers already utilizing Copilot CLI, the Copilot app, or supported IDE integrations, the introduction of local inference could seamlessly integrate into their existing Copilot workflow alongside other AI-assisted tools. The option for explicit local-model selection may prove beneficial for developers who wish to leverage a local endpoint, while the auto orchestration feature aims to streamline the decision-making process for AI-assisted coding by determining the most suitable environment for each request automatically.

From a business perspective, the immediate implications of this rollout extend beyond AI coding assistance, focusing on matching the appropriate AI capabilities to the specific needs of each development workflow. While local processing may be advantageous for teams equipped with compatible machines and seeking a local option within their current development tools, cloud inference remains integral to Microsoft's routing strategy.

Consequently, the announcement suggests a hybrid approach rather than an entirely local-only Copilot experience.

One unresolved aspect of the rollout is the pricing model for local inference. While a previous social media post suggested that local calls would incur no inference charges, Microsoft's official announcement remains silent on this matter. Pricing details, subscription treatment, usage limits, and the potential absence of additional costs for local workflows are not specified.

Companies evaluating the integration of AI coding assistance into their software delivery processes should exercise caution and await further guidance from Microsoft regarding the financial implications of local inference.

For organizations contemplating the implementation of AI coding assistance within their software development processes, the technical details of local inference carry significant weight. The decision to incorporate local and cloud AI capabilities must align with the practical requirements of development tasks, ensuring a cohesive and efficient workflow.

Microsoft Scalevise can assist companies in mapping practical development tasks, connecting AI tools to reliable workflows, and reducing repetitive operational work through AI workflow automation services. A thorough assessment can identify the most suitable integration points for local and cloud AI capabilities, allowing businesses to leverage these technologies without disrupting current workflows.

In summary, the introduction of local inference for MAI Code 1.1 Flash in GitHub Copilot represents a significant step forward in AI-assisted coding. By offering developers the choice between local and cloud processing, Microsoft aims to cater to a diverse range of development environments and use cases. As the rollout progresses, the practical implications for developers, businesses, and the broader software ecosystem will become increasingly evident, highlighting the importance of aligning AI capabilities with the unique needs of each development workflow.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Tokyo Seniors Learn How to Use AI

Tokyo hosted an employment support event for older workers on October 6, offering job consultations, interviews, hands-on work experience and training in artificial intelligence as the number of…

More from Wednesday 7 October →