Microsoft and Nvidia Move Windows Toward a Local Runtime for AI Agents


NVIDIA Blog
other
NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents
Windows Developer Blog
other
Microsoft Execution Containers: Policy-driven containment for AI agents
Windows Experience Blog
other
Building Windows for hybrid intelligence
MXC is GA
Microsoft Execution Containers are now generally available as a policy-driven containment layer for AI agents on Windows.
Local 120B models
Microsoft says Surface Laptop Ultra can run AI models exceeding 120 billion parameters locally with up to 128 GB of unified memory.
October shipping
RTX Spark laptops are available for preorder now and are scheduled to become available beginning October 16.
Microsoft and Nvidia are recasting Windows PCs as local execution environments for AI agents, pairing new RTX Spark hardware with Microsoft Execution Containers, an operating-system-level containment layer that Microsoft says is now generally available for agent workloads on Windows 11.12
The companies announced the shift on October 7 at a Windows AI and Surface event in San Francisco. Nvidia said RTX Spark laptop preorders opened that day, with laptops available October 16 and compact desktops planned for November.1 Microsoft framed Windows as a “hybrid intelligence” platform where agents can run locally when appropriate, call cloud models when needed, and operate under enterprise controls for containment, identity and manageability.3
For enterprise developers and IT security teams, the practical change is that Microsoft is not just adding AI features to Windows. It is building infrastructure intended to let agents keep running on PCs, access local files and tools, and be governed by policies separate from the agents themselves.23
RTX Spark systems are designed to bring larger local AI workloads to Windows laptops and small desktops. Nvidia said the platform combines an Nvidia Blackwell RTX GPU with up to 6,144 cores and an up to 20-core Nvidia Grace CPU, connected at 600 GB/s.1 The company said RTX Spark can deliver up to one petaflop of FP4 AI performance and up to 128 GB of unified memory, allowing some large models to run without sending data to the cloud.1
Microsoft’s Surface Laptop Ultra uses the same RTX Spark foundation. Microsoft said the device supports up to 128 GB of unified memory, up to one petaflop of AI performance and local models exceeding 120 billion parameters.4 Surface Laptop Ultra starts at $2,599, with availability beginning October 16. The Surface RTX Spark Dev Box is listed at $5,999 and is scheduled to begin shipping to U.S. customers in November.4
The broader RTX Spark PC lineup includes systems from ASUS, Dell, HP, Lenovo, MSI and Microsoft Surface. Microsoft said those laptops are available for preorder and begin shipping October 16.3 Tom’s Hardware also reported that Surface Laptop Ultra, Surface RTX Spark Dev Box and partner RTX Spark systems were available for preorder, with laptops from several manufacturers shipping October 16.7
Microsoft is also previewing a higher-end tier: DGX Station for Windows. Nvidia said the deskside system uses the GB300 Grace Blackwell Ultra Desktop Superchip, with 748 GB of coherent memory and up to 20 petaFLOPS of FP4 AI compute, enough to run models at up to trillion-parameter scale locally.1
That product is not the mainstream laptop story, but it shows the direction of travel: Windows desktops are being positioned as local AI infrastructure for developers, researchers and enterprises.
The central Windows change is Microsoft Execution Containers, or MXC. Microsoft describes MXC as a policy-driven execution layer for untrusted code or dynamically generated workloads. In agent scenarios, developers can use it to contain model-generated output, plug-ins, tools, the agent harness or the full agent workload.2
MXC is now generally available and is meant to define what an agent can access at runtime, including files, network destinations, tools and user-interface resources.2 The policy is kept outside the agent workload’s control, so the model or generated code cannot expand its own permissions.2
Microsoft presents this as a necessary boundary for agents because they can operate differently from traditional applications. They may run continuously, use tools, write code, read and modify files, and act across systems without a user watching each step.3 Microsoft’s argument is that an agent should not inherit the full authority of the signed-in user simply because it is acting on that user’s behalf.2
MXC supports multiple containment levels. Process containers provide lighter-weight sandboxing across Windows 11, macOS and Linux. Windows-only session containers run an agent in a separate operating-system-isolated session, with its own local agent identity and isolated desktop, clipboard, user interface and input boundaries. Windows Subsystem for Linux containers support Linux-first toolchains on Windows, while experimental microVMs provide stronger hardware-backed isolation for higher-risk workloads.2
For developers, the immediate integration model is a JSON configuration schema and SDK that declares workload requirements. Microsoft says MXC then maps those requirements to the appropriate containment backend.2
For IT teams, the important distinction is that organizational policy can further constrain what developers request. Microsoft said Intune policy support for MXC process containers on Windows 11 is coming, allowing administrators to control how container creation requests are evaluated and what resource boundaries are enforced.2
Microsoft’s Windows strategy relies on more than faster local inference. The company is designing for agents that can keep working in the background and combine local context, local models and cloud models.
In its Windows strategy post, Microsoft said Copilot will gain local context, local actions and local models on Copilot+ PCs, with user permission. The company said its Autopilot experience is intended to become a persistent, proactive and personal agent that can continue working while the user focuses elsewhere.3
Those Copilot hybrid intelligence features are expected to begin rolling out on Copilot+ PCs in the coming months, meaning not every capability announced with the strategy is available immediately.3
The same local-plus-cloud approach is expected to affect developer workflows. Microsoft said GitHub HydraFusion will be extended to Windows so GitHub Copilot can route work to local models when appropriate and to cloud models when needed, with an experimental preview later in October.3 Axios reported that Microsoft’s pitch is partly economic: developers want more intelligence than cloud budgets can comfortably support, so the company is emphasizing a hybrid model that uses local compute where it can be effective.5
Microsoft’s security model has three stated pillars: containment, identity and manageability.23 MXC provides the containment layer available now. It can govern file-system access, network connectivity, process settings and whether a workload can interact with the desktop.2
The identity layer is partly still ahead. Microsoft said Windows will soon enable Microsoft Entra to distinguish agent activity from user activity in Microsoft Agent 365.2 The goal is attribution: security teams should be able to determine which agent took an action and apply controls to that agent without necessarily blocking the human employee’s access.2
Manageability is also being tied to Agent 365 and Intune. Microsoft said Agent 365 controls will extend to local on-device agents, letting IT manage MXC containers, apply policies and monitor agent activity.2 HotHardware reported that Microsoft positioned Agent 365 as an enterprise control plane for discovering, monitoring and governing agents across Windows and cloud environments, with integrations involving Entra, Defender and Purview.6
MXC also includes operational modes intended to help organizations build least-privilege policies. Enforcement mode blocks ungranted access. Learning mode blocks and records denied activity. Permissive mode allows activity while recording what policy would have denied, helping developers and administrators observe what resources an agent tries to use before tightening controls.2
The announcement does not eliminate the need for cloud AI. Microsoft’s own framing is hybrid: local models for responsiveness, cost control and data locality; cloud models for tasks that need broader or more powerful capability.3
But the center of gravity changes if Windows PCs can host local models with more than 120 billion parameters, keep agents running in isolated sessions, and route work between local and cloud models under policy control.34
For developers, the near-term question is whether MXC becomes a standard packaging and execution target for Windows agents, especially coding agents and tool-using assistants. Microsoft said GitHub Copilot, OpenAI Codex, Replit, LM Studio and other agents or frameworks already support MXC, with more support planned.2
For security teams, the question is whether Microsoft’s controls are strong and observable enough for real enterprise deployment. The company is drawing a line between the user and the agent, and between the agent’s objective and its authority. If implemented consistently, that separation could reduce the risk of local agents overreaching into files, networks or credentials they do not need.
That is the bigger shift behind the device launch. RTX Spark gives Windows PCs more local model capacity. MXC gives agents a place to run under operating-system control. Together, Microsoft and Nvidia are trying to make Windows a managed local runtime for AI agents, not just a client that sends prompts to the cloud.

Tab has emerged from stealth with a personal AI assistant that users can message through iMessage or WhatsApp to shop, book, call and manage reminders. Its architecture — a phone number, computer, wallet and approval gates — shows how consumer agents are shifting from chat interfaces to delegated operators.

Google Research’s new Population Dynamics Foundation Model work suggests reusable geospatial embeddings can improve or match conventional inputs across several public-health tasks. The engineering case is strongest when the embeddings supplement existing epidemiological models, not when they are treated as a replacement for surveillance data, local validation or deployment governance.

Cisco disclosed critical NX-OS vulnerabilities affecting Nexus 3000 and 9000 switches, including NGOAM flaws that can allow unauthenticated remote code execution with root privileges or denial of service. With no formal workarounds and only limited temporary mitigations, network teams may need to prioritize staged fabric upgrades over routine endpoint-style patching.

Anthropic’s Claude Haiku 5.5 makes its small-model tier cheaper and faster for high-volume tasks, underscoring a broader shift in AI product architecture. For builders, the next margin battle is less about flagship reasoning models and more about routing repetitive agent work to low-cost models that can run at scale.
Microsoft Execution Containers
A Microsoft containment layer that lets developers and IT define what files, networks, tools and interfaces an agent workload can access.
Hybrid intelligence
Microsoft’s term for routing AI work between local models on a PC and cloud models depending on capability, cost, latency and policy needs.
Unified memory
A memory architecture shared between the CPU and GPU, giving local AI workloads access to a larger common memory pool.
Session container
A Windows-only MXC containment option that runs an agent in a separate isolated session with its own desktop, clipboard and input boundaries.
Comments