Skip to main content

The Computer Comes Home

RTX Spark, local agents, and the return of compute we can own

A laptop on a quiet desk becomes a miniature local compute environment while a distant data centre recedes into the background.
  • NVIDIA RTX Spark matters less as a fast new laptop chip than as a sign that capable AI inference, persistent agents, large unified memory and security policy are becoming properties of the personal computer itself. Apple, AMD and DGX Spark set parts of the pattern; RTX Spark brings CUDA, Windows on Arm, broad OEM support and an agent runtime into one consumer-PC platform.
  • The headline figure is one petaflop. The strategically important figure may be 128 GB. Large unified memory changes which models, contexts and agent workloads can remain local instead of being forced through a cloud API.
  • Local AI changes the economics of intelligence. A cloud API rents capacity one request at a time. A local machine converts part of that variable spend into owned capacity, energy, depreciation and operational responsibility.
  • Locality can improve privacy, latency, continuity and control. It does not automatically create sovereignty. CUDA, Windows, Arm, MediaTek, TSMC, firmware, model distribution, drivers and global memory supply remain dependencies.
  • Agents make the endpoint more important because useful agents need proximity to files, credentials, applications, identity and long-running context. That is why Microsoft and NVIDIA are building agent identities, execution containers, policy runtimes and audit layers into the operating system around the chip.
  • The likely future is not local or cloud. It is a hierarchy: routine and private work stays local; parallel work spreads across nearby owned machines; specialised or frontier tasks escalate to larger private or public infrastructure when required.
Is RTX Spark already on sale?

Not at the evidence cutoff. NVIDIA announced RTX Spark on 31 May 2026 at GTC Taipei and confirmed at IFA 2026 that the first Windows PCs arrive in October, from ASUS, Dell, HP, Lenovo, Microsoft and MSI, with others following. Retail pricing and independent benchmarks of production systems were still missing on 30 September 2026.

Does one petaflop make a laptop a supercomputer?

No. NVIDIA advertises up to one petaflop of FP4 AI performance. That is specialised low-precision tensor throughput, useful for quantised inference. It is not one petaflop of arbitrary computing and not a universal measure of application performance.

Why does 128 GB matter more than the petaflop?

An AI workload runs well only if the model and its working state fit where the accelerator can reach them. With up to 128 GB of coherent unified memory, larger models, longer contexts, retrieval indexes and several agents can stay on the machine instead of crossing the network for every interaction.

Does local AI mean no cloud and no subscriptions?

Only for a fully local stack. An open-weight model on your own hardware avoids paying a provider for every token, but a local agent can still use subscription software, licensed models, cloud search, hosted email, SaaS APIs and frontier-model escalation. The likely pattern is a hierarchy, not a replacement.

Is a local AI computer sovereign?

Not by itself. Locality can improve privacy, latency, continuity and cost control. The machine still depends on CUDA, Windows, Arm, MediaTek, TSMC, firmware, drivers, model distribution and global memory supply. The better question is which capabilities you could still operate, repair, replace and understand if one supplier changed tomorrow.

What does this mean for Swiss organisations?

Local AI lets a firm, hospital, archive or public body do useful work without sending every document into a shared remote inference service, and it lowers the threshold for owning AI capacity. But buying local AI computers is procurement, not sovereignty. The practical question is which critical AI capabilities can be made locally operable, portable and recoverable.

Your trail
Reading tools
Publication record
Author
Thierry Gilgen
Edition
1
Published
2026-09-30
AI assistance
Not recorded
Editorial review
Not recorded
Structured source record
No structured sources recorded. Inline citations are separate.
Edition change
Published content updated
SHA-256
6fb08687651cabd2595b43e602c92bbb317bf0e43f6e525ebccba1c1ba07927f

A checksum identifies the recorded text and metadata. It does not certify the truth of a claim or the contents of external links.

Thierry Gilgen. The Computer Comes Home. Edition 1. 2026-09-30. https://www.thierry-gilgen-ict.ch/field-notes/the-computer-comes-home

Edition history

Thierry Gilgen. The Computer Comes Home. Edition 1. 2026-09-30. https://www.thierry-gilgen-ict.ch/field-notes/the-computer-comes-home

A LinkedIn post stopped me this week.

It described NVIDIA's RTX Spark as the moment the traditional PC architecture ends: CPU, GPU and memory collapsing into one system, a petaflop of local AI performance, large models running without a cloud server, and personal agents living continuously on the machine.

The wording was dramatic.

Some of it was too dramatic.

But the instinct underneath it was right.

Something important is happening to the personal computer.

For most of the last fifteen years, the direction of travel in computing seemed obvious. We moved computation away from the desk and into somebody else's data centre. Servers became cloud regions. Software became services. Storage became accounts. The most capable artificial intelligence became an API call.

The machine in front of us remained important, but increasingly as a window onto infrastructure somewhere else.

AI accelerated that shift. The smartest thing my laptop could do was often not happening on my laptop at all.

RTX Spark points in the opposite direction.

Not back to the 1990s. Not away from the cloud. Not toward some fantasy of technological isolation.

Toward a different division of labour.

Training can remain inside enormous AI factories. Frontier reasoning can still be called when needed. But memory, context, model inference, tools and persistent agents can increasingly sit next to the person whose work they serve.

That changes performance.

It also changes ownership.

And ownership changes the sovereignty question.

Four eras of computing, left to right: a central mainframe feeding thin terminals, a standalone personal computer, many small endpoints all pointing to one remote block, and a personal computer with two nearby owned nodes inside a red authority boundary that escalates first to private compute and only then to a distant cloud block. Height stands for distance from the user.
The direction is not back from cloud to PC. It is toward a hierarchy in which more intelligence can remain close to the user while larger infrastructure still exists above it. Reading the figure, left to right: mainframe (central compute, thin terminal); personal computer (local applications and local state); cloud (remote services, endpoint as client); agentic hierarchy (local device, nearby owned nodes, escalation to private compute or cloud). Red marks the authority boundary around the local devices in the final era. Conceptual diagram based on the essay’s argument; no external data.

First, correct the launch story

The social-media version compresses several separate developments into one dramatic moment.

RTX Spark was not unveiled yesterday. NVIDIA announced the platform on 31 May 2026 at GTC Taipei around COMPUTEX. What became current again in September is the move from announcement to product: NVIDIA used IFA 2026 to confirm that RTX Spark Windows PCs are arriving in October and to add more of the surrounding local-agent stack, including NVIDIA PAIR for routing inference across machines on a local network.12

The architecture itself also deserves precision.

RTX Spark combines a Blackwell RTX GPU with a 20-core Arm-based Grace CPU and supports up to 128 GB of coherent unified LPDDR5X memory. NVIDIA's developer documentation describes a 256-bit memory interface with roughly 300 GB/s of bandwidth on the high-end system.3

That does not mean CPU, GPU and 128 GB of RAM have literally become one monolithic piece of silicon. It means the CPU and GPU operate over a shared coherent memory architecture instead of the traditional arrangement in which a CPU owns system RAM and a separate GPU owns a physically separate VRAM pool.

That distinction matters.

So does the petaflop.

NVIDIA advertises up to one petaflop of FP4 AI performance. This is specialised low-precision tensor performance, not one petaflop of arbitrary computing. It is an important capability for quantised inference, but the marketing number should not be confused with a universal measure of application performance.1

The claim that the system can run 120-billion-parameter models with very long context is also NVIDIA's platform claim, not yet an independent benchmark across retail systems. It is technically plausible that 120B-class quantised models fit into this memory class: OpenAI's gpt-oss-120b checkpoint, for example, is about 60.8 GiB and was designed to fit into an 80 GB accelerator.4 But model size, context length, KV-cache design, quantisation, batch size and inference engine all change what “runs locally” means in practice.

The same caution applies to the gaming claim. NVIDIA says RTX Spark can exceed 100 frames per second at 1440p in AAA titles, while recent PC coverage notes that the headline can rely heavily on DLSS and frame generation rather than raw raster performance.5

And “no monthly subscription” is not a property of the chip.

If I run an open-weight model, local tools and local storage, then yes: inference can happen without paying a model provider for every token. But a local agent can still use subscription software, licensed models, cloud search, hosted email, SaaS APIs, remote storage or frontier-model escalation.

The correct statement is more interesting than the viral one:

RTX Spark makes a much larger share of useful machine intelligence economically and technically possible on hardware the user owns.

That is enough.

The number that matters is 128 GB

The launch is marketed around a petaflop.

I keep coming back to the memory.

AI has a brutal property: if the model and the working state do not fit where the accelerator can reach them efficiently, theoretical compute stops being the interesting constraint.

For years, consumer GPU discussions were dominated by how much VRAM a card had. Eight gigabytes. Twelve. Sixteen. Twenty-four. Plenty for gaming, often painfully little for large local models.

Unified memory changes that geometry.

RTX Spark's high-end configuration can expose up to 128 GB to a coherent CPU/GPU architecture.3 AMD's Ryzen AI Max+ systems already demonstrated why that matters: AMD has shipped systems with 128 GB of memory and explicitly positioned them for 128B-class and, in newer “agent computer” messaging, even larger local models under the right quantisation and software conditions.6 Apple's current M5 Ultra takes the same general architectural idea much further at the workstation end, with up to 512 GB of unified memory and 1.2 TB/s of memory bandwidth.7

So the unified-memory idea is not new.

What is changing is how normal it is becoming to design a personal computer around AI memory requirements instead of treating local AI as a side workload for unused GPU capacity.

This has consequences beyond loading a large language model.

A useful agent needs more than weights. It may need a long context, embeddings, visual encoders, speech models, retrieval indexes, tool state, multiple concurrent model instances, speculative decoding, caches and perhaps several subagents operating at once.

Memory becomes a strategic resource because memory determines how much of that working set can remain close to the accelerator.

A tall outlined memory field representing a 128 GB unified capacity envelope, divided into unequal blocks for model weights, context state, vision and speech models, retrieval, concurrent agents and a system reserve, with a red bracket along its full height. To the left, four small columns show 8, 12, 16 and 24 GB consumer graphics memory at the same scale.
Peak compute tells us how quickly the chip can calculate. Memory capacity determines how much of the working intelligence can stay beside it. Reading the figure: the tall envelope is the 128 GB unified capacity envelope. Blocks, top to bottom: model weights; KV and long-context state; vision and speech models; embeddings and retrieval; agent concurrency and working memory; operating-system and application reserve. The red bracket marks what can remain local; the small double arrows mark movable boundaries. The four small columns show consumer VRAM of 8, 12, 16 and 24 GB at the same scale. Allocation is illustrative and depends on model and runtime; block sizes are not percentages. Sources: up to 128 GB of coherent unified memory, 256-bit interface, about 300 GB/s from NVIDIA, Windows on Arm Porting Guide — System Overview (accessed 30 September 2026), and NVIDIA’s launch announcement of 31 May 2026. VRAM tiers as discussed in the essay.

The model that fits can be more useful than the nominally smarter model that must cross the network for every interaction.

That is why 128 GB may matter more than the launch slide with one petaflop written on it.

The petaflop tells us what the chip can calculate.

The memory tells us how much intelligence can actually live there.

The PC is becoming an agent host

A conventional application waits.

I launch it. I click something. It processes a request. I close it.

A useful agent has a different operating shape.

It watches. It remembers. It receives events. It works while I am doing something else. It may split a task into subproblems. It may call tools, search files, generate code, open documents, invoke models, wait for a reply and continue later.

That makes the physical location of the agent surprisingly important.

An agent that lives beside my files, applications, microphone, local databases and devices can observe and act with lower latency and without continually exporting raw context. It can continue to perform bounded work when the public internet is unavailable. It can keep sensitive intermediate state off a third-party inference service. It can also become far more dangerous if it is given the same authority as the human account running it.

The industry has noticed the second half of that sentence.

Microsoft is turning agent execution into an operating-system concern. Its Microsoft Execution Containers (MXC) are designed to let developers declare what an agent may access and have Windows enforce those limits at runtime. Microsoft is also adding distinct agent identities and policy controls so that autonomous software does not need to inherit the full authority of the human user.89

NVIDIA's OpenShell works at a similar boundary. It places agents in sandboxes, denies access by default, controls filesystem and network reach, brokers credentials and routes inference according to policy. The important design principle is that the restriction is enforced outside the agent's own reasoning process. The model is not asked politely to obey the rule; the environment prevents the action.10

A left-to-right control flow. A person holds a red revocation switch that connects to an agent identity badge, which leads into a red policy sandbox containing the agent process and the model. A broker in the sandbox wall fans out to files, network and tools before a final action. An audit strip runs beneath the whole flow and records every stage.
The rule should not live in the model’s good intentions. It should live in the environment that can deny the action. Reading the figure, left to right: human authority, agent identity, policy sandbox (containing the model and the agent process), broker, then files, network and tools, then actions; underneath, the audit log. Red marks the policy boundary and the revocation switch controlled by the human. Sources: sandboxes, default deny, policy enforced outside the agent process, credential brokering and audit from NVIDIA, OpenShell (accessed 30 September 2026); Microsoft Execution Containers and agent identity from Microsoft, Build 2026 (2 June 2026) and Windows Agentic (accessed 30 September 2026); MXC in Windows preview builds from Microsoft Support, Windows 11 preview update KB5124006 (22 September 2026).

This is a much bigger change than an “AI button”.

The operating system historically knew about users, processes, files and applications.

It is beginning to know about agents as operational principals.

That means identity, permissions, isolation, audit and revocation have to be redesigned around software that can pursue objectives over time.

The AI PC is therefore not merely a PC with a fast neural accelerator.

It is a machine that is being redesigned to host delegated actors.

From personal computer to personal compute fabric

The most interesting NVIDIA announcement in September may not be the laptop at all.

It may be PAIR.

NVIDIA Personal AI Router connects compatible RTX PCs, DGX Spark systems and supported Macs on the same local network and exposes them as a place to send local inference requests. PAIR does not combine their memory into one giant GPU and it does not split one model across the house. It routes independent requests to whichever eligible node is ready.11

That distinction is exactly why it is relevant to agents.

A single chat is often sequential. A multi-agent workflow is naturally parallel.

One worker can inspect a repository while another analyses documents, another tests a hypothesis and another verifies sources. If all of them queue behind the same local inference server, the agent system becomes slower as it becomes more ambitious.

If those requests can spread across idle machines, the architecture changes.

NVIDIA demonstrated a five-subagent workload that completed in 8 minutes 48 seconds across three paired machines compared with 18 minutes on a single RTX Spark laptop. That is a vendor demonstration, not a universal performance result, but it illustrates the intended topology.11

This is a small conceptual step from a home network and a large conceptual step from the personal computer.

The personal computer used to be one box.

The cloud then made the box almost irrelevant to where computation happened.

The emerging local-AI model can make the room itself a small compute environment: laptop, workstation, mini PC, NAS, perhaps a dedicated AI node, each contributing capacity under one local policy boundary.

IDC saw the same pattern across IFA 2026, where vendors demonstrated AI NAS systems and clustered mini PCs aimed at agents, RAG and local model workloads. Their conclusion was not that every home becomes a data centre. It was that local AI infrastructure is becoming a product category of its own.12

That is a meaningful shift.

The endpoint is no longer merely a client of infrastructure.

It can become infrastructure.

Renting intelligence versus owning capacity

Cloud AI created a wonderful economic abstraction.

I do not need to own the accelerator. I do not need to cool it, patch it, replace it or keep it busy. I send a request and pay for the service.

That model is extraordinarily efficient when demand is uncertain, workloads are bursty, the best model changes every month, or the required capability is too large to own economically.

But the abstraction has a price.

Every useful token remains someone else's production capacity.

Every persistent agent becomes a recurring consumer of somebody else's inference infrastructure.

Every background check, classification, planning loop, summary, embedding, tool decision and retry can become a metered event.

The local machine changes the accounting.

Once I have bought the hardware, another local inference request is not free. It consumes electricity, hardware life, storage, engineering time and opportunity. The system depreciates. Models may outgrow it. Memory can become the bottleneck. A broken driver can waste a day.

But the marginal economics are different.

A purchased machine gives me a block of capacity I can allocate continuously without asking whether another ten thousand internal model calls justify another API bill.

A conceptual cost chart with no values. A rented line rises steadily from zero, and an owned line starts with a tall initial step and then rises more slowly. Both carry broad uncertainty bands, and a red dashed ellipse marks the region where they cross, without a number.
Local inference is not free. It changes the cost structure from pure metering toward capital, energy, depreciation and operational responsibility. Reading the figure: the horizontal axis is cumulative machine work or inference volume, the vertical axis cumulative cost. The straight line from the origin is rented inference (cloud or API); the step followed by a shallower line is owned local capacity (a capital step, then energy and operations). The tinted band (rented) and the hatched band (owned) show uncertainty. The red dashed ellipse marks a conceptual crossover, deliberately not numbered: no break-even value, no currency. Conceptual chart based on the essay section on renting intelligence versus owning capacity; not based on measured cost data. RTX Spark retail pricing was not available at the evidence cutoff (Ars Technica, 1 June 2026).

That matters most for agentic systems because agents manufacture requests.

A human may ask one question.

A capable agent may answer it by making fifty model calls, launching three subagents, retrieving ten documents, checking the result twice and monitoring a future event.

The economics of one response and the economics of delegated work are not the same.

This is where local compute connects to the idea I explored in The Compounding Class.

Access to an answer is different from the ability to commission machine work. And the ability to commission machine work is different again from owning the infrastructure on which that work runs.

Local AI hardware lowers the distance between those layers.

It turns part of machine intelligence from a service expense into a productive asset.

Not for everyone.

Not for every workload.

But enough to change the boundary.

This does not kill the cloud

It is tempting to tell the story as a pendulum.

Mainframe. PC. Cloud. Local AI.

History is rarely that tidy.

The more plausible future is additive.

Gartner's September 2026 forecast still describes hyperscaler and service-provider AI infrastructure as the largest single area of AI spending, with the data-centre buildout continuing at extraordinary scale.13

At the same time, endpoint vendors are putting more memory and accelerator capacity into personal machines.

Both can be rational because they solve different problems.

Training a frontier model remains a data-centre activity. Serving millions of users benefits from pooled infrastructure. A rare, extremely difficult reasoning task may justify a cloud model that is much stronger than anything practical to run on a laptop. Enterprise systems may prefer centrally governed private clusters. Consumer devices may need cloud escalation when battery, memory or latency constraints are reached.

Apple's Private Cloud Compute architecture expresses exactly this split from another direction: process locally where possible, then move selected harder requests into protected cloud infrastructure when the device is not enough.14

NVIDIA and Microsoft are building routing into their agent stack for similar reasons. OpenShell can route model requests according to policy; PAIR can route them across local nodes; cloud endpoints remain available when permitted.1011

The future therefore looks less like “the cloud versus the PC” and more like a hierarchy of intelligence.

The closest adequate compute wins.

A small model handles routine work.

A larger local model handles private or sustained work.

Nearby owned machines absorb parallel jobs.

Private infrastructure handles organisational workloads.

Frontier cloud systems receive the tasks that genuinely justify them.

This is not decentralisation in the ideological sense.

It is locality as an engineering decision.

Four stacked layers: a personal device at the bottom, then a local fabric of owned nodes, then racks of private organisational compute, and large public-cloud halls at the top. Six requests leave the device and fewer pass each red policy gate; only one reaches the cloud. A wedge on the left widens upward and a wedge on the right widens downward.
The closest adequate compute wins. Locality becomes a policy decision rather than an ideology. Reading the figure, bottom to top: device (private, lowest latency, constrained capacity); local fabric (PAIR, workstation, mini-PC or NAS-style nodes); private organisational compute (on-premises or sovereign/private cloud); frontier public cloud (maximum capability, highest external dependency). The left wedge stands for scale and capability, the right wedge for proximity and control. Each red gate means: escalate only when needed and permitted. Sources: local routing across owned devices from NVIDIA, PAIR Virtual Inference Router (3 September 2026) and PAIR FAQ; policy-based inference routing from NVIDIA, OpenShell (accessed 30 September 2026); local-first with selective cloud escalation from Apple Security Research, Expanding Private Cloud Compute (8 June 2026); AI NAS and clustered mini-PCs from IDC at IFA 2026 (September 2026).

What local compute gives back

There are real sovereignty gains here.

The first is continuity.

If a model, its weights, its inference engine and the necessary tools are already on the machine, an API outage or pricing change does not immediately remove the capability. Some functions can continue without an external model provider.

The second is data locality.

A contract, codebase, family archive, research corpus or client document does not need to leave the local environment merely to be summarised, classified or searched. This does not eliminate every data risk, but it removes an entire class of transmission by default.

The third is latency.

The round trip to a remote model disappears for tasks that fit locally. That matters for voice, coding, vision, interaction with applications and agents that make many small decisions.

The fourth is cost control.

Capacity can be planned as a capital purchase rather than only as metered inference. Organisations can decide where the break-even lies for repetitive workloads.

The fifth is model choice.

Open-weight models can be retained at a known revision, tested, fine-tuned, replaced and operated after a vendor changes its public API strategy.

The sixth is authority over execution.

When the agent runtime, policy and data are local, it becomes technically possible to place the authority boundary around the user's environment rather than around the provider's service.

Those are significant improvements.

They are also not sovereignty by themselves.

Local is not sovereign

A machine can sit under my desk and still belong, operationally, to a supply chain I do not control.

RTX Spark illustrates the paradox beautifully.

The platform is local.

Its dependencies are global.

NVIDIA contributes the Blackwell GPU architecture, CUDA and the AI software stack. MediaTek says it contributed CPU, memory-controller, power and connectivity expertise and explicitly references its manufacturing partnership with TSMC.15 The CPU uses the Arm architecture. The operating system is Windows. Drivers, firmware, compiler toolchains, model formats, inference engines and distribution hubs remain part of the usable system.

The result may give me more control over where intelligence runs while making one vendor's software ecosystem even more important to how it runs.

A dependency stack in which a small computer holds the upper layers from application down to SoC and memory. Below a red horizontal cut, wider layers continue under the machine: silicon IP, fabrication and packaging, and energy, network, replacement supply and engineering competence.
A model in the building is only one layer. Physical possession can reduce exposure without removing the supply chain beneath it. Reading the figure, top to bottom: agent or application; model weights; inference engine; CUDA or accelerator runtime; operating system, identity and policy; firmware and drivers; SoC and memory; Arm IP, MediaTek design inputs and NVIDIA IP; semiconductor fabrication and packaging; energy, network, replacement supply and engineering competence. The red line separates what is physically local (above) from what remains externally dependent (below). Sources: co-design and manufacturing dependencies (CPU, memory controller, power and connectivity contributions; TSMC manufacturing partnership) from MediaTek, 1 June 2026; Arm CPU architecture and unified memory from NVIDIA, Windows on Arm Porting Guide — System Overview (accessed 30 September 2026).

That is the exact distinction behind Open Is Not Sovereign.

A model can be open-weight and still depend on a concentrated runtime.

A computer can be physically mine and still be difficult to operate outside one accelerator ecosystem.

A local agent can avoid a cloud API and still depend on a Microsoft identity layer, signed drivers, model registries and software updates.

A machine can continue working during an internet outage and still be impossible for my organisation to replace within a reasonable time if the hardware supply disappears.

Sovereignty is not a location field.

It is a system property.

The correct question is therefore not:

Is this AI local?

It is:

If one important supplier, service or software layer changed tomorrow, which capabilities could I still operate, repair, replace and understand?

RTX Spark improves some answers to that question.

It does not answer all of them.

NVIDIA is attacking the centre of the PC

For decades, the Windows performance PC had an almost ritual architecture.

Intel or AMD supplied the CPU.

NVIDIA or AMD supplied a GPU.

System memory sat on the CPU side. VRAM sat on the graphics side. PCIe connected the worlds.

The operating system and application ecosystem were overwhelmingly x86.

RTX Spark attacks that arrangement from several directions at once.

The CPU is Arm-based. The GPU is integrated into the same coherent platform. Memory is unified. CUDA is native. The same device is pitched simultaneously at gaming, creation and agentic inference. And the platform is shipping through ASUS, Dell, HP, Lenovo, Microsoft and MSI, with others following.12

This does not mean Intel and AMD disappear.

AMD already has a strong answer in Ryzen AI Max, where x86 compatibility, large unified memory and capable integrated graphics make local AI possible without leaving the traditional Windows instruction set.6

Apple demonstrated the strategic power of vertical integration years earlier: own the SoC architecture, memory model, operating system, developer framework and hardware design, then optimise the whole machine as one product. Current Apple silicon continues to push that approach into local AI.7

Qualcomm has already pushed Windows toward Arm with Snapdragon.

What NVIDIA changes is the competitive centre of gravity.

For the first time, the company that became the dominant platform for AI acceleration is attempting to own the complete high-performance PC compute substrate as well: CPU, GPU, interconnect, memory architecture, compiler/runtime and agent stack.

The old question was which CPU belongs at the centre of the PC.

The new question may be whether the PC still has a single centre at all.

The real product is the stack

A chip can be copied in concept.

A stack is harder.

RTX Spark is interesting because NVIDIA is not selling a block diagram. It is bringing CUDA, TensorRT, RTX graphics, Windows on Arm tooling, llama.cpp and vLLM optimisation, local-agent integrations, OpenShell, PAIR and OEM designs around the same platform.121011

That matters strategically.

A technically superior accelerator with poor model support is a benchmark result.

A platform that developers can target, package, secure, monitor and deploy becomes infrastructure.

This is where NVIDIA's strength can become somebody else's dependency.

CUDA has value because it works, because developers know it, because libraries target it and because models are routinely optimised for it.

Every one of those reasons is legitimate.

Together they create path dependence.

The more complete the platform becomes, the easier it is to choose.

The easier it is to choose, the more expensive it becomes to leave.

That does not make the choice wrong.

It makes exit architecture part of the design.

If I were evaluating local AI infrastructure for an organisation, I would therefore care about more than tokens per second. I would ask whether the models can move, whether data formats are portable, whether the agent policy can be reproduced elsewhere, whether the inference API is standard enough to redirect, whether staff understand a second stack, and whether a failed device can be replaced by something that does not share the same dependency.

The fastest machine is not necessarily the most sovereign machine.

The sovereign system is the one whose capability survives a change of machine.

Agents make sovereignty more personal

There is another reason this matters.

The closer AI gets to the individual, the less abstract the sovereignty problem becomes.

A cloud chatbot knows what I type into it.

A persistent personal agent may know my files, appointments, contacts, preferences, accounts, devices, messages, projects and routines. It may be able to act on some of them.

That agent is not merely a productivity tool.

It becomes an interface to my digital life.

Running more of that intelligence locally can be an important protection. The raw personal context can remain under a boundary I control. Sensitive retrieval does not need to be repeated through an external service. A private model can work over information that I would never intentionally upload to a general-purpose assistant.

But proximity increases authority as well as privacy.

A local agent can be closer to credentials, private keys, source code, local applications and authenticated sessions than a cloud chatbot ever was.

That is why agent identity, sandboxing and audit are not optional add-ons to the AI PC.

They are part of what makes the category viable.

The personal AI computer only becomes a sovereignty tool if the person can still define the agent's authority, observe its actions, revoke it, replace it and recover without it.

Otherwise we have moved the dependency from a cloud account into a box on the desk.

That would be locality without agency.

What this means for Switzerland

The Swiss sovereignty conversation often starts too high in the stack.

We ask where the cloud region is, where the data is stored, whether the provider is Swiss and whether a contract contains the right jurisdiction clause.

Those questions matter.

Local AI creates another option below them.

A law firm, engineering company, archive, industrial business, research group, bank, hospital or public institution can increasingly perform useful AI work without sending every document and every intermediate reasoning step into a shared remote inference service.

That can reduce external exposure and improve continuity.

For smaller organisations, the more radical change is economic. Workloads that previously implied a rack server or rented GPU can increasingly fit into a workstation, mini PC or high-end laptop. The technical threshold for owning useful AI capacity is falling.

That is strategically attractive for a small country built around specialised companies, confidential knowledge and expensive labour.

But Switzerland should not mistake procurement for sovereignty.

Buying thousands of local AI computers does not create a domestic semiconductor industry. It does not create an alternative to CUDA. It does not guarantee access to replacement hardware during a supply shock. It does not create local model competence by itself.

The useful policy and architecture question is narrower and more practical:

Which critical AI capabilities can we make locally operable, portable and recoverable even when the wider supply chain remains international?

That is achievable.

And it is more serious than pretending independence means manufacturing every transistor ourselves.

The cloud made compute invisible. Agents make it visible again.

Cloud computing taught an entire generation to stop thinking about machines.

That was one of its triumphs.

Provision an instance. Call an endpoint. Add a subscription. Scale when needed. The physical computer disappeared behind an abstraction.

AI is reversing part of that abstraction because intelligence has locality requirements.

Memory matters.

Latency matters.

Privacy matters.

Energy matters.

Identity matters.

The physical proximity between the agent and the systems it can act upon matters.

And once an agent is expected to work continuously, the question of who owns the machine that keeps it alive becomes economically meaningful again.

This is why I think RTX Spark deserves more attention than another processor launch.

Not because NVIDIA has ended the traditional PC in one evening.

It has not.

Not because one petaflop on a laptop makes the cloud obsolete.

It does not.

And not because a locally running model makes its owner sovereign.

It does not.

RTX Spark matters because it joins several trends that had been developing separately and makes them legible as one product category:

large unified memory;

accelerated local inference;

personal agents;

operating-system containment;

local compute routing;

and a PC industry increasingly willing to treat AI as a resident workload rather than a remote service.

The traditional personal computer ran my applications.

The next one may host my delegated machine labour.

That is a different relationship with the device.

It is also a different relationship with the companies behind it.

The computer comes home

The cloud turned the computer into a window onto somebody else's machine.

Agentic AI may turn it back into a machine worth owning.

Not because the future is offline.

Not because every model will fit under the desk.

Not because local hardware frees us from dependencies.

Because an increasingly important layer of intelligence can once again live close to the person, organisation and data that give it authority.

That gives us choices we did not have when every useful model call had to cross a provider boundary.

We can keep some context local.

We can own some capacity.

We can continue some work without an API.

We can decide which tasks deserve the cloud instead of assuming the cloud by default.

We can build agents whose policy boundary belongs to the environment they serve.

And we can begin treating compute not merely as a utility we rent, but as a capability we may sometimes choose to possess.

That still does not make the machine sovereign.

The point of sovereignty was never the machine.

The point is preserving meaningful choice above it.

So the question I would ask about RTX Spark is not whether it is the future of the PC.

The better question is what kind of future becomes possible when the intelligence working for us no longer has to live entirely on somebody else's computer.

For the first time in the cloud era, the answer may be sitting on the desk.

Research note

RTX Spark systems were announced in May 2026 and, as of the evidence cutoff on 30 September 2026, NVIDIA says the first systems will arrive in October. Retail pricing, independent production-system benchmarks, sustained thermals, battery behaviour, software compatibility and real-world performance of the 120B / long-context claims remain incomplete. Vendor performance statements in this volume are identified as such and should not be read as independently reproduced results.

The broader argument does not depend on RTX Spark becoming the fastest product in every benchmark. AMD, Apple, DGX Spark and other local-AI systems already demonstrate that high-memory endpoint inference is an industry trend. RTX Spark is treated here as a particularly visible convergence of Windows, CUDA, unified memory and agent infrastructure.


Sources

Architecture, specifications, availability and security features come primarily from vendor documentation and announcements. Performance figures from NVIDIA, AMD, Apple and PC makers are vendor claims, not independent benchmarks. RTX Spark retail systems were not yet broadly shipping at the evidence cutoff, so pricing, sustained performance, battery behaviour and software compatibility remain open. Analyst material from Gartner and IDC is forecast or conference observation, not shipment data. Evidence cutoff: 30 September 2026, Europe/Zurich.

Footnotes

  1. NVIDIA, NVIDIA and Microsoft Reinvent Windows PCs for the Age of Personal AI, 31 May 2026. https://nvidianews.nvidia.com/news/nvidia-microsoft-windows-pcs-agents-rtx-spark ↩ ↩2 ↩3 ↩4

  2. NVIDIA, Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026, 3 September 2026. https://blogs.nvidia.com/blog/local-ai-ifa-next-gen-agents-nv-pair-rtx-spark/ ↩ ↩2 ↩3

  3. NVIDIA, Windows on Arm Porting Guide — System Overview, accessed 30 September 2026. https://docs.nvidia.com/rtx-spark/rtx-spark-porting-guide/latest/overview.html ↩ ↩2

  4. OpenAI, gpt-oss-120b & gpt-oss-20b Model Card, 5 August 2025. https://openai.com/index/gpt-oss-model-card/ ↩

  5. PCWorld, Nvidia’s RTX Spark PCs launch in October with bold 100fps gaming promises, 3 September 2026. https://www.pcworld.com/article/3225974/nvidias-rtx-spark-pcs-launch-in-october-with-bold-100fps-gaming-promises.html ↩

  6. AMD, FAQs: AMD Variable Graphics Memory, VRAM, AI Model Sizes, Quantization, MCP and More!, 29 July 2025; AMD, Agent Computers, accessed 30 September 2026. https://www.amd.com/en/blogs/2025/faqs-amd-variable-graphics-memory-vram-ai-model-sizes-quantization-mcp-more.html and https://www.amd.com/en/products/processors/consumer/agent-computers.html ↩ ↩2

  7. Apple, Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute, 25 August 2026. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/ ↩ ↩2

  8. Microsoft, Build 2026: Furthering Windows as the trusted platform for development, 2 June 2026. https://blogs.windows.com/windowsdeveloper/2026/06/02/build-2026-furthering-windows-as-the-trusted-platform-for-development/ ↩

  9. Microsoft, Windows Agentic — Build and run agents locally on Windows, accessed 30 September 2026. https://developer.microsoft.com/en-us/windows/agentic ↩

  10. NVIDIA, NVIDIA OpenShell — Open, Secure Runtime for AI Agents, accessed 30 September 2026. https://www.nvidia.com/en-us/ai/openshell/ ↩ ↩2 ↩3

  11. NVIDIA, NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network, 3 September 2026. https://developer.nvidia.com/blog/nvidia-pair-virtual-inference-router-expands-available-compute-on-your-local-network/ ↩ ↩2 ↩3 ↩4

  12. IDC, IDC on the Ground at IFA 2026: AI Set the Design Agenda for Every Device Category, September 2026. https://www.idc.com/resource-center/blog/idc-on-the-ground-at-ifa-2026-ai-set-the-design-agenda-for-every-device-category/ ↩

  13. Gartner, Gartner Forecasts Worldwide AI Spending to Grow 49.5% in 2026, 16 September 2026. https://www.gartner.com/en/newsroom/press-releases/2026-09-16-gartner-forecasts-worldwide-ai-spending-to-grow-49-point-5-percent-in-2026 ↩

  14. Apple Security Research, Expanding Private Cloud Compute, 8 June 2026. https://security.apple.com/blog/expanding-pcc/ ↩

  15. MediaTek, MediaTek Collaborates with NVIDIA on RTX Spark to Power the Next Wave of Windows PC Experiences, 1 June 2026. https://www.mediatek.com/press-room/mediatek-collaborates-with-nvidia-on-rtx-spark-to-power-the-next-wave-of-windows-pc-experiences ↩

Leave a note in the margin

Your submission is stored privately until reviewed or deleted. Only an edited, accepted note can appear publicly. Do not include confidential information. Attribution is optional.