News

NVIDIA DGX Spark in Malaysia: a guide for companies

The strongest argument for buying an AI machine outright has very little to do with performance. It’s paperwork.

Malaysia’s amended Personal Data Protection Act came into force in stages through the first half of 2025, and it changed what it costs, in effort rather than ringgit, to send customer data to a hosted AI service. Every one of those requests is a cross-border transfer. You need a lawful basis for it, you may need an assessment on file to justify it, and the ceiling on penalties for getting the principles wrong moved from RM 300,000 to RM 1,000,000.

A DGX Spark sidesteps the question entirely, because the data never leaves the room. EMARQUE stocks seven versions of it, with current Malaysian pricing listed on the DGX Spark range.

That is the case for one. What follows is the case against, and what the hardware actually does once the compliance argument has finished doing its work.

The rule that changed

Three pieces of the amended Act matter if you’re weighing hosted AI against local.

The old whitelist approach to cross-border transfers is gone, replaced by a risk-based framework, and the Commissioner published the Cross Border Personal Data Transfer Guidelines on 29 April 2025. If you rely on the destination country offering adequate protection, you need a Transfer Impact Assessment on file. It stays valid for up to three years before somebody has to sit down and do it again.

From June 2025, organisations carrying out large-scale processing have to appoint a data protection officer, and that obligation reaches processors, not only controllers.

The penalty ceiling moved at the same time: maximum fine from RM 300,000 to RM 1,000,000, maximum imprisonment from two years to three.

None of this makes cloud AI illegal, and anyone telling you otherwise is selling something. What it does is convert a technical decision into a documented one. That cost rarely appears in the spreadsheet comparing API pricing against hardware, and it is the cost that grows quietly as usage spreads across departments.

Not only a Malaysian anxiety

IBM’s Institute for Business Value surveyed a thousand senior executives across sixteen countries between February and April this year, with Oxford Economics. Sixty-eight percent said meeting data residency and sovereignty requirements across geographies was difficult.

Two other findings from the same study are worth sitting with, because they describe a problem most companies discover late. Seventy-one percent said switching their primary AI vendor or model would be hard. Eighty-one percent said a seven-day outage at that vendor would cause severe or critical disruption.

Read those together and the compliance argument and the dependency argument turn out to be the same argument wearing different clothes.

What the machine is

A desktop computer built around NVIDIA’s GB10 Grace Blackwell chip, with 128GB of memory shared between the CPU and GPU rather than split between them.

That shared pool is the entire reason the thing exists. A consumer graphics card at the top of the range carries 32GB, which caps the size of model you can hold. 128GB puts models on a desk that otherwise need a server.

It runs off a wall socket

About 240W, through a USB-C brick.

This is the least interesting specification on the sheet and one of the most consequential. No rack, no server room, no three-phase supply, no dedicated cooling, no conversation with the building manager. For a company that has never bought infrastructure before, the absence of all that is often what makes the purchase possible at all.

Where the throughput actually is

Most reviews of this machine test it the way a hobbyist would use it, one person typing one prompt, and come away underwhelmed. That is the wrong test for a business, and it produces the wrong conclusion.

StorageReview measured it under concurrent load instead. Serving 128 simultaneous requests, it returned 924 tokens per second on Llama 3.1 8B at FP4, and 611.7 on gpt-oss 20B at NVFP4. Mistral Small 3.1 24B managed 319.7 tokens per second at the same concurrency, and Qwen3 Coder 30B reached 482.6 at a batch size of 64. Compute measured 99.8 TFLOPs in BF16 and 207.7 in FP8.

Those are throughput numbers, not latency numbers, and the distinction decides whether you should buy one. This is a machine for serving a team, an application, or an overnight batch. It is not a machine for making one impatient executive feel fast.

Where it runs out

Memory bandwidth, not compute.

On decode-heavy work, the kind that produces long output from a short prompt, throughput falls off noticeably, because the limit is how quickly weights move rather than how quickly arithmetic happens. Treat that as a characteristic of the design rather than a defect, but do go in knowing it. If your workload is mostly long-form generation for a small number of users, the economics look worse than the headline figures suggest.

One more thing from the same testing, unglamorous but expensive to discover late: the internal M.2 slot is small and slow. Budget for external storage if you intend to move real datasets around.

Fine-tuning, not training

Companies asking whether they can “train their own model” almost always mean fine-tuning, and the answer there is yes.

128GB is enough to do full-parameter fine-tuning of an 8B model, which is the step that separates a generic assistant from one that knows your product catalogue, your policies, and the particular way your people write. Training a frontier model from scratch is a different exercise on very different hardware, and nothing in this price class comes close.

Adding a second one

The units cluster, which matters more for the first purchase than for the second.

Each has ConnectX-7 networking built in. Two connect directly with a single 200Gbps cable. Three form a ring using both ports on each unit. Four or more need a managed switch with 200Gbps-class ports, and at four you are pooling 512GB of unified memory. NVIDIA ships a Cluster Assistant that handles the configuration.

The practical point is that buying one is not a dead end. Start with a single box, prove the workload, add the second when the queue justifies it rather than when a vendor suggests it.

Seven boxes, one chip

Every variant runs the same GB10 silicon and the same 128GB. Nobody is selling you a faster one.

What differs is storage, chassis, thermals and warranty, and warranty is the one that gets skipped in the comparison. Base terms vary materially between the seven manufacturers, and some offer a paid extension where others do not. On a machine at this price, running as infrastructure rather than as a toy, that gap is worth more than the badge on the front. Get the term in writing before you commit to a brand.

Availability moves around too. Of the seven, the GIGABYTE AI TOP ATOM and the Lenovo ThinkStation PGX are the two currently held in depth. The rest are on the DGX Spark range.

When not to buy one

Three situations, and they are more common than the marketing suggests.

If nobody in your organisation objects to company data reaching a hosted service, a subscription is cheaper, faster to start, and easier to abandon. The hardware earns its place when the data is the constraint, not when the enthusiasm is.

If what you want is a frontier model of your own, this is the wrong machine, and the gap is measured in data centres rather than in ringgit.

And if no one internally wants to own it, do not buy it. Infrastructure without a named owner becomes an expensive shelf, and no amount of specification fixes an absent champion.

The question that decides it

Can this data legally leave the country?

If the answer is no, local inference stops being a preference and becomes a requirement, and the machine starts paying for itself in avoided compliance work as much as in avoided API bills. If the answer is yes, be honest about that. Plenty of companies asking about this hardware do not need it, and the ones that say so out loud tend to make better decisions later.

EMARQUE is a solution provider for the DGX Spark family in Malaysia, and every unit ships with a complimentary EMARQUE AI Utility, installed on request, which handles the distance between unboxing the thing and having a model actually running on it.

The DGX Spark range lists every variant and what is in stock. If you would rather describe the workload than read specifications, send it over WhatsApp and the AI team will come back with a scoped answer.

Specifications, measured figures and regulatory details verified August 2026. This is general information about hardware, not legal advice; confirm your own obligations under the PDPA with a qualified adviser.