NVIDIA DGX Spark AI Supercomputer Malaysia

Build, fine-tune and run large AI models on hardware you own, with nothing leaving your building.

NVIDIA DGX Spark is a compact personal AI supercomputer built on the GB10 Grace Blackwell Superchip, with 128GB unified memory and up to 1 petaFLOP of AI performance. It suits AI developers, research teams, and businesses holding data they cannot send offshore.

Every DGX Spark from EMARQUE includes a complimentary pre-sales consultation. We size the unit to the models you actually plan to run, plan the deployment, and handle warranty and delivery here in Malaysia. More than 300 Malaysian companies have deployed with us.

EMARQUE AI Utility, the complimentary local AI software included with every DGX Spark from EMARQUE, shown with its dashboard, chat and model library.

Every DGX Spark from EMARQUE includes the EMARQUE AI Utility, installed on request.

Compare

Storage and warranty, model by model

Every DGX Spark shares the identical NVIDIA GB10 Grace Blackwell platform: 20-core Arm CPU, Blackwell GPU, 128GB unified memory, 273 GB/s bandwidth, up to 1 petaFLOP FP4, dual ConnectX-7 200GbE and NVIDIA DGX OS. Local RMA is handled by EMARQUE on every model.

NVIDIA DGX Spark models available in Malaysia, compared by storage capacity and warranty term
Brand & modelStorageWarranty
GIGABYTE AI TOP ATOM4TB (Gen4 / Gen5)3-year
MSI EdgeXpert4TB (Gen5 NVMe)3-year
Dell Pro Max with GB102TB / 4TB1-year (3-yr upgrade)
HP ZGX Nano G1n4TB1-year (3-yr upgrade)
Lenovo ThinkStation PGX4TB1-year (3-yr upgrade)
ASUS Ascent GX101TB / 2TB / 4TB1-year
NVIDIA DGX Spark Founders Edition4TB1-year

All seven run the same NVIDIA AI software stack, and EMARQUE can cluster up to four units for larger models. We have been building and deploying computing systems in Malaysia since 2016: sizing, deployment, local warranty handling, fast delivery and financing are all handled here.

FAQ

Questions, answered.

What you need to know before choosing NVIDIA DGX Spark: capabilities, comparisons, the upgrade path, and how EMARQUE helps Malaysian businesses deploy it on-premise. Reviewed and fact-checked September 2026.

  • Who is DGX Spark for?

    DGX Spark suits AI developers, data scientists, research teams, software teams building AI agents, businesses needing on-premise AI, and universities or labs running local AI workloads. The common need: build, test, fine-tune, and run AI models locally instead of relying on cloud GPUs.

  • What are people actually using DGX Spark for in 2026?

    The dominant use is private local LLM inference. Teams run models such as GPT-OSS-120B, Qwen and DeepSeek through Ollama or vLLM behind a browser front end, keeping prompts and documents entirely on hardware they own. The 128GB of unified memory is what makes this work: it holds models a consumer graphics card physically cannot, including FLUX 2 at full precision for image generation.

    Agentic AI has been the fastest-moving use through 2026. NVIDIA released Nemotron 3.5 Lightning tuned for DGX Spark as an agent-serving system, and in August 2026 Perplexity demonstrated a fully local agent stack running on a single unit, with the orchestrator, subagents and harness all on-device. Alongside that, teams use DGX Spark for LoRA and full fine-tuning of models up to roughly 70 billion parameters at 4-bit precision, for ComfyUI image and video work, and as an always-on internal intelligence layer running analytics and forecasting over private company data.

    One honest caveat, because it matters at this price. DGX Spark is not the fastest box per ringgit. Independent reviewers have measured comparable throughput from cheaper AMD Strix Halo mini PCs, and developer Simon Willison summarised the launch consensus as great hardware with an ecosystem still maturing. What you are buying is memory capacity plus the CUDA software stack, which essentially every major AI framework targets first. That is also why the machine keeps improving after purchase: a 2026 llama.cpp update alone delivered roughly a 35% average uplift on mixture-of-experts models with no hardware change.

  • What AI models can DGX Spark run, and can it fine-tune them?

    NVIDIA positions DGX Spark for models up to ~200B parameters. In real-world use, EMARQUE's testing and client deployments show ~120B parameters as the practical optimum on a single Spark; beyond that, throughput and latency degrade. NVIDIA's NVFP4 quantisation format reduces model memory use by roughly 40% while maintaining accuracy, which stretches what a single unit fits in practice.

    A single Spark fits well:

    • 7B–14B: excellent
    • 30B–32B: suitable
    • 70B: workload-dependent
    • ~120B: real-world ceiling
    • ~200B: pushes the limits
    • 405B and above: needs a 2–4 unit cluster or a step-up, covered under clustering below

    Fine-tuning: yes. Small-to-mid model fine-tuning, LoRA and QLoRA, private data adaptation, prototyping before scaling. Training from scratch: only small experiments. Large foundation model training, heavy multi-user inference or production AI serving need RTX PRO 6000 Blackwell, DGX Station GB300 or server-class systems.

  • What real-world performance can I expect from DGX Spark?

    DGX Spark is built for capacity, not peak speed. Its 128GB unified memory holds models a gaming GPU cannot, but its memory bandwidth is modest at 273 GB/s against an RTX 5090's 1,792 GB/s, and that sets the token-generation rate. From independent mid-2026 reviews, single-user, varying by model, quantisation and framework:

    • 8B–20B models: interactive, from tens of tokens per second single-user to hundreds batched
    • 70B models: comfortably usable, roughly 35–45 tokens per second
    • 120B models such as GPT-OSS 120B: around 38 tokens per second

    Those figures are single-user. Batched serving is a different picture: independent testing measured up to roughly 505 tokens per second on GPT-OSS-120B across a two-unit cluster at batch 64.

    The machine keeps getting faster after purchase, which matters on hardware you intend to keep. StorageReview's coverage of the CES 2026 enterprise update reports up to 2.5× on key workloads against launch, and around 8× on video generation, where an eight minute job came down to roughly one minute. NVIDIA's own published figure is 2.6× on Qwen-235B across two units, using NVFP4 with speculative decoding against FP8. The June 2026 release added multi-node clustering for up to four units through the Cluster Assistant in NVIDIA Sync.

    For peak speed on models that fit in 32GB, an RTX 5090 or RTX PRO 6000 is still faster. For running 70B to 120B models locally, DGX Spark is the compact option.

  • How is DGX Spark different from NVIDIA RTX Spark?

    The names are close and the tech press regularly confuses them, but they are different platforms. DGX Spark is the compact AI development machine on this page: NVIDIA's GB10 Grace Blackwell Superchip, 128GB unified memory, DGX OS with NVIDIA's AI software stack, available in Malaysia now across seven brands.

    RTX Spark is a separate family of Windows AI laptops and compact desktops built on the RTX Spark Superchip, with up to 1 petaflop of FP4 AI performance and up to 128GB unified memory, announced for Fall 2026 from ASUS, Dell, HP, Lenovo, Microsoft Surface and MSI. It runs Windows rather than DGX OS, and is aimed at AI-accelerated everyday computing rather than dedicated model development.

    If you want to develop, fine-tune and serve models locally today, DGX Spark is the one you can order. The NVIDIA RTX Spark lineup has details on what is coming.

  • How does DGX Spark compare to a custom AI PC with an RTX 5090?

    DGX Spark is built for AI development with NVIDIA's preconfigured AI software stack, not gaming. For gaming or general workstation use, choose an EMARQUE Gaming PC.

    For AI workloads against a custom PC built around an RTX 5090 with 32GB of GDDR7: DGX Spark wins when models exceed ~32GB. Its 128GB unified memory handles 70B–120B-class models that simply will not fit on a single RTX 5090, and the AI stack is preconfigured out of the box. It also draws far less power: a 240W external supply against 575W for the RTX 5090 GPU alone. The RTX 5090 wins for smaller models of 30B and under, with roughly 6.6× the memory bandwidth and higher raw FP8 and FP4 throughput for faster token generation, plus broader workstation flexibility for rendering and mixed workloads.

    Different sweet spots. EMARQUE builds both, and the EMARQUE AI team will tell you honestly which fits your workload better.

  • Which of the seven DGX Spark brands should I choose?

    EMARQUE stocks seven DGX Spark variants: NVIDIA Founders Edition, ASUS, Dell, HP, Lenovo, MSI and GIGABYTE. All seven are built on the same GB10 platform with the same 128GB unified memory, so raw performance is effectively identical. StorageReview benchmarked Dell, GIGABYTE and HP two-node clusters and measured them within a few tokens per second of each other, concluding that buyers should choose on chassis design, thermal behaviour, warranty terms and support relationship rather than benchmark deltas.

    That is how EMARQUE advises too. The practical differences are storage configuration, chassis and cooling character, and each manufacturer's warranty terms, which vary by brand. The comparison table above this FAQ sets the variants side by side, and the EMARQUE team will match a unit to your workload, budget and support expectations as part of the standard consultation.

  • DGX Spark, a Spark cluster, RTX PRO 6000 or DGX Station: which should I buy?

    • DGX Spark: local AI development, LLM and VLM testing, RAG, fine-tuning small-to-mid models.
    • 2–4 unit DGX Spark cluster: larger models via pooled memory and distributed local AI, covered under clustering above.
    • RTX PRO 6000 Blackwell: stronger raw GPU performance plus workstation flexibility for rendering, simulation and mixed AI-plus-creator workloads.
    • DGX Station GB300: when DGX Spark is too small and a workstation is not enough. Far more compute, trillion-parameter models, multi-user workloads, serious on-premise AI infrastructure.
    • Multi-GPU AI servers: rack-scale deployments beyond a single workstation.
    • Gaming? An EMARQUE Gaming PC, not DGX Spark.

    The EMARQUE AI team sizes this against your actual workload. Talk to the EMARQUE AI team on WhatsApp.

  • What warranty does DGX Spark come with in Malaysia?

    Warranty is the specification buyers most often skip, and it varies more than the hardware does. Every unit runs the same GB10 platform, so on a machine you plan to keep, the warranty term is often the more meaningful difference between two models at a similar price.

    • GIGABYTE AI TOP ATOM and MSI EdgeXpert: 3 years
    • NVIDIA DGX Spark Founders Edition and ASUS Ascent GX10: 1 year
    • HP ZGX Nano G1n, Dell Pro Max with GB10 and Lenovo ThinkStation PGX: 1 year, with a paid extension to 3 years available

    EMARQUE handles the local RMA routing with the manufacturer on every model, so a fault is dealt with in Malaysia rather than through an overseas support queue. The comparison table above sets the terms side by side.

    One clarification worth making: EMARQUE's 90-day one-to-one exchange and lifetime labour support apply to EMARQUE-built PCs. They do not apply to the DGX Spark family, which rests on the manufacturer warranty plus EMARQUE's local handling.

  • How much does NVIDIA DGX Spark cost in Malaysia?

    EMARQUE lists every DGX Spark variant with live Malaysian pricing on this page, currently from around RM 24,000 for an entry configuration up to roughly RM 36,000 for the largest 4TB builds. The product cards above carry the current figure for each brand and storage tier, and that is the number to trust: DGX Spark pricing moves with supply and exchange rates, so fixed figures quoted elsewhere go stale quickly.

    Three things drive the spread. Storage is the largest factor, with 1TB, 2TB and 4TB tiers sitting on the same GB10 platform and the same 128GB unified memory, so compute performance does not change with price. Chassis and brand account for a smaller difference. Warranty is the factor buyers most often overlook: GIGABYTE and MSI ship with a three-year warranty, while NVIDIA Founders Edition, ASUS, Dell, HP and Lenovo ship with one year, with paid multi-year upgrades offered on some. A cheaper unit on a one-year warranty is not always cheaper across three years of ownership.

    Listed prices are for local stock with local warranty handling and EMARQUE setup support included. Financing and corporate purchase orders are available.

  • Can I buy DGX Spark in Malaysia now, and how fast is delivery?

    Yes. DGX Spark is shipping, and EMARQUE holds local stock in Malaysia, with per-variant availability shown on each product page.

    In-stock units dispatch within 1–3 business days with insured nationwide courier delivery. Walk-in collection is available at EMARQUE HQ in the Klang Valley by appointment; DGX Spark is not stocked at the Penang collection point. Pre-deployment validation testing is available on request before dispatch.

  • Can I connect multiple DGX Spark systems together?

    Yes. Since NVIDIA's June 2026 update you can cluster up to four DGX Spark units, not just two. Each unit has two ConnectX-7 200Gb/s ports, and NVIDIA's Cluster Assistant in NVIDIA Sync configures the cluster:

    • 2 units: one QSFP cable, no switch. 256GB unified memory, models up to ~405B parameters.
    • 3 units: direct ring topology using both ports, no switch. 384GB unified memory.
    • 4 units: via a managed 200GbE QSFP switch. 512GB unified memory, enough for large mixture-of-experts models and multi-agent pipelines.

    A four-unit build is still office-scale: the whole cluster draws under 1,200W, within an ordinary office circuit. Two configuration details catch self-builders. Switch auto-negotiation often sets ports to 50Gb/s and must be set to 200Gb/s manually, and real-world throughput of around 100Gb/s per link is expected behaviour from the ConnectX-7's PCIe interface, not a fault. Firmware and software also need to match across every node.

    EMARQUE supplies and configures full clusters as a service from the EMARQUE AI team: units, switch, cabling and Cluster Assistant setup, scoped and quoted after a discussion of the deployment. The result is a working multi-node system rather than a box of parts.

  • What does DGX Spark need to run, and can it be rack mounted?

    Very little, which is much of the point. A single unit draws around 240W through a USB-C power adapter, so it runs from an ordinary wall socket. No rack, no server room, no three-phase supply and no dedicated cooling. A four-unit cluster still stays under 1,200W in total, which an ordinary office circuit handles.

    Two practical notes. The I/O is USB-C, so most existing keyboards, mice and older peripherals need a USB-C to USB-A adapter. And DGX Spark is a desktop unit rather than a rackmount server, so multi-unit deployments usually sit on a shelf rather than in rails. EMARQUE can advise on shelving, switching and cabling when scoping a cluster.

  • What software comes with DGX Spark, and what happens if setup goes wrong?

    DGX Spark arrives with NVIDIA's DGX OS and preinstalled AI software stack: CUDA, container runtimes, NIM microservices, AI Workbench and pre-built framework images including PyTorch and vLLM. NVIDIA AI Enterprise is a separate entitlement, available as a purchase or a 90-day trial, not something bundled with the machine.

    Setup friction is among the most discussed DGX Spark topics on NVIDIA's own forums. First-boot network configuration, monitor detection, SSH access and cluster networking are where new owners lose their first week. That is the part buying locally changes. Every EMARQUE unit can be delivered ready to run, with the complimentary EMARQUE AI Utility installed on request, so the machine is usable from first boot without touching a command line.

    After handover, support is Malaysia-based and comes through your assigned EMARQUE account manager rather than a ticket queue. Larger jobs, such as a multi-unit cluster or a bespoke local AI stack with private RAG pipelines and model serving, are scoped and quoted by the EMARQUE AI team. Where a manufacturer hardware fault arises, EMARQUE handles the local RMA routing with the brand concerned.

  • Why run AI on-premise in Malaysia instead of the cloud?

    For Malaysian organisations the deciding factor is usually regulatory rather than technical. Under the amended PDPA, sending personal data to a hosted AI service abroad is a cross-border transfer. It needs a lawful basis, and where you rely on the destination country offering comparable protection, the Commissioner's Cross Border Personal Data Transfer Guidelines of 29 April 2025 expect a Transfer Impact Assessment on file, valid for up to three years. Breaching the Data Protection Principles now carries a fine of up to RM 1,000,000, up to three years' imprisonment, or both.

    Running the workload on hardware inside your own building removes the cross-border question entirely, along with the number of external parties handling the data. On-premise also holds up on cost against metered GPU billing, on latency with no round trip, and on control of the model and stack.

    Our guide to DGX Spark for Malaysian companies works through the detail, including the breach notification duties, the Data Protection Officer threshold and the cases where buying one is the wrong call.

    General guidance, not legal advice. Confirm your obligations with your Data Protection Officer or counsel.