Samsung SDS Launches FuriosaAI NPU Cloud Service

For practitioners, cloud access to inference-focused NPUs broadens the infrastructure choices available for serving AI workloads without dedicated accelerator procurement. Samsung SDS launched NPU-as-a-Service on Samsung Cloud Platform using FuriosaAI's second-generation RNGD accelerator, according to ChosunBiz. The service offers configurations of one, two, four, or eight NPU cards and operates in Samsung SDS's sovereign-cloud environment, ChosunBiz reports. KM Journal reported that Samsung SDS also intends to use the accelerator for its FabriX AI agent platform and Brity Works collaboration suite, and that it has installed chips at data centers in Seoul and Hwaseong.
Inference compute enters the cloud menu
For practitioners, subscription access to inference-oriented accelerators can make workload placement a more practical engineering decision than a capital-procurement exercise. Industry context: hybrid GPU-NPU environments commonly separate high-throughput model training from steady-state inference, where throughput per watt, software compatibility, batching behavior, and operational cost can be decisive evaluation criteria.
Samsung SDS launched NPU-as-a-Service, or NPUaaS, on the Samsung Cloud Platform using FuriosaAI's second-generation RNGD neural processing unit, according to ChosunBiz. The offering lets customers subscribe to NPU servers in configurations of one, two, four, or eight cards, and can connect with Samsung Cloud Platform storage, compute, and high-speed networking services, ChosunBiz reports.
ChosunBiz describes RNGD as a South Korean inference accelerator for workloads such as answer generation, document analysis, and image classification. The publication also reports that Samsung SDS is providing the service in a sovereign-cloud environment intended to support organizations, including public-sector users, with stringent security requirements.
KM Journal reported that Samsung SDS had installed the chips at data centers in Sangam, Seoul, and Dongtan, Hwaseong. It also reported that Choi Jung-jin, a Samsung SDS group leader, announced at the 2026 K-NPU Tech Wave event that the company had secured customers and would expand capacity in stages through the end of the year. The same report states that Samsung SDS intends to use the accelerator for its FabriX AI agent platform and Brity Works collaboration suite.
The engineering question is workload fit
Editorial analysis - technical context
an NPUaaS option does not eliminate the need for accelerator qualification. Teams evaluating an inference service typically need to test model framework support, compiler and runtime maturity, quantization paths, supported operators, multi-card scaling, observability, and failure recovery alongside headline price-performance figures. These checks are particularly important for production LLM inference, where model architecture, context length, request concurrency, and batching policies can alter cost and latency materially.
KM Journal reported that Samsung SDS is developing a system to compare cost and performance before assigning workloads to GPUs or NPUs. If deployed as described, such placement tooling would align with a broader infrastructure pattern: organizations increasingly require measurable routing policies rather than assuming one accelerator type is optimal for all AI tasks.
For practitioners
the reported card-level configurations provide a concrete basis for capacity testing across smaller pilots and larger serving deployments. Useful public indicators to monitor include supported model families, availability of standard serving stacks, benchmark methodology, pricing, regional capacity, and the extent to which sovereign-cloud controls apply to data, model artifacts, and operational logs.
Key Points
- 1Samsung SDS added FuriosaAI RNGD NPU capacity to its cloud, expanding accelerator options for AI inference deployments.
- 2Reported one-to-eight-card configurations let teams size NPU trials and serving capacity without purchasing dedicated on-premises servers.
- 3Industry patterns show hybrid GPU-NPU deployments require workload-specific benchmarking across latency, throughput, runtime compatibility, and total operating cost.
Scoring Rationale
The launch is a notable commercial deployment of a South Korean inference accelerator through a major enterprise cloud platform. It is relevant to practitioners assessing non-GPU serving infrastructure, although public technical specifications, pricing, and benchmark results remain limited in the available reporting.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

