GEEKOM squeezes a 250K-token AI workload into four small boxes
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
Geekom has released the DeepSeek V4 Flash model across four A9 Mega mini PCs, forming a distributed cluster aimed at enterprise AI workloads.
The setup connects the devices through USB4 rather than relying on a traditional data center server.
Each A9 Mega runs on the AMD Ryzen AI Max+ 395 chip, which combines 16 Zen 5 CPU cores with Radeon 8060S graphics and unified memory in one compact chassis
The four-node configuration brings a combined 512GB of RAM to the task, with Ubuntu, ROCm, and DwarfStar software distributing the optimized model across the machines.
An OpenAI-compatible API links applications and AI agents to the cluster, while USB4 removes any need for a proprietary switch or server rack.
The arrangement allows organizations to keep prompts, documents, source code, and credentials within local infrastructure rather than routing them through a public cloud.
Businesses could theoretically build private knowledge assistants capable of searching contracts, manuals, and internal reports without external exposure.
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
For agent-based systems, Geekom says the cluster can process tools, policies, memory, and logs before executing any action, with a technical path that has operated using contexts up to 250K tokens.
Testing reported by Geekom recorded approximately 14.61 tokens per second at single concurrency, while P95 time to first token reached about 0.42 seconds.
The company says the configuration also provides greater capacity for long prompts, rather than concentrating solely on faster short-response generation.
Users can begin with one or two A9 Mega systems before expanding the configuration to four nodes as workloads increase over time.
Each machine can operate independently, while connected systems can contribute to distributed inference when greater computing capacity becomes necessary for demanding workloads.