Local Agentic AI Machines for Qwen3.8-27B
A practical look at local agentic AI hardware for Qwen3.8-27B, from Apple Silicon and the RTX 5090 to DGX Spark, Intel Arc Pro, and AMD Radeon AI PRO.
Qwen3.8-27B can run at its full 262,144-token context on a 32 GB GPU.
That does not make every 32 GB machine equally good for agentic coding.
A coding agent is not a single clean inference benchmark. It reads a repository, fills and reuses a prompt cache, calls tools, receives terminal output, edits files, and keeps going. The machine may also be running an embedding model, reranker, browser, IDE, containers, and sometimes image generation at the same time.
For this workload, the important hardware characteristics are:
- enough memory for model weights, KV cache, runtime buffers, and the rest of the system;
- memory bandwidth for decode speed;
- compute throughput for prefill and prompt ingestion;
- a mature software stack for the models and agent tools being used; and
- enough headroom that full context is usable rather than merely possible.
The interesting part is that the current hardware options reach those goals in very different ways. Apple offers large unified-memory pools and very high bandwidth. NVIDIA offers the strongest CUDA ecosystem and the fastest single-GPU path when a model fits. DGX Spark trades bandwidth for a much larger coherent memory pool. Intel and AMD now offer 32 GB workstation-class alternatives that are especially interesting on Linux.
The short version
For Qwen3.8-27B, these are the most interesting local-agentic-AI configurations right now:
| Platform | Approx. price / class | Memory | Bandwidth | Main advantage | Main trade-off |
|---|---|---|---|---|---|
| Mac Studio M5 Max, 40-core GPU | $5,099 | 128 GB unified | 614 GB/s | Large memory pool in a quiet all-in-one workstation | About half the memory bandwidth of M5 Ultra |
| Mac Studio M5 Ultra, 64-core GPU | $5,499 | 96 GB unified | 1.2 TB/s | Very high bandwidth plus a much larger GPU and CPU | Less total memory than the 128 GB M5 Max; no CUDA |
| RTX 5090 | $1,999 launch MSRP; current street pricing is far higher | 32 GB GDDR7 | 1.79 TB/s | Fastest single-GPU path here when the workload fits; strongest software support | 32 GB is tight for full-context agentic use; extreme current pricing |
| DGX Spark | premium desktop AI appliance | 128 GB unified | 273 GB/s | Large coherent memory pool with NVIDIA's CUDA stack | Much lower bandwidth; Arm64 host |
| Intel Arc Pro B70 | 32 GB workstation GPU | 32 GB ECC GDDR6 | 608 GB/s | Large VRAM for a single lower-cost card | Software path is more backend-sensitive than CUDA or MLX |
| Radeon AI PRO R9700 | 32 GB workstation GPU | 32 GB ECC GDDR6 | 640 GB/s | Strong memory capacity and improving ROCm/Vulkan support | Linux is the cleaner AI path; Windows support is less complete |
The RTX 5090 deserves a pricing footnote. NVIDIA launched it at $1,999, and at anything close to MSRP it is the most aggressive performance option in this group. In September 2026, however, ordinary retail captures have been around the mid-$4,000s, while some third-party listings have spiked into the $7,000-$9,000+ range. That changes the comparison dramatically: the GPU can cost as much as, or more than, an entire high-memory workstation.
The two Mac configurations are also worth looking at together. The 128 GB M5 Max has the larger memory pool, but the 96 GB M5 Ultra roughly doubles memory bandwidth, increases the GPU from 40 to 64 cores, and provides a larger CPU. For a 27B agentic coding model that already fits comfortably inside 96 GB, the Ultra's extra bandwidth and compute can matter more than the Max's additional 32 GB of capacity. The M5 Max becomes more interesting when fitting a larger model is more important than maximizing inference speed.
Why full context barely fits in 32 GB
Qwen3.8-27B uses a hybrid architecture. Most of its layers use a fixed-state linear-attention design, while 16 layers use conventional full attention. Only those full-attention layers grow a normal KV cache with every token.
With a strong 4-bit quantization, the rough full-context allocation is:

| Allocation | Approximate memory |
|---|---|
| UD-Q4_K_XL model weights | 17.6 GB |
| Q8 KV cache at about 262K tokens | 9 GB |
| Runtime context and scratch buffers | 1.5 GB |
| Total allocated | 28.1 GB |
| Nominal headroom on a 32 GB card | 3.9 GB |
The rough budget is:
usable model memory = total memory - system use - weights - KV cache - runtime buffers
For a conventional attention layer, KV-cache memory grows roughly as:
tokens × layers × KV heads × head size × 2 (K + V)
× bytes per value × parallel slots
Grouped-query attention reduces the KV-head count. Q8 or Q4 caches reduce the bytes per value. Hybrid, recurrent, and sliding-window architectures can reduce how much history needs a conventional cache. Multiple concurrent requests move in the other direction: two slots can require close to twice the cache allocation.
Runtime memory is also real. Backends reserve temporary buffers, compute graphs, kernels, prompt caches, and sometimes an MTP or speculative-decoding model. Agent frameworks add tool state and may keep more than one sequence alive. On a desktop GPU, the display and ComfyUI can take memory too.
That is why 32 GB is a target to hit, while 64–96 GB is room to work.
Prefill and decode are different workloads
Local-LLM performance is often collapsed into one tokens-per-second number. That hides the two phases that matter to a coding agent.
Prefill is the model reading the prompt. Repository maps, source files, terminal history, documentation, and retrieved chunks all pass through the model before it produces the first token. Compute throughput and optimized attention kernels matter here. A huge context can take minutes to ingest even when it fits in memory.
Decode is the model writing the answer one token at a time. For a dense 27B model, decode is usually dominated by repeatedly streaming weights and active state from memory. This is where memory bandwidth becomes an especially useful predictor.
In one long-context RTX 5090 test, a short prompt decoded at about 62 tokens per second. Around 130K occupied tokens, prefill took roughly 84 seconds and decode fell to about 42 tokens per second. Near 262K, prefill took roughly 267 seconds and decode landed around 30 tokens per second.
The lesson is not that full context is bad. It is that allocating a 262K context window is not the same as filling it, and filling it is not free.
For an agentic workflow, I would normally run 64K or 128K, enable prefix caching, and let retrieval decide what enters the prompt. I would keep 262K available for repository-wide analysis, large migrations, and long autonomous runs. That gives the agent room without paying the maximum prefill cost on every request.
M5 Max 128 GB versus M5 Ultra 96 GB
These two Mac Studio configurations are unusually close in price but optimized for different things.
The M5 Max configuration at $5,099 provides 128 GB of unified memory, an 18-core CPU, a 40-core GPU, and 614 GB/s of memory bandwidth. Its advantage is capacity: that extra 32 GB can matter for larger models, multiple resident models, or workloads that simply exceed the practical limits of 96 GB.
The base M5 Ultra at $5,499 provides 96 GB of unified memory, a 30-core CPU, a 64-core GPU, and 1.2 TB/s of memory bandwidth. For Qwen3.8-27B, 96 GB is already far beyond the memory required for a Q4 model plus a long KV cache, so the Ultra's additional bandwidth and GPU throughput are likely to be more noticeable during everyday agentic coding than the Max's extra capacity.
Both machines benefit from unified memory: the CPU and GPU share the same pool instead of requiring separate copies of the model. That does not mean the entire memory pool is available to the LLM—macOS, applications, and inference buffers use it too—but both configurations leave much more working headroom than a 32 GB discrete GPU.
For serving, MLX-LM is the Apple-native path, while llama.cpp with Metal, Ollama, and LM Studio provide flexible alternatives. The trade-off is software compatibility: CUDA-only extensions, NVIDIA-specific kernels, and many research projects still target NVIDIA first.
RTX 5090: fastest when it fits
The RTX 5090 is the highest-throughput single GPU in this group for workloads that fit inside its 32 GB of VRAM. It combines roughly 1.79 TB/s of memory bandwidth with very high tensor throughput and the software ecosystem against which most local-AI tooling is still measured.
CUDA builds of llama.cpp are mature, while vLLM, TensorRT-LLM, PyTorch, FlashAttention, experimental kernels, ComfyUI, and most training workflows tend to support NVIDIA first. That matters for an agentic workstation because the same machine can move between inference, coding, image generation, and fine-tuning without changing ecosystems.
The limitation is capacity. Qwen3.8-27B at Q4 can fit with a full long-context cache, but 32 GB does not leave the same room for multiple agent slots, MTP/speculative models, vision components, or another GPU-heavy application. Full 262K context is therefore more of a tightly managed mode than it is on a 96-128 GB unified-memory system.
Pricing is the other major issue. The RTX 5090 launched at $1,999, but September 2026 retail tracking has commonly placed available cards around $4,500-$5,000, and unusually high third-party listings have reached roughly $7,000-$9,000+. At MSRP, its performance case is extremely strong. At several times MSRP, the comparison shifts toward complete systems with much larger memory pools.
DGX Spark: more CUDA memory, less bandwidth
DGX Spark approaches the problem from the opposite direction: 128 GB of coherent LPDDR5X memory, a Grace Blackwell GPU, a 20-core Arm CPU, and NVIDIA's software stack in a compact appliance.
Its strength is memory capacity and CUDA compatibility. Models that cannot fit on a 32 GB GeForce card become practical, and the system has room for larger contexts, multiple model components, and experimentation with much larger models.
Its main constraint is 273 GB/s of memory bandwidth. That is far below both the M5 Ultra and RTX 5090, so dense-model decode can be slower even though Spark has much more capacity. The Arm64 CPU is another consideration because some x86-only tools, binaries, or Python wheels may require alternate packages or containers.
For Qwen3.8-27B specifically, Spark is best understood as a capacity-first NVIDIA system rather than a raw-throughput competitor to the 5090.
Intel Arc Pro B70 and AMD Radeon AI PRO R9700
Intel's Arc Pro B70 and AMD's Radeon AI PRO R9700 are the two interesting 32 GB workstation alternatives.
The B70 provides 32 GB ECC GDDR6 and 608 GB/s, with inference paths through llama.cpp SYCL/Vulkan, OpenVINO, and Intel XPU tooling. The R9700 provides 32 GB GDDR6 and 640 GB/s, with ROCm/HIP on Linux and Vulkan as a portable fallback.
Both have enough VRAM for Qwen3.8-27B at useful quantizations and long context sizes, but neither has NVIDIA's breadth of copy-paste CUDA recipes. Linux is the most attractive environment for both—especially for ROCm on AMD—while Windows support depends more heavily on the exact backend and framework.
These cards become especially interesting when their complete-system cost is well below an RTX 5090 or high-memory Mac Studio.
Software compatibility at a glance
| Platform | Best local serving paths | Agent and development compatibility | Main friction |
|---|---|---|---|
| M5 Max / M5 Ultra | MLX-LM, llama.cpp Metal, Ollama, LM Studio | Excellent macOS development environment; easy OpenAI-compatible serving | No CUDA; vLLM and CUDA-only extensions are not native paths |
| RTX 5090 | llama.cpp CUDA, vLLM, TensorRT-LLM, Ollama, LM Studio | Broadest support for PyTorch, agent research, ComfyUI, and new kernels | 32 GB ceiling, power, price |
| DGX Spark | NVIDIA containers, CUDA, vLLM, TensorRT-LLM, llama.cpp | Strong supported AI stack and large coherent memory | Arm64 package gaps; modest bandwidth |
| Arc Pro B70 | llama.cpp SYCL/Vulkan, OpenVINO, XPU tooling | Strongest as a Linux inference workstation | More backend-specific tuning than CUDA or MLX |
| Radeon R9700 | ROCm/HIP, llama.cpp Vulkan, selected vLLM paths | Strong Linux AI path with a portable Vulkan fallback | Windows support remains more fragmented |
An OpenAI-compatible HTTP endpoint is the useful boundary. Once the model server exposes one, the coding agent does not need to know whether the model lives on MLX, CUDA, SYCL, or ROCm. The backend still determines performance and reliability, but the client can remain portable.
Sources
- Apple M5 Mac Studio announcement and starting prices
- RTX 5090 September 2026 market-price tracking
- RTX 5090 recent retail stock history
- Apple Mac Studio configurations and US pricing
- Apple Mac Studio technical specifications
- Apple MLX project
- NVIDIA GeForce RTX 5090 specifications
- NVIDIA DGX Spark specifications
- NVIDIA DGX Spark system overview
- Intel Arc Pro B-series specifications
- AMD ROCm Linux system requirements
- AMD ROCm Windows support matrix