Hardware Insights

  • Dec. 11, 2025 / Hardware Insights

    We Tested Devstral 2 (24B & 123B) — Here’s the Hardware You Actually Need

    Mistral AI has just released its new coding model, Devstral 2. We’ve been using its predecessor, Devstral Small, locally for code completion and have been very impressed with its performance. Early reports on Devstral 2 put it on par with other top models like Kimi K2 and Deepseek v3.2, so we were eager to get...

    devstral 2 llm hardware options gpus laptops mini pc
  • Dec. 9, 2025 / Hardware Insights

    Best Unified Memory Computers: Mini PCs, Laptops & Desktops for LLMs (2026)

    Unified memory has become one of the most important features for anyone running local LLMs in 2025. Instead of splitting memory between CPU RAM and GPU VRAM, unified architectures pool it into one high-bandwidth space that both the CPU and GPU can access. This matters because LLM inference is memory-bound long before it becomes compute-bound....

    computer models with unified memory for local llm
  • Nov. 17, 2025 / Hardware Insights

    Best Black Friday 2025 GPU Deals for Local LLM Users

    We’re tracking GPUs that make sense for LLM workloads and monitoring their prices now through Black Friday 2025, and we’re grouping them by VRAM since memory capacity determines which models and context lengths they can run, with bandwidth playing a major role in real-world throughput.

    llm capable gpus on discount on black friday
  • Nov. 11, 2025 / Hardware Insights

    Building a Multi-GPU LLM Workstation: Choosing the Right Motherboard for 6 – 10 GPUs

    If you want to run larger local models like Qwen3 235B A22B or GLM-4.6 355B fully in VRAM, you quickly run into the problem of scale. Even with 4-bit quantization, Qwen3 235B A22B is about 135 GB and GLM-4.6 355B is roughly 206 GB. On budget-tier GPUs such as RTX 3090 (24 GB VRAM), that...

  • Nov. 10, 2025 / Hardware Insights

    GPT-OSS 120B: Offloading MoE Layers to CPU Boosts RTX 3090 and 5090 Performance

    I’ve been testing the --n-cpu-moe flag in llama.cpp to see how much it improves performance with large Mixture of Experts models. The standard method of splitting layers between the GPU and CPU can be slow for these models. This flag offers a more targeted approach by moving just the expert layers to system RAM while...

    rtx 3090 and rtx 5090 stading on top of moe layers
  • Nov. 4, 2025 / Hardware Insights

    Running vLLM for Local LLMs on Mixed GPUs? MIG Might Just Make It Work.

    When I recently helped set up an LLM inference server for a client, I ran into a problem that may sound familiar to anyone mixing different GPUs. I had an RTX Pro 6000 Workstation (95 GB VRAM) and an RTX 5090 (32 GB VRAM). The goal was simple: run vLLM setup without wasting available memory....

  • Nov. 3, 2025 / Hardware Insights

    Inside PewDiePie’s $41,000 AI PC: 424GB of VRAM for Local LLMs

    When one of YouTube’s biggest creators decides to build a personal AI supercomputer, the local LLM scene takes notice. PewDiePie’s journey into AI hardware has produced a multi-GPU, 424GB VRAM workstation that many enthusiasts dream of. While his budget is far beyond the average builder, his component choices and setup offer a valuable blueprint for...

    PewDiePie’s custom open-frame AI PC build showing 10 GPUs installed on the left and NVIDIA System Management Interface on the right listing eight RTX 4090 48GB cards and two RTX 4000 Ada 20GB cards, totaling 424GB of VRAM.
  • Nov. 2, 2025 / Hardware Insights

    The Definitive GPU Ranking for LLMs: Token Generation & Prompt Processing Performance

    At Hardware Corner, we set out to create a data-driven benchmark hierarchy for local LLM inference – focusing on the two workloads that define real-world performance: prompt processing and token generation. Using llama.cpp’s latest llama-bench on Ubuntu 24.04 with CUDA 12.8, we measured a wide range of GPUs across model sizes, context lengths, and quantization...

  • Oct. 24, 2025 / Hardware Insights

    Best PC Builds for Local LLMs: From 7B to 123B Models

    This guide presents several PC build options at different price points for enthusiasts looking to run large language models (LLMs) on their local machines. These are templates designed for performance and value in LLM inference. You can adjust them based on component availability and your specific budget. At the moment, RAM prices are unusually high,...

    a desktop pc with dual rtx 3090 GPUs connected with SLI