AIAI Tech Engine

AI Tech Engine Lab

Local AI benchmarks from measured hardware

DGX Spark · EVO-X3 · Mac · CUDA — same needles for tok/s and render_ok.

Latest benches

2026-09-04

Lab control plane: Mac Studio · Spark · EVO+4080

Open WebUI + status board on Mac Studio for Spark and EVO+4080. Bundled with the Flash-Next 4080 spill quality×cost result (1083s vs 5321s).

2026-09-04

Flash-Next visual compare @16k: Spark vs AMD vs AMD+4080

Three fixed prompts (game/arch/space), max_tokens 16k. Mean wall: Spark 247s, AMD 410s, AMD+4080 442s. Once the model fits, 4080 spill adds capacity—not speed.

2026-09-04

Visual compare pack v1: game · architecture · product · space × AMD / AMD+4080 / Spark

Twelve fixed-component prompts to compare Qwen3.8, Flash-Next, GLM-Flash, DeepSeek-Flash on amd / amd+4080 / spark with checkpoints and a /10 rubric.

2026-09-04

Visual compare pack v2: playable vs still + DeepSeek on AMD+4080

v1 auto-scores lied. v2 splits playable/still, adds negatives, and uses AMD+4080 only for models that do not fit AMD VRAM — DeepSeek ~101G.

2026-09-03

Flash-Next + RTX 4080 spill: Spark vs EVO-X3 quality×cost

Flash-Next 3-stage race with RTX 4080 Vulkan spill on EVO. Spark wall_sum 1083s vs EVO+4080 5321s. EVO wins arch3d first-pass; angry fails 12/12 on EVO.

2026-09-02

DGX Spark vs EVO-X3: Flash-Next same-prompt bench

DGX Spark (NVFP4/SGLang) vs EVO-X3 (ROCmFP4/Vulkan) on Flash-Next with the same Angry Birds and 3D architecture HTML prompts. Decode ~37 vs ~27.5 tok/s.

2026-09-02

DGX Spark vs EVO-X3: Quality×Cost 3-stage race

Quality×cost tabs: Flash-Next (Spark 1694s vs EVO 3374s), Qwen3.8-27B (800s vs 1370s), DeepSeek-V4-Flash Spark-only (1844s). EVO DeepSeek GGUF is incoherent.

2026-09-01

One DGX Spark box: Qwen3.8-27B · Flash-Next · DeepSeek-v4-Flash (REV-441)

Same REV-441 needle (~2.5k–82k) and Angry Birds HTML on one DGX Spark GB10 across Qwen3.8-27B-NVFP4, Flash-Next, and DeepSeek-v4-Flash. ~82k TTFT: Flash-Next 0.50s · Qwen27B 4.1s · DeepSeek 6.8s. All needles matched.

2026-08-31

DGX Spark vs M1 Max vs M2 Max: same REV-441 needle, same Angry Birds HTML | AI Tech Engine

Same REV-441 needle (~2.5k–82k) and Angry Birds-style HTML on DGX Spark Qwen3.8-27B-NVFP4, M1 Max (prior post), and M2 Max Ornith/Qwen. Spark 82k TTFT 4.1s. All needles matched. GIFs included.

2026-08-31

Qwen3.8-Flash-Next memory tiers: what fits in 16GB / 64·96GB / planned 128GB | AI Tech Engine

Reading the 2026-08-26 Qwen3.8-Flash-Next (125B-A6B + 51B n-gram) from Unsloth GGUF, community MLX file sizes, and the official card only. Mapped onto the AI Tech Engine lab’s RTX 4080 16GB / M1 Max 64GB / M2 Max 96GB / planned 128GB. No unmeasured tok/s.

2026-08-30

M1 Max 64GB timings: Qwen3.8-27B (MTPLX · omlx) vs Ornith-1.5-35B | AI Tech Engine

On the same MacBook Pro (M1 Max 64GB) we run Qwen3.8-27B through MTPLX and omlx DFlash, and Ornith-1.5-35B-A3B through omlx, on the same needle and the same tokens. All three paths hit the needle. The same Angry-Birds-style HTML is playable only from Ornith. We do not use other people’s card benches.

2026-08-30

The 2026 local AI lab: 16GB CUDA vs 64/96GB Mac vs planned 128GB boxes | AI Tech Engine

How the AI Tech Engine lab is actually wired. What the RTX 4080 16GB, M1 Max 64GB, and M2 Max 96GB do today, and why the planned Strix Halo and DGX Spark 128GB boxes are for 70B-class work and local agents. No unmeasured TPS.

2026-08-30

Qwen3.8-27B and August 2026 open weights: what this lab can load today | AI Tech Engine

We check Qwen3.8-27B and Meta Muse Glimmer 30B against Hugging Face and official docs, then map them onto the AI Tech Engine lab’s M1 Max 64GB / M2 Max 96GB / RTX 4080 16GB and the planned 128GB boxes. No unmeasured tok/s.

2026-02-23

ChatGPT Plus vs Claude Pro vs DeepSeek API: squeezing the most work out of ~₩30,000 a month | AI Tech Engine

A 2026 subscription-cost guide for developers and operators. We compare code generation, long-document analysis, and output per dollar — not vendor slogans.

2026-02-23

Is making money with AI actually easy? The sour truth behind those ‘₩10 million a month side hustle’ videos | AI Tech Engine

YouTube and social feeds are packed with AI side-hustle receipts. We unpack the ‘anyone, 10 minutes a day’ myth, and the open-chat / paid-course funnel sitting under the video.

2026-02-23

2026 local-LLM PCs that do not waste the budget: how to pick a GPU | AI Tech Engine

RTX 4090, 4080 Super, used dual 3090s, Mac Studio M3 Max. What actually fits at each VRAM size, and which builds make sense at each budget — a buying map, not a lab fleet list.

2026-02-23

A $0/month serverless AI media stack (Astro + Cloudflare Pages) | AI Tech Engine

Skip WordPress and a fat Vercel bill. Pair a static site generator with Cloudflare Pages and serve a tech blog from 300+ edge cities in well under a second.

2026-02-23

Why ~80% of enterprise AI automation projects fail: the gap between a PoC and production | AI Tech Engine

Agents and pipelines that look perfect in a demo clip fall apart the day they hit production. The failure modes, and the guards that actually belong in the stack.