Fourth in a series. Follows the DeepSeek-V4-Flash single-3090 note. Living document; numbers update as the campaign runs. Last updated 2026-08-16. A revision history is at the bottom, and it is worth reading: this note has retracted one headline and reversed one verdict since rev 1.
Read entry →
We have been running low-bit quantization experiments on CPU only. It works, and it is slow enough that the feedback loop hurts — a context-depth sweep is most of an afternoon. The GPUs we already own are doing production work, so borrowing one was not an option. What we did have spare was an older dual-socket ML350 Gen9 tower with 512 GB of RAM, which is a good shape for MoE expert offload and had no GPU at all. So: a second-hand RTX 3090, blower-style, into an enterprise tower that was never designed for one.
Read entry →