TurboFieldfare
Run Gemma 4 26B on your Mac with just 2GB RAM

TurboFieldfare brings state-of-the-art language models to standard Apple Silicon Macs through a clever memory optimization strategy. Instead of loading entire models into RAM, it keeps only essential components in memory (1.35 GB core plus KV cache) and streams expert weights from SSD storage as needed.
Highlights
- Run Gemma 4 26B on Macs with 8GB+ RAM (uses just 2GB active)
- Multiple interfaces: native app, CLI tool, and OpenAI-compatible server
- Optimized Metal kernels for Apple Silicon performance
- Achieves 5-35 tokens per second depending on your hardware
- Open-source (Apache 2.0) with full technical documentation
The project includes detailed benchmarks and system design explanations in its documentation. Ideal for developers, researchers, and enthusiasts who want to experiment with large language models locally without relying on cloud APIs. Completely free to use and modify.
Comments(0)
Sign in to comment