[ICML ’26 Effort] Efficient Qwen: Making Qwen3.5-4B Faster on a Single A10G
My ICML 2026 AdaptFM effort to optimize Qwen3.5-4B for low-latency inference on a single NVIDIA A10G through 100+ runtime, compiler, and quantization experiments.
June 30, 2026