TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
TACO reduces OPT-13B optimizer state 174× versus AdamW8bit and enables 30B fine-tuning on one H100.
TACO is a ternary column-wise one-sparse optimizer that follows Muon's operator-norm steepest-descent view while remaining compatible with AdamW-pretrained models. On OPT-13B it cuts persistent optimizer state from 27.7 GB to 0.16 GB (174× versus AdamW8bit) and peak training memory from 80.6 GB to 27.5 GB, with comparable accuracy and runtime. The method enables full-parameter fine-tuning of 30–32B models on a single 80 GB H100 across several families and tasks.