AI5 min read
Deep Tech
One routing pass to rule them all: 7x faster LLM decoding without the quality trade-off
CLSA uses one routing pass per sequence instead of one per layer, cutting KV-cache overhead while keeping token-level selectivity. Benchmarks show 17.1x throughput improvement at 128K context.
2026-07-26