Qwen3
3 published articles
AI Research
Memory Decoder: a 6.9B bolt-on memory makes a 410M model beat Pythia-12B
Memory Decoder at Scale splits long-term memory from reasoning in LLMs. A 6.9B pretrained memory lifted Pythia-410M past Pythia-12B on 17 benchmarks with 39% fewer parameters; 1.7B domain memories added 9+ points to Qwen3 Base at every scale tested.
2026-07-31
Competitive programming
NousCoder-14B just opened the coding RL black box that OpenAI and DeepMind keep locked
Nous Research drops NousCoder-14B, a 14B competitive programming model with a fully open RL pipeline. The 68% Codeforces solve rate is notable, but the real story is that anyone can now replicate the stack.
2026-07-16
Long-context LLMs
Jet-Long's bifocal attention just killed the fixed-scaling trade-off for long-context LLMs
Jet-Long adapts RoPE rescaling dynamically per sequence length, using a local window and a long-range window merged via inclusion-exclusion. On Qwen3 models up to 128K context, it beats existing zero-shot baselines by over 2 percentage points on RULER and achieves the lowest perplexity on PG-19, all while adding less than 4% generation overhead.
2026-07-12