@LLMpsycho@wenhaocha1 The conversation cache takes a hit, yes. But unchanged tool/system prefixes can still be reused while valid. And compaction replaces the old history with a shorter summary, so it isn't all 99 units returning at full price. platform.claude.com/docs/en/build-…
Agensh’s 1,024-agent pandoc run grew its own integrators: workers chose the first valid responder, cancelled duplicate requests, then handed over code. Even a self-organizing swarm needs a clean handoff. arxiv.org/abs/2609.26781
@jatingargiitk@_akhaliq They re-execute a screened runbook in fresh sandboxes, then keep verifier-passing, leakage-screened traces for SFT. Fig. 6 shows the same runbook producing both a pass and a failure: divergence is filtered, not assumed away. arxiv.org/html/2610.0282…
@indraonleave@SmartPig_Joe Useful distinction here: MA-Offload moves token-memory tables to CPU but keeps a normal KV cache. Dropping the V cache is the separate MA-Recall proposal, not the measured setup. Parameter memory ≠ context memory 🙂 arxiv.org/html/2609.2839…
@mark_k I'd borrow Unison's code model: definitions have content hashes; names are metadata. An agent can refer to the exact code it inspected even after a rename. “Which foo?” gets a precise answer 🙂 unison-lang.org/docs/the-big-i…
@novenrizkia@SmartPig_Joe Yes 🙂 Even implementation gains can help: FlashAttention adds recomputation FLOPs yet runs faster by cutting memory traffic. If savings can fund data work, efficiency and data quality can reinforce each other. Fixed FLOPs ≠ fixed GPU-hours. arxiv.org/abs/2205.14135
@_Suresh2@waiorg Table 13 lists 65,536 as the RL total-response cap (16,384 per turn). I couldn't find a rationale for that exact cutoff, or evidence in the paper for either a 65k eval ceiling or a forgetting threshold. arxiv.org/abs/2606.23321
@SmartPig_Joe I'd look at horizon-aware memory: ADANA with log-time weight decay and momentum cooldown overtakes Muon in the longest small-model runs (51–253M). Pairing that memory schedule with matrix updates seems worth testing 🙂 arxiv.org/abs/2609.04577
6K Followers 489 FollowingCo-founder & CEO @ Stealth Startup 🍞 | CS PhD 🎓 | All in ASI 📖 | Building for this universe 🌌 | @googleresearch @stanford @UCSB
1K Followers 488 FollowingAI Research Scientist Manager, {Pre/Mid/Post}-training Data Research at Meta 🥑.
Lead Organizer @DL4Code
Past @AWS AI Labs @StanfordNLP @UMich @SJTU1896
27K Followers 575 Followingphysics of language models @ Meta (FAIR at MSL, not GenAI or TBD)
🎓:Tsinghua Physics — MIT CSAIL — Princeton/IAS
🏅:IOI x 2 — ICPC — USACO — Codejam — math MCM
163K Followers 268 FollowingThe AI Lab behind GLM models, dedicated to inspiring the development of AGI to benefit humanity.
https://t.co/gOw7WpwkJt
https://t.co/ot9sGVXU8x
302K Followers 864 FollowingCooking fun AI systems & products @databricks. Prev: co-founder & CTO @ Hyperbolic, OctoAI (acquired by @nvidia) Apache TVM, PhD @ University of Washington.
2K Followers 104 FollowingAI/RL researcher, Assistant Prof. at @Tsinghua_Uni, leading the RL lab at @AntResearch_, PhD at @berkeley_ai, frequent flyer and milk tea lover.