A research study evaluating ALTK-Evolve across eight language models on the AppWorld benchmark showed that memory dosage must be tailored to model tiers, with strong models benefiting from full guideline injection and weaker models performing best with curated retrieval.
Aug 18, 2026
13d agoKey Details
- Evaluated ALTK-Evolve across eight models from 30B dense models to frontier proprietary systems on the AppWorld benchmark (585 multi-step tasks).
- Strong models with headroom such as DeepSeek-V3.2 improved task completion by +9.5 percentage points with full self-mined guideline sets.
- Weaker models such as gpt-oss-120b achieved a +16.1 percentage point task completion gain with curated retrieval while adding only +5% token overhead.
- Saturated models like GLM-5 showed no measurable gain from added agentic memory.
- Prompt caching was identified as a key mechanism to keep full guideline sets affordable in production environments.