A new paper from Google introduces SKILL.state, a method designed to drastically reduce the token usage of AI agents during long interactions. Traditional agents maintain extensive conversation histories as part of their input, leading to escalating token counts and costs. SKILL.state proposes a shift from this history-based approach to a structured representation of the agent's current state and the latest observation.
During its reasoning process, the agent identifies and stores crucial information for future steps within this state representation, effectively discarding the lengthy conversation history. This allows the input size to remain relatively constant, even over extended sessions. Benchmarks using Gemini-3-Flash demonstrated that SKILL.state achieved 0.94 accuracy with only 65,000 tokens, a significant improvement over a stateful baseline that required 1.1 million tokens for comparable accuracy.
Google Paper Slashes Agent Token Use by 94% with State Tracking
17 stories · 4 sources
#deepseek #ai #benchmarksOther digests
- 2026-09-19 — Vals AI Aims to Set New Standard for Neutral AI Benchmarking
- 2026-09-18 — ChatGPT Co-Creator's New AI Model Promises Faster, Cheaper Software Intelligence
- 2026-09-17 — Bonsai 2 27B Model Achieves Near-Lossless Compression, Significantly Reducing Footprint
- 2026-09-16 — Open-Source AI Models Gain Ground, Challenging Proprietary Systems
- 2026-09-15 — Google's Gemini 3.8 Models Now Available with Enhanced Thinking Capabilities
- 2026-09-14 — New AI Models Enhance Speech Synthesis and Recognition, While Others Flood Social Media
- 2026-09-13 — AI Leaders Push Boundaries: OpenAI Solves Millennium Problem, Moonshot AI Eyes $2B Revenue
- 2026-09-12 — OpenAI Claims Millennium Prize Problem Solution Amidst AI Advancement
- 2026-09-11 — Moonshot AI Aims for $2 Billion Revenue with Kimi Models
- 2026-09-10 — OpenAI's Navier-Stokes Model Release Includes Formal Mathematical Proof
- 2026-09-09 — OpenAI Claims Millennium Prize Problem Solution Amidst Scrutiny
- 2026-09-08 — DeepMind Unveils AlphaGenome Atlas; Qwen Quantization Benchmarks Revealed
- 2026-09-07 — AI Models Show Judgmental Tones; Open-Source Tool Integrates Free AI Models
- 2026-09-06 — DeepSeek Coder 3.5 Achieves Top Benchmark Scores, Outperforming GPT-6 Astra
- 2026-09-05 — DeepSeek Coder 1.0 Challenges GPT-4 on Coding Benchmarks
- 2026-09-04 — OpenAI CEO Apologizes for GPT-6 Astra Access Issues; Corporate America Embraces Open-Source AI
- 2026-09-03 — AI Coding Assistants: Claude vs. OpenAI Under Scrutiny
- 2026-09-02 — Google Unveils Gemini 3.8 Flash, Emphasizing Enhanced Reasoning Capabilities
- 2026-09-01 — Anthropic Cuts Claude Costs, AfterQuery Achieves Rapid Unicorn Status
- 2026-08-30 — AI Agents Aim for $10 Profit; Memory Accuracy Benchmarks Revealed
- 2026-08-29 — Google Paper Slashes Agent Token Use by 94% with State Tracking
- 2026-08-28 — AI Cost Reduction Explores Human-LLM Interaction, Analytical Handbook Emerges
- 2026-08-27 — AI Models Tested on Recursive Self-Improvement Benchmarks
- 2026-08-26 — AI Boom Fuels Record Profits for World's Most Valuable Company
- 2026-08-25 — Community-Run AI Discord Launches, Aims for Transparent Moderation and SOTA Local Models
- 2026-08-24 — Gemini 3.7 Outperforms 3.6 Despite Similar Release, Users Debate AI Model Performance
- 2026-08-23 — AI Agents Consume Five Times More Tokens Than Humans
- 2026-08-22 — DeepMind Alumni's AI Agent Faraday Shows Edge in Research Replication
- 2026-08-21 — GTA 6 Developer Rockstar Reportedly Furious Over Leaks, Premiere Date Unchanged
- 2026-08-20 — OpenAI Competes with Anthropic for Business AI Users Amidst Model Release Volatility