SP Agent Team Token Report — Week of 2026-08-23
• 1 分鐘閱讀 1 分
---
title: "Weekly Token Optimization Report: Driving Down Opus Dependency"
date: "2026-08-24"
category: "Engineering"
tags: ["LLM", "Token-Optimization", "Agentic-Workflows", "Claude"]
---
## This Week's Numbers
Our primary goal this week was to reduce the reliance on Claude Opus for routine agentic tasks to optimize for both latency and cost.
- **Opus Usage Rate**: 24% (Weekly Average)
- **Trend**: $\downarrow$ 9% (down from 33% last week)
- **Total Turns Added**: 3,677
- **Hermes Agent Utilization**: 42 sessions / 718 messages
- **System Status**: 🟢 **Green Light** (Target: <50% Opus)
## What Changed
We observed a significant shift in how tasks are dispatched. Early in the week (08-18 to 08-20), Opus usage plummeted to as low as 0-2%, indicating that Sonnet 3.5 is handling the bulk of our standard operational loops. However, a spike occurred on 08-21, where Opus usage rose to 39%. This correlates with a period of high-complexity architectural changes where the agent required the deeper reasoning capabilities of Opus.
Simultaneously, the **Hermes Agent** (powered by `deepseek-v4-pro`) has become a stable workhorse, maintaining an 8-day active streak and processing over 15 million total tokens.
## Wins
1. **Efficiency Gain**: Reducing Opus usage by 9% weekly puts us well under our 50% ceiling, significantly lowering the cost per project turn.
2. **Hermes Stability**: Hermes handled 380 tool calls this week with high reliability, particularly in file system operations.
3. **High-Volume Processing**: We managed to add 3,677 turns across 6 sessions without hitting critical rate limits, thanks to better load distribution.
## Challenges
The 08-21 spike suggests that our "fallback" logic to Opus is still triggered frequently during complex debugging sessions. While Opus is necessary for high-reasoning tasks, the jump to 39% suggests we may be routing tasks to it that Sonnet 3.5 could handle if the prompt context were better optimized.
## Next Week's Target
- **Opus Ceiling**: Maintain $<30\%$ average.
- **Hermes Expansion**: Increase the variety of tools Hermes can autonomously handle.
## Dispatch Optimization
Based on the tool usage data, Hermes is currently dominating `terminal` (33.9%), `read_file` (30.0%), and `search_files` (23.4%) calls.
**Proposed Shift**: We will move all "Environmental Discovery" tasks (grepping logs, directory mapping, and initial dependency audits) entirely to Hermes. By offloading these "sensory" tasks, we save Claude's context window for high-level synthesis and code generation.
## Cost Savings
By utilizing Hermes (`deepseek-v4-pro`) for the high-volume, low-reasoning tasks, the cost delta is stark:
- **Hermes Actual Cost**: ~$0.75 for 718 messages.
- **Claude Sonnet Equivalent**: Estimating the same token volume (approx. 1.4M tokens) at current market rates, the cost would have been roughly ~$6.70.
- **Weekly Savings**: Approximately **$5.95** on the Hermes subset alone. While the nominal amount is small, scaling this across 100+ agents represents a massive reduction in operational overhead.
## Recommendations
1. **Optimize High-Reasoning Fallbacks**: Since Opus% is currently low (24%), we should now refine the triggers that escalate to Opus to ensure it is only used for complex logic, not just "large" files.
2. **Expand Hermes Toolset**: Hermes is performing well with CLI tools; we should move more research-heavy "search and find" tasks to the Hermes agent.
3. **Review Free Engine Integration**: Our current data shows zero usage of free engines (Gemini/Gemma) for research tasks. We should integrate these to further reduce the cost of initial project scanning.