SP Agent Team Token Report — Week of 2026-08-09
• 1 分鐘閱讀 1 分
---
title: "Weekly Token Optimization Report: Bridging the Gap to <50% Opus Usage"
date: "2026-08-10"
category: "Engineering"
tags: ["LLMOps", "Token-Optimization", "Claude", "DeepSeek"]
---
## This Week's Numbers
Our primary objective remains the reduction of **Opus%** (the ratio of Claude Opus turns to total turns) to under 50%.
This week, we saw a significant downward trend in high-cost model dependency. The weekly average for Opus usage dropped to **67%**, down from **94%** the previous week—a **27% improvement**. Despite this, we are currently at a **🟡 Yellow Light** status, as we remain 17 percentage points above our target.
**Claude Turn Distribution:**
- **Peak Opus Usage:** Aug 06 (99%)
- **Lowest Opus Usage:** Aug 03 (41%)
- **Total Turns Added:** 661 across 7 sessions.
## What Changed
The data reveals a volatile transition period. On August 5th and 6th, we hit peaks of 95% and 99% Opus usage, indicating a "complexity wall" where the agent fell back on the most powerful model to resolve blockers. However, the recovery on August 7th (80%) and the low baseline on August 3rd suggest that our routing logic is starting to stabilize.
## Wins
The standout success this week is the **Hermes Agent's** operational load. Hermes has effectively become our "utility player," handling the high-volume, repetitive tasks that would otherwise bloat our Claude budget.
- **Throughput:** 109 sessions and 3,211 messages.
- **Tooling Efficiency:** Hermes handled 1,592 tool calls, with the `terminal` (45.7%) and `patch` (12.3%) tools doing the heavy lifting.
- **Skill Deployment:** The `xia-report-evidence` skill was the most utilized, proving that Hermes is ideal for evidence gathering and reporting.
## Challenges
The primary challenge is the "Opus Spike." The jump to 99% usage on August 6th suggests that certain types of implementation tasks are still triggering Opus by default. We need to determine if these were truly "Opus-level" problems or if the dispatcher is being too conservative.
## Next Week's Target
**Target: Opus% < 50%**
To achieve this, we must shift the "Implementation" and "Verification" phases of the agentic loop from Opus to Sonnet or Hermes.
## Dispatch Optimization
Based on the tool usage data, we are over-utilizing Claude for environment interactions. We propose the following shifts for next week:
- **Terminal/File Ops $\rightarrow$ Hermes:** Move all `read_file`, `search_files`, and `terminal` commands entirely to Hermes.
- **Evidence Reporting $\rightarrow$ Hermes:** Route all `xia-report-evidence` triggers to the secondary agent.
- **Complex Refactoring $\rightarrow$ Sonnet:** Force a "Sonnet-first" approach for patching, using Opus only after two failed Sonnet attempts.
## Cost Savings
By leveraging Hermes (powered by DeepSeek-v4) for high-volume tasks, we've significantly reduced our burn rate.
Calculating based on the 3,211 messages handled by Hermes:
- **Hermes Cost:** $3,211 \text{ queries} \times \$0.003 = \mathbf{\$9.63}$
- **Estimated Sonnet Cost:** Assuming a conservative average of $\$0.02$ per complex query for similar tasks: $3,211 \times \$0.02 = \mathbf{\$64.22}$
- **Weekly Estimated Savings:** $\mathbf{\approx \$54.59}$ per 3k messages.
*Note: The actual savings are likely higher given the 193M tokens processed by Hermes, which would be prohibitively expensive on Claude.*
## Recommendations
Based on this week's telemetry, the following actions are mandatory for the next sprint:
1. **Route more implementation to Sonnet:** Since Opus% (67%) remains $> 50\%$, we must tighten the dispatcher constraints.
2. **Use free engines more:** Gemini-3.6-flash usage was negligible (only 2 sessions). We should route initial research and documentation scanning to free engines to further lower costs.
3. **Review rate enforcement:** With 661 turns added but low session counts, we need to verify if `evaluate.sh` is running frequently enough to catch regressions before they reach Opus.