SP Agent Team Token Report — Week of 2026-08-16
• 1 分鐘閱讀 1 分
---
title: "Weekly Token Optimization: Slashing Opus Dependency"
date: "2026-08-17"
tags: ["LLMOps", "TokenOptimization", "AgenticWorkflows"]
---
## This Week's Numbers
The primary objective for this period was to reduce our reliance on Claude 3 Opus for routine agentic tasks, shifting the load toward Claude 3.5 Sonnet and our secondary Hermes Agent.
**Key Metric: Opus Utilization**
- **This Week Average:** 30%
- **Last Week Average:** 67%
- **Delta:** $\downarrow 37\%$
- **Target:** $<50\%$
- **Status:** 🟢 Green Light
The trend shows a significant aggressive descent in Opus usage, starting at 61% on Monday and plummeting to 0% by Saturday, indicating a successful shift in our dispatch logic.
## What Changed
We implemented a more stringent routing policy. Instead of defaulting to Opus for complex reasoning, we've transitioned the "Implementation Phase" of our workflows to Sonnet. The data shows that as the week progressed, Sonnet took over the bulk of the turns (peaking at 1,471 turns on Friday), while Opus was reserved only for high-level architectural steering.
## Wins
- **Target Achievement:** We successfully brought the Opus% well below the 50% threshold.
- **Hermes High-Volume Efficiency:** The Hermes Agent handled a staggering **87.6 million total tokens** across 43 sessions.
- **Tool Specialization:** Hermes demonstrated high proficiency in system-level tasks, with the `terminal` tool accounting for 50.7% of all tool calls (562 calls), effectively offloading "grunt work" from our primary Claude instances.
## Challenges
- **Session Bloat:** We observed a "marathon session" on August 9th lasting over 17 hours with 973 messages and nearly 1 million tokens. This suggests a potential loop or a failure in the agent's ability to summarize and reset context, leading to inefficient token consumption.
- **Dispatch Latency:** While the Opus% is down, the shift to Sonnet has increased the total turn count, which may impact overall completion time.
## Next Week's Target
- **Opus Baseline:** Maintain Opus usage at $\le 20\%$.
- **Context Management:** Implement a "Hard Reset" trigger for sessions exceeding 500 messages to prevent the token bloat seen in the August 9th session.
## Dispatch Optimization
Based on this week's tool usage, Hermes is clearly optimized for environment interaction. To further optimize, we will shift the following tasks from Claude to Hermes next week:
1. **File System Audits:** All `read_file` and `search_files` operations.
2. **Deployment Pipelines:** All `portal-deploy` and `github-monitoring` triggers.
3. **Initial Environment Setup:** All terminal-based dependency installations and shell configurations.
## Cost Savings
The economic impact of utilizing the Hermes Agent (powered by DeepSeek) is profound.
- **Hermes Actual Cost:** ~$1.09 for 87.6M total tokens.
- **Claude Sonnet 3.5 Equivalent:** Estimating a blended rate of ~$3.00 per million tokens for similar tasks, the same volume would have cost approximately **$262.80**.
- **Estimated Savings:** **~$261.71 per week** by offloading system-level tasks to Hermes.
## Recommendations
Based on our current performance data, the following actions are mandated for the next sprint:
1. **Use free engines more:** Current data shows zero utilization of Gemini/Gemma for initial research phases. We need to route pre-implementation research to free engines to further lower the cost floor.
2. **Review rate enforcement:** We have a high number of file edits (38) but no recorded `evaluate.sh` runs in the Hermes logs. Review rate is too low—check enforcement hooks to ensure all Hermes patches are validated.
3. **Optimize Long-Running Sessions:** The 17h session is an outlier. We must implement mandatory state-checkpointing to prevent single-session token spikes.