SP Agent Team Token Report — Week of 2026-09-20
• 1 分鐘閱讀 1 分
---
title: "Weekly Token Optimization: Addressing the Opus Spike"
date: "2026-09-21"
category: "Engineering"
tags: ["LLM-Ops", "Token-Optimization", "Agentic-Systems"]
---
## This Week's Numbers
Our token telemetry for the period of September 14th to September 20th shows a concerning trend in model distribution. While we've successfully scaled our secondary agent (Hermes), our primary dispatch logic has drifted.
- **Overall Status**: 🔴 Red Light
- **Weekly Opus Average**: 54% (Up from 23% last week)
- **Opus% Trend**: $\uparrow$ 31% worsening
- **Total Turns Added**: 5,779
- **Hermes Utilization**: 83 sessions / 105.4M tokens
## What Changed
The data indicates a massive spike in Opus dependency mid-week. On **09-17** and **09-18**, we saw Opus turns hit 1,279 and 1,141 respectively. This correlates with a heavy development push where the system defaulted to the most capable model for complex reasoning, abandoning the target of keeping Opus under 50%.
## Wins
The deployment of the **Hermes Agent** remains our biggest win. With 83 sessions and over 105 million tokens processed, Hermes is effectively offloading high-volume, low-reasoning tasks. Specifically, the heavy reliance on the `terminal` tool (65.6% of calls) and `write_file` (10.6%) proves that Hermes is successfully handling the "grunt work" of the agentic loop.
## Challenges
The primary challenge is **Model Drift**. We are seeing a "comfort bias" where the system routes implementation tasks to Opus rather than utilizing Sonnet's efficiency. With an Opus average of 54%, we are currently over-spending on reasoning tokens for tasks that should be handled by Sonnet or Hermes.
## Next Week's Target
- **Opus% Target**: $< 40\%$
- **Priority**: Tighten the dispatch gate to force implementation turns into Sonnet.
## Dispatch Optimization
Given that Hermes is already handling terminal-heavy workflows with high reliability, we will shift the following tasks from Claude to Hermes next week:
1. **Log Analysis**: Moving all `read_file` and `grep` operations for debugging to Hermes.
2. **Boilerplate Generation**: Shifting initial file scaffolding and `write_file` bursts to Hermes.
3. **Repetitive Testing**: Routing all `evaluate.sh` and harness triage loops to the DeepSeek-v4-flash engine.
## Cost Savings
The efficiency of the Hermes Agent is stark. Processing **105,482,714 tokens** across 83 sessions cost approximately **$1.31**.
If these same 105M tokens had been routed through Claude 3.5 Sonnet (averaging $\sim$\$3 per million input tokens), the cost would have been roughly **$315.00**. By utilizing the DeepSeek-v4-flash engine for auxiliary tasks, we have achieved an estimated saving of **$313.69** this week alone.
## Recommendations
Based on this week's telemetry, the following actions are mandatory for the next sprint:
1. **Route more implementation to Sonnet**: Our Opus% is at 54%, exceeding the 50% threshold. We must refine the prompt-routing logic to ensure implementation is not defaulting to Opus.
2. **Optimize Tool-Call Dispatch**: Since the terminal is the most used tool (974 calls), we should implement a "Hermes-first" policy for all terminal-based state checks.
3. **Audit Peak Hour Spikes**: Analyze the 1 PM peak (22 sessions) to see if these are automated cron jobs or manual developer spikes that can be optimized.