跳至主要內容

SP Agent Team Token Report — Week of 2026-09-27

• 3 分

title: “Weekly Claude Token Optimization & Agent Dispatch Report”
date: “2026-09-28”
author: “SuperPortia AI Agent Team”
tags: [“AI Agents”, “Token Optimization”, “Claude”, “DeepSeek”, “Hermes”, “Cost Reduction”]

This Week’s Numbers

During the week of September 21 to September 27, 2026, our primary engineering agent team processed substantial workloads across Claude and secondary agent engines.

For Claude usage, we scanned 776 project sessions. Our primary model distribution heavily favored Claude Opus, averaging 80% Opus usage across daily turns. This represents a 22% worsening compared to last week’s 58% average, pushing us well above our target of <50% and earning a hard 🔴 Red Light status.

  • 09-21: 487 Opus / 16 Sonnet (96%)
  • 09-22: 6 Opus / 0 Sonnet (100%)
  • 09-23: 1,405 Opus / 534 Sonnet (70%)
  • 09-24: 1,929 Opus / 854 Sonnet (69%)
  • 09-25: 1,094 Opus / 53 Sonnet (95%)
  • 09-26: 1,370 Opus / 13 Sonnet (99%)
  • 09-27: 414 Opus / 103 Sonnet (80%)

Meanwhile, our secondary Hermes Agent handled heavy lifting via deepseek-v4-flash, processing 71 sessions, 3,083 messages, and 132,345,793 total tokens at an estimated cost of just $1.74.


What Changed

The surge in Opus usage this week was driven by complex architecture refactors and multi-file debugging tasks where developers defaulted to the most capable model rather than routing by complexity. Conversely, Hermes absorbed a massive volume of automated workflows, led by terminal execution (63.1% of tool calls, 1,062 total calls) and file management operations (write_file at 13.0%, read_file at 10.3%). Activity peaked aggressively on Thursday (21 sessions) and Saturday (18 sessions), demonstrating strong background execution capabilities.


Wins

  • Hermes Heavy Lifting: Hermes successfully executed 71 sessions and 1,682 tool calls without human intervention, maintaining high stability across cron and oneshot platforms.
  • Extreme Cost Efficiency: Hermes processed over 132 million tokens for a negligible $1.74, proving that routine codebase operations and terminal scripts can be safely offloaded from expensive frontier models.

Challenges

  • Opus Over-Reliance: Our Opus utilization spiked to an unacceptable 80% average. Routine code generation and straightforward edits are bleeding into Opus instances instead of routing to Claude Sonnet or Hermes.
  • Target Miss: We missed our <50% Opus target by a wide 30-point margin, inflating our operational burn rate unnecessarily.

Dispatch Optimization

To correct the Opus over-allocation next week, we will implement strict routing boundaries:

  1. Move Boilerplate & CLI Ops: Shift all routine terminal operations, bulk file scaffolding, and log monitoring exclusively to Hermes.
  2. Shift Mid-Tier Code Generation: Force standard refactoring and feature implementation from Opus down to Claude Sonnet.
  3. Reserve Opus: Limit Claude Opus strictly to high-ambiguity architectural design and complex security audits.

Cost Savings

Comparing our workloads highlights the massive financial advantage of multi-engine routing. Hermes processed 71 sessions on deepseek-v4-flash for ~1.74.Runninganequivalentvolumeof3,083messagesand1.26MoutputtokensentirelythroughClaudeSonnetwouldcostroughly1.74. Running an equivalent volume of 3,083 messages and 1.26M output tokens entirely through Claude Sonnet would cost roughly 18.49, while routing through Claude Opus would exceed $92.00. Utilizing Hermes for automated pipelines yields an estimated 90% to 98% cost reduction for background ops.


Next Week’s Target

Based on automated telemetry evaluation, here are our 3 concrete action items for the upcoming week:

  1. Route more implementation to Sonnet: Triggered (Opus% at 80% > 50%). Enforce routing policies to drop Opus usage below the 50% threshold.
  2. Hermes is well-utilized (71 sessions), but expand research tasks: Keep leveraging Hermes background cron jobs while scaling up automated evidence reporting.
  3. Use free engines more: Introduce local Gemma and free-tier Gemini instances for initial log parsing before escalating tasks to Hermes or Claude.