200+ Google Workspace engagements delivered, plus measurable time, cost, and token savings from 10+ production systems. Every number backed by real delivery records or code.
| Technique | Savings | How it works | Applied to |
|---|---|---|---|
| Thinking Budget Control | 40-50% | Disabled thinking tokens on structured JSON calls | Job Agent, AI Orchestrator |
| Transcript Trimming | 8-25K tokens/call | Capped document reads at 20K chars, video transcripts at reasonable limits | AI Orchestrator |
| Model Downgrade | 97% cost | Flash produces identical quality at 1/4 price, 3x faster | AI Orchestrator, Job Agent |
| Cached Digest | 30-60s per request | Firestore-cached action digest, refreshed every 30 min | AI Chat Bot |
| Tiered Provider Fallback | $0 baseline, $0.05/day worst case | Five-layer chain: gpt-oss-120b:free → paid retry → llama:free → Gemini (daily-capped at 100) → keyword. Never one degraded provider away from a bill spike. | Crypto Agent, AI Orchestrator, Hermes |
| Cache & Dedup at Source | 5x calls cut, 50% prompt tokens | Same input × identical output × N calls = N-1 wasted dollars. Memoize analyses by content hash, dedup repeating tool outputs before the model sees them. Sentiment went 91 calls/hour to 24; digest prompts shed half their size. | Crypto Agent, AI Orchestrator, Hermes |
| Model Tier per Dispatch | Up to 50x per call | Haiku ($0.30/Mtok) for bulk grep, Sonnet ($3/Mtok) for routine code, Opus ($15/Mtok) only for spec and audit. Don't pay Opus rates to grep 50 files. | AI Orchestrator, Job Agent, Crypto Agent |