Blog
Notes on cloud and AI cost management.
Posts about the FinOps metrics worth caring about, how AI spend lands on the same ledger as the cloud bill, what CloudQuell ships, and how teams move from reactive cost surprises to a real operating cadence.
Three frontier models in 72 hours: what the September 2026 launches cost
Anthropic, Google, and OpenAI all shipped between 1 and 3 September 2026 — Claude Fable 5.1 and Mythos 5.1, Gemini 3.8 Flash, and GPT-6 Astra. List prices at the top of the market have converged on $10 / $50; the real differences are now in cache rates, a scheduled Gemini increase, and rollout limits. Here are the numbers and what they change on your bill.
The hidden cost of not knowing who owns your cloud spend
A cloud cost you can see but can’t attribute is a cost nobody is accountable for reducing, explaining, or forecasting. Here’s what unowned cloud spend actually costs: slower anomaly investigations, weaker forecasts, and optimization recommendations that never get implemented, along with practical ways to establish ownership across the bill.
What Is OpenAI Codex Persistent Mode? Everything We Know So Far
OpenAI is reportedly testing a “Persistent mode” for Codex — one that keeps the agent working until you put it to sleep and lets it create its own follow-up tasks. Here’s what the code actually shows, what OpenAI has confirmed, and what’s still unknown.
AWS Cost Explorer vs. FinOps Platforms: When Do You Need More?
AWS Cost Explorer is a strong, free, AWS-native tool for analyzing where your cloud spend went. A dedicated FinOps platform earns its place when the problem shifts to ownership, allocation, governance, and multi-account, multi-cloud, and AI cost visibility across teams. Here’s how to tell which one you actually need.
What Is Cursor Origin? Cursor’s New GitHub Alternative Explained
Cursor shipped Origin, an agent-native Git forge, in early beta on August 17, 2026. Here’s what it actually does today, how it compares to GitHub, what it costs, and what’s still coming.
Why your AWS bill keeps growing even when traffic doesn’t
Many AWS costs are decoupled from request volume — they accrue on provisioned capacity, stored bytes, and background data movement, not on how many users hit your app. Here are the specific line items that climb while traffic stays flat, and where to look first.
AWS cost allocation tags: best practices that survive contact with a real bill
The listicles say tag everything. Real bills are never fully tagged. The practice that survives: measure coverage, close the gap with rules, and stop chasing 100%. Activation lag, case-sensitivity, the untaggable services, and what good coverage looks like by org size.
Chargeback vs showback: pick by accounting reality, not maturity
The SERP is full of posts that rank the two on a maturity ladder. The FinOps Foundation says neither is more mature. Both fail the same way — until the bill sums to teams. The real decision is your accounting model, not your FinOps score.
The FOCUS spec, explained: one billing format for every cloud
FOCUS is the FinOps Foundation’s open spec that gives AWS, Azure, and GCP billing data the same column names and meanings. What it standardizes, the current version, who’s adopted it, and what it does — and doesn’t — change about multi-cloud cost work.
How to track OpenAI API costs (per project, per model, per feature)
The OpenAI dashboard tells you what you spent, not why. Three levels of cost attribution, the DIY playbook in full, and the point where a tool earns its keep.
Anthropic API cost tracking: workspaces, token types, and cache math
Claude bills five token types, and the cache math can flip which model is cheapest. Workspace attribution, the cost API, and where a tool earns its keep.
LLM cost management tools, compared: observability vs APM vs FinOps
Three kinds of tool track LLM spend, and they aren't substitutes: observability, APM add-ons, FinOps platforms. Which fits your team, where each is weak.
Set a budget on your AI spend before the invoice does
A retry loop is a 100× hour, not a 20% month — a monthly budget catches it too late. Why provider caps are blunt, and what per-model alerting does instead.
OpenAI vs Anthropic API pricing
OpenAI vs Anthropic API pricing, tier by tier — with the two things that actually decide your bill and that the sticker price hides: tokenizer skew and the cache-write premium. At least once this year, the cheaper option flips.
LLM API pricing comparison (all providers)
A cross-provider LLM API price comparison — flagship, workhorse, and reasoning tiers side by side, in USD per 1M tokens — plus the caveats (preview pricing, off-peak vs peak, long-context surcharges, third-party hosting, tokenizer skew) that reorder the ranking once you read the fine print.
Cheapest LLM API per million tokens
The cheapest LLM API is not one model — it depends on your input:output ratio. Here is the budget tier ranked by blended price (updated for Qwen-Flash's September 2026 price cut, which pulls a current-generation model level with the cheapest legacy one), a worked example at 10M+2M tokens/month, and the fine print that decides the real winner.
Connect Claude, Cursor, or any AI agent to your AWS cost data over MCP
CloudQuell runs a hosted MCP server that exposes your cost data as 43 tools over OAuth. Point Claude, Claude Code, or Cursor at it and ask your AWS bill questions in plain language — no API keys to manage, no CSV exports, the same plan entitlements as the portal.
Most AWS cost anomaly alerts are noise. Alert on the ones that matter.
A single global threshold on total spend fires on every weekday-versus-weekend wiggle and teaches your team to ignore it. Per-service baselines, first-time-spend detection, and one-click scoped alert rules are how you get signal instead.
Put an owner on every AWS dollar: tag coverage, cost centers, and allocation rules
Chargeback fails when the bill does not sum to teams. The fix is a coverage metric that tells you how much is unallocated, allocation rules for shared and untagged cost, and a cost-center hierarchy that rolls up to the whole bill.
Why we don't charge a percentage of your cloud spend
A vendor billing 2–3% of your cloud bill earns more when your bill grows — the opposite of what a cost tool should be paid to do. CloudQuell charges a flat monthly fee. Here is the incentive math, and what it costs at scale.
Introducing Effective Coverage %: the FinOps metric that combines coverage and utilization
Commitment utilization on its own lies. Coverage on its own lies. The honest answer requires combining them. Here is the metric, the formula, and how to read it.