What I've Actually Noticed
This isn't a free tier complaint. I'm paying for all four of these platforms. I'm not hitting usage walls and getting bumped to a weaker model. I'm a paying customer on each one, and what I'm describing is a quality shift I've observed over time on platforms I've been using long enough to know what they used to feel like. ChatGPT is where it's most pronounced for me. I used it extensively to build this site, and I know what it felt like when it was working at its best: direct, decisive, implementing rather than discussing. What I'm experiencing more recently is something different. More preamble. More restating of the question before answering it. More outlining of what it's about to do rather than just doing it. More hedging at the end. The output is longer. The usefulness per word has gone down. Gemini has been a similar story, and honestly a more frustrating one. The back and forth required to get a clean output has increased significantly. What should be a single exchange becomes three or four, not because the task is complex but because the first response over-explains, qualifies, and then asks for confirmation rather than committing to an answer. Copilot and Claude sit differently in my experience, but I'm not here to hand out free passes. The question I'm asking applies to the whole industry, and readers should draw their own conclusions based on their own use.Why This Matters Beyond Annoyance
Overexplaining feels like a minor frustration when you're an individual user. But when you understand how these platforms bill at scale, the pattern starts to look less like a quality control issue and more like a structural incentive. That's the part I want to dig into.
How AI Billing Actually Works at Scale
Most people using consumer AI subscriptions pay a flat monthly fee and don't think much beyond that. But at the enterprise level, where large organisations deploy these tools across hundreds or thousands of employees, the billing model is fundamentally different. And understanding it changes how you look at the overexplaining problem. Enterprise AI contracts are priced in two ways. The first is per seat, meaning a fixed monthly charge for each user who has access. ChatGPT Enterprise, for example, has no published price and is negotiated directly with OpenAI, with 2026 contracts averaging around $60 per user per month, a reported 150-seat minimum, and annual prepay, putting the realistic entry point near $108,000 a year. That's just to get through the door. The second billing model is per token. A token is roughly three quarters of a word. Every word your AI assistant generates, every sentence of preamble, every restatement of your question, every hedge at the end of a response, costs tokens. And at the enterprise API level, the flagship GPT-5.5 costs $5 per million input tokens and $30 per million output tokens. Now here's the question worth sitting with. If a model is trained or tuned to produce longer, more verbose responses, to outline what it's about to do before doing it, to add qualifications and summaries and follow-up offers at the end of every reply, it generates more output tokens. In a flat-fee consumer subscription, that costs the company more to run and delivers less value to the user. But in a token-billed enterprise context, every unnecessary word is a billable unit. As one analysis of enterprise AI pricing put it, the real enterprise cost is never just the per-seat rate. It is seats plus API plus coding-agent credits plus a model price that resets upward every quarter. What if the verbosity isn't a quality control failure? What if it's a feature of the billing model? I want to be careful here. I am not accusing anyone of anything. I don't have internal documents. I don't have proof of intent. What I have is a pattern of behaviour that is consistent with a financial incentive, and a question about whether anyone is looking at that alignment closely enough.The Mechanism: How It Could Work
Let's think through how this would actually function, if it were happening. You wouldn't need a deliberate decision to make the AI worse. You would just need the optimisation targets during training and fine-tuning to reward certain behaviours. If a model is trained on human feedback and human raters consistently score longer, more thorough-sounding responses as higher quality, the model learns to be verbose. That's not a conspiracy. That's a training artefact. But the question of who defines "thorough" and what the downstream billing consequences of thoroughness happen to be is worth asking. If the people defining quality metrics are working for companies whose enterprise revenue scales with token output, the incentive to define quality as verbosity is at least present. Whether it has been acted on, deliberately or otherwise, is something only the companies involved can answer honestly. The bloated response also has a second effect beyond token billing. It fills the context window faster. Every AI conversation has a limit to how much text it can hold in memory at once. A model that uses twice as many words to say the same thing fills that window in half the time. Once the window fills, older context drops out. The conversation degrades. The user starts a new session. In a usage-based system, more sessions mean more billing events.Is Anyone Watching?
This is the part of the question that concerns me most. Regulators are beginning to pay attention to AI, but the focus is largely on discrimination, safety, and content moderation. The EU AI Act goes into full effect in August 2026, requiring risk management systems, technical documentation, conformity assessments, and human oversight frameworks. In the US, states including Texas, New York, California, and Illinois are entering 2026 with new AI laws targeting transparency and consumer protection for AI systems. These are important. But none of them are looking at the question I'm raising, which is not about harmful content or biased decisions. It's about whether the quality of a paid AI product is being quietly calibrated in ways that serve the vendor's billing model rather than the user's actual needs. Consumer protection law has dealt with this pattern before in other industries. A printer manufacturer that ships firmware updates making third-party ink cartridges fail. A streaming service that throttles resolution to reduce server costs while charging the same subscription fee. These are not hypothetical cases. They resulted in investigations and settlements. Regulators are less interested in aspirational ethics statements and more focused on demonstrable controls and accountability. But "demonstrable controls" assumes someone is asking the right questions. Right now, I'm not sure anyone is asking this one.What Would Change My Mind
I want to be fair here because I think good editorial requires it. There are legitimate explanations for what I'm describing. Model updates change behaviour in ways that aren't always improvements, and companies don't always communicate those changes clearly. Safety and alignment fine-tuning can make models more cautious and hedging as an unintended side effect. The AI landscape has been moving so fast that quality inconsistency could genuinely be a product of chaos rather than strategy. If an AI company published transparent documentation of how their models are tuned for response length and verbosity, what the optimisation targets are, and how those targets relate to their billing models, that would go a long way. Not because transparency proves innocence, but because the absence of it in an industry billing at this scale is itself worth noticing.
The question I'm really asking: In an industry where output length is directly tied to revenue at scale, who is verifying that the AI you're paying for is optimised for your benefit rather than for the number of words it generates?
