Quota is an architecture constraint
Quota is an architecture constraint: five-hour and weekly windows force multi-harness routing so exhausting one provider never stops the work.
I do not run six coding harnesses because I cannot pick a favorite model. I run them because subscriptions have five-hour and weekly windows, and routing partly exists to load-balance across those windows so that exhausting one provider never stops the work.
That is not a model-quality decision. It is most of why a multi-harness setup exists in the first place.
People write routing charts as if the only question is “which model is best at this kind of task.” Capability matters. Claude models are still the ones I trust most for frontend planning and design. Codex still has the best computer-use and browser verification of the set. None of that helps when the window for that provider is closed and a Ready ticket is waiting on a human gate behind it.
Under a small shared AI budget, the constraint is not abstract. Personal and work subscriptions refill on different clocks. A heavy grooming pass on Opus 5 and a long implementation pass on GPT 5.6 Sol can burn different windows. If every phase of the pipeline insists on the same provider, one exhausted quota freezes the whole line. If phases can move, work continues: light tasks to GPT 5.6 Luna on high reasoning while the heavier window recovers, frontend planning on Claude while backend reasoning stays on Sol, review on whichever model did not plan the feature.
The last rule doubles as a reliability practice. Separating planner and reviewer is good epistemology. It is also good quota hygiene, because it keeps two different usage pools in play for the phases that cost the most attention.
I keep the routing table dated and short on purpose. These choices change every few weeks. What does not change is the frame: capability starts a route, licensing topology decides whether the route is even legal, and quota decides whether it is available this afternoon. When people ask why I bother adapting another harness, the honest answer is usually the third axis. A second surface for the same model is redundancy. A second provider with its own five-hour and weekly window is capacity.
If your agent setup only fails when the model is wrong, you are optimizing the easy axis. Mine fails when a window empties mid-pipeline. Routing around that is the job.