Physical Address
London, UK
Physical Address
London, UK

Anyone who has sat still on the M25 knows the feeling. A lane is added at great expense to ease the congestion, and before long the traffic is as bad as it was. The economists Gilles Duranton and Matthew Turner studied road building in US cities and found that traffic rose roughly in proportion to the capacity added. More capacity did not buy less delay. It quietly invited more of it.
Testing functions can behave in much the same way. When delivery is under pressure, the instinct is to add hands. Double the QA headcount, clear the backlog, release faster. Yet six months on, deployment frequency looks much as it did, the backlog is no smaller, and the coordination burden has grown. The extra capacity went somewhere. It just did not turn into faster, safer releases.
So here is the takeaway, stated plainly. Beyond a fairly modest size, the evidence suggests that adding people tends to lower the output of each person, and that the added capacity often buys little extra speed. The bottleneck is rarely the number of people. It is coordination, communication, and clarity of purpose. If you are buying headcount in the hope of buying confidence to release, you are usually paying for the first and receiving less of the second.
Testing looks divisible. Split the cases, add the people, finish sooner. It is the same seductive logic behind the assembly line fallacy covered earlier in this series: the belief that skilled cognitive work can be chopped into parcels and staffed at will.
The problem is that people are not interchangeable units, and the work does not partition cleanly. Fred Brooks named this in 1975 in The Mythical Man-Month, and half a century later the observation still holds. Adding people to late software work tends to make it later. The reason is arithmetic. Every person you add opens a potential new line of communication with everyone already there, and the number of those lines grows by the formula n(n-1)/2. A team of five carries ten potential communication channels. Double it to ten and you do not double the channels, you more than quadruple them, to forty-five.
Each channel is a standing tax: a clarification, a handover, a status meeting, a requirement read two different ways. Add enough people and the time spent coordinating the work starts to eclipse the time spent doing it. Amazon’s answer to the same problem was the two-pizza team, the rule, associated with Jeff Bezos and described in Bryar and Carr’s Working Backwards, that if two pizzas will not feed the group, the group is too large.
The evidence points the same way. Quantitative Software Management (QSM) analysed completed projects in its database and found that, for medium-sized systems, teams of around three to seven people were the optimum. In a later comparison, large teams bought little or no faster delivery, yet cost three to four times as much and delivered two to three times as many defects. A large-scale study of open-source projects by Scholtes and colleagues found that output per contributor tends to fall as team size climbs. They gave the pattern an old name, the Ringelmann effect.
That effect is older than software itself. The French agricultural engineer Maximilien Ringelmann had people pull on a rope, alone and then in groups, and measured the force, publishing the results in 1913. Eight people did not pull eight times as hard as one. They pulled roughly four times as hard. Per person, effort had about halved as the group grew.
Later psychologists separated two forces behind it. Coordination loss, because aligning effort is genuinely hard. And motivation loss, because individual contribution becomes harder to see. When accountability blurs, people ease off without meaning to, quietly assuming the collective will carry it.
A caution before applying this to testing. Almost all of this evidence comes from software development teams, open-source projects and laboratory tasks, not from testing functions, so I am applying it by inference. It is also not unanimous. One study of open-source projects found the opposite, with output growing faster than headcount (Sornette and colleagues, 2014), and the authors of the two studies have since debated why their results differ. Volunteers who choose their own tasks are also a long way from a managed test function. Read the pattern as a strong warning, not a law.
With that caveat in mind, a large testing team is fertile ground for exactly this. The more testers assigned to an area, the less any single one of them feels they own its risk. Coverage is assumed to be someone else’s remit. Gaps hide in the overlap. This is not a failing of the individuals involved. It is a predictable property of how the work has been structured, and structure is a leadership decision rather than a personal one.
We have made the point elsewhere in this series that you cannot hire your way out of quality debt. Scaling a testing team while blind to these dynamics simply adds a second bill on top of the first.
The new capacity does not arrive free. It arrives with onboarding measured in months, during which your most capable testers stop testing and start explaining. It arrives with more meetings, more handovers, and slower feedback loops. And it arrives with diluted ownership, so that risk intelligence, the very thing you were trying to buy, becomes thinner rather than richer.
Return to the measure that matters. In the first post in this series we drew the line between activity and insight: a wall of green tells you what was executed, not whether the business is safe. Headcount is the same category error, one level up. It is an input. It counts hands, not outcomes. Deployment frequency, change failure rate, and time to recover, the measures at the centre of a decade of DORA research, tell you whether you can ship quickly and safely. None of them improves automatically because the team got bigger.
There is a great deal of noise at present about AI making delivery faster, generating more code, more tests, more throughput, and about testing becoming the bottleneck that holds it all up. The reflex is to meet more with more. Buy capacity, staff up, keep pace.
The DORA research suggests that reflex deserves caution. Its 2024 report associated higher AI adoption with lower delivery throughput and stability, and concluded that improving one part of the process does not automatically improve delivery without the basics, such as small batches and robust testing. The 2025 report shows the picture moving. The throughput association reversed and is now positive, but AI adoption still has a negative relationship with delivery stability. DORA’s summary is that AI does not fix a team, it amplifies what is already there. Strong teams get better. Struggling teams find their problems intensified. More output has not automatically meant more confidence to release.
I would argue headcount works the same way. Add people to a system with unclear ownership and slow feedback, and you amplify that too. The lesson is not to panic-buy capacity. It is to build capability. A small team of skilled testers who understand the business risk is likely to out-deliver a large team executing scripts, and to do it with far less coordination drag.
This is the efficiency argument that leaders keep, rightly, asking for. Properly stated, efficiency is not the lowest cost per test or the most tests per tester. As the assembly line piece put it, that is optimising for efficiency while remaining blind to effectiveness. Real efficiency is the least effort required to reach sufficient confidence to release. Sometimes that means fewer, better people. Almost always it means clearer ownership and faster feedback rather than more hands.
The uncomfortable part is that an oversized testing team looks busy. Cases are executed, reports are filed, the function is visibly well staffed. The waste stays invisible because it takes the form of coordination and dilution, and neither ever appears as a line item.
So the question is not how many testers you have. It is this.
THE LEADERSHIP QUESTION
“How do we raise the capability of the team we already have, rather than the size of it?”
Put it to your test leadership directly: what would raise the floor of the current team enough to release with confidence next quarter? If the first answer is more headcount, press a little further. More headcount to do what, and how would we see it in deployment frequency and change failure rate? Push instead towards better skills, clearer ownership, faster feedback, and better testability in the architecture. More often than not, the honest answer is that headcount was never the constraint.
This is not an argument for understaffing, nor for treating testers as a cost to be trimmed. It is an argument against mistaking size for strength.
Scale the capability, not the crowd. Keep teams small enough that ownership is unambiguous and feedback is quick. Invest in the judgement that lets a tester see business risk rather than merely execute a script. Measure the outcomes that show whether you can release with confidence, and stop celebrating the inputs that only tell you how many people are in the room.
Because a bigger testing team is not automatically a safer one. Past a certain point, it is a slower and more expensive way to arrive at less. The organisations that win the delivery race are not the ones with the most testers. They are the ones whose testers are skilled enough that they do not need an army to feel safe.
Sources