Scaled founder Simon Penson has spent two decades building, scaling and investing in B2B businesses – first as an agency founder who sold to IPG, and since as an investor, Chair and NED. Here, in his own words, is what he’s seeing from the companies automating hardest.

I sit on a small handful of boards across agencies, SaaS businesses and a couple of companies trying to become both at once. Through Haatch I have had a window into more than a hundred early stage B2B companies over the past decade, and before any of that I built an agency from a performance and content shop into a business of 300+ people that IPG acquired. I mention this not to polish my credentials but to explain the vantage point, because everything that follows comes from pattern recognition across a great many board packs, quarterly reviews and late-night conversations with founders wrestling with the same question at the same moment.

That question, in one form or another, is what remains for the humans if the models keep improving at their current rate, and it deserves a better answer than either the doom-laden headlines or the breathless vendor decks are currently offering.

Here is the strange thing I keep observing, both in my own business and among the most aggressive adopters across the portfolio. The companies automating hardest, running agents across coding, content, customer service, reporting and research, have not shrunk their teams into oblivion. If anything they have more genuinely important human work to do than they did two years ago, and that work commands a higher price. The shape of it has changed beyond recognition, but the volume and value have both moved upwards.

I have stopped treating this as a temporary quirk of the adoption curve and started treating it as a structural feature of the technology, and understanding why it happens is the single most useful strategic exercise a B2B leadership team can undertake this year. So let me walk through what I have seen, why the pattern holds, and what I would be doing about it in your chair.

The mood in the room

Before getting into the mechanics it is worth naming the emotional weather, because fear has become a strategic variable in its own right. The consensus among chief executives, investors and the AI laboratories themselves is bleak for knowledge work, with credible voices warning that half of all entry level white collar roles could disappear within a few years, and some of the sharpest investors in the world admitting publicly that it is the extraordinarily high-skill jobs being automated first rather than the junior ones. I now sit in board meetings where the unspoken question behind every agenda item is whether the business model itself has a shelf life measured in quarters rather than decades. That fear is rational as far as it goes, but fear makes for a terrible strategist, and the pattern on the ground tells a far more useful story than the headlines manage to.

What full automation actually looks like from the inside

I gave a keynote earlier this year called The Extinction Event, arguing that the convergence of software and services creates a binary outcome for agencies and professional services firms, and I stand behind every word of it. What the most automated businesses in my orbit have taught me since is something more nuanced about what adaptation actually means once the applause dies down and Monday morning arrives.

Inside these businesses, two distinct modes of working with AI have settled into place, and the difference between them matters enormously for anyone planning an operating model.

The first mode is the one everybody predicted, which is agents behaving as employees. These are agents living in Slack with names and defined jobs, drafting sales proposals, monitoring pipelines, triaging support tickets and compiling the first cut of the board report before a human has finished their coffee. In the strongest implementations I have seen, an embedded support agent will resolve well over a third of all customer conversations without a person ever touching them, freeing the customer service lead to spend the week improving the system rather than answering tickets. Any business not yet running some version of this is carrying cost that its competitors have already stripped out.

The second mode is stranger and considerably more important. It involves humans and agents working together inside the same workspace on complex, original problems: a founder and an agent going back and forth on a repositioning exercise, an engineer directing three coding agents at once, a consultant pressure testing a valuation model with an agent while the client is still on the call. This is collaboration rather than delegation, and the output from a skilled operator working this way sits on a different planet from anything either party produces alone.

What rarely makes it onto the conference slide is how astonishingly high maintenance both modes turn out to be. One business I know well built an automation to produce presentation decks, and by the time it worked reliably it required more than twenty distinct skills, a pile of supporting scripts and a meaningful compute bill for every deck produced. Agents drift and go stale without warning, needing someone to point them at the right problem, review whatever comes back, catch the confident errors and convert the output into an actual decision. The organisations that handed every employee a personal agent quietly rowed back within months to shared, team-level agents with named owners, because unowned agents decay into uselessness with remarkable speed.

Every agent, in other words, needs a human wrapped around it, framing the work at the start and judging it at the end, and the further an agent drifts from a person who owns its output, the worse it performs. I have not yet come across an exception, and I have been looking.

Why automating everything creates more expert work, not less

That is the first-order explanation for the paradox, but there is a deeper engine underneath it, and once seen it cannot be unseen. It runs as a loop with five stages, each of which I have watched play out across the portfolio over the past two years in almost exactly this sequence.

The loop begins with the models making yesterday’s competence cheap. They are trained on the accumulated residue of human work, the code, copy, decks and support tickets that competent people have produced over decades, and the result is that skills which were rare and expensive five years ago, such as writing a decent pull request or a serviceable strategy memo, are now available to almost anyone for pennies.

Cheap competence then gets adopted everywhere at extraordinary speed. Inside the automated businesses I see, the operations person now ships production code, the engineer drafts the product guide and the marketer builds their own dashboards, and output volume explodes across the whole organisation as a result. Some popular open source projects now receive more contributions in six weeks than the giants of the pre-AI era managed across an entire year.

Abundance, in turn, creates sameness. Because everyone is drawing on the same handful of models trained on broadly the same corpus, the default output converges on a recognisable house style, and buyers have developed a nose for it with impressive speed. The market has even coined a word for it, which is slop, and the important thing about slop is that it is not any single mistake but visible sameness repeated everywhere, the inevitable product of millions of people using the same tool without thinking especially hard.

Sameness then creates a premium on difference. Once competent-but-generic becomes effectively free, the only work that commands a price is work that is specific, alive and unmistakably built for this client in this market at this moment, and the standard for what counts as impressive ratchets upwards relentlessly. Capabilities that floored people in a demo six months ago now read as table stakes.

The final stage resolves the paradox, because difference can only come from expert human judgement. The models know a great deal about work that has already been done, but they know nothing, natively, about what needs doing right now in your specific situation, with your specific client, given everything you understand about the account that never made it into any training data. When the operations person ships code, somebody senior has to review it, and when everyone can produce a report, somebody has to know which report is worth producing. The judgement about what matters, what to build and what good looks like in this context remains stubbornly and structurally human.

I have started calling the people who thrive at this level the framers, the ones who decide which problem is worth solving and can tell whether it has been solved well. The models will climb whatever frame you hand them, and faster with every release, but they cannot replace the person doing the framing, because the moment a goal becomes explicit enough to hand over, someone with judgement had to make it explicit, and supplying that judgement turns out to be the enduring job.

I recognise that a version of this argument sounds like wishful thinking, since we have been promised before that the machines would do the work while we all became philosophers, so let me be precise. I am not suggesting that every job survives or that the transition will be gentle. What I am saying is that inside the businesses automating hardest, the binding constraint on growth has visibly shifted from execution capacity to judgement capacity, meaning the number of people who can look at a live situation, decide what matters, brief it into a system and hold the quality bar on whatever comes back. That capacity is scarce and expensive, demand for it is rising, and an organisation chart still built around execution as the bottleneck is staffed for a world that is ending.

What the benchmark charts are really measuring

The obvious objection to all of this is the benchmark data, which is genuinely alarming on first contact, with frontier models jumping from single-digit scores to near-human performance on graduate-level reasoning and real-world economic tasks inside twelve months. I have felt the same flutter of chart-induced panic as everyone else reading those reports over breakfast.

Having watched several portfolio businesses build their own internal benchmarks, however, I have learned to read the charts differently. Every benchmark measures a model operating inside a frame that a human designed, and the framing does a remarkable amount of the heavy lifting: change the wording of the task even slightly, or strip out the hints a thoughtful evaluator buried in the prompt, and the score collapses towards zero, while enriching the framing sends it soaring. The headline number measures the model plus an enormous quantity of smuggled human intelligence, because somebody decided what the task was, which constraints applied and what a good result would look like before the model produced a single token.

What happens when a benchmark saturates is telling as well, because the frontier does not vanish but moves up a level. Where yesterday’s test asked whether the model could execute a codebase rewrite, tomorrow’s will ask whether it can decide if a rewrite is warranted at all, scope it sensibly, manage the risk and judge the result, which is precisely the work senior people were doing all along. The finish line keeps being redrawn at the altitude where human judgement lives, and at every capability level reached so far, value has migrated towards the people directing the capability rather than evaporating from the system.

Seven moves I would make this quarter

Diagnosis without action is merely anxiety with a word count, so here is how I would map this onto a real business, whether you run an agency, a SaaS company or something between the two. This is, in essence, the conversation I am currently having in every boardroom I sit in.

1. Audit your delivery stack with brutal honesty

Work through every service line and recurring deliverable, then sort each one into one of two piles, the first containing genuinely irreplaceable judgement and the second containing repeatable process that a well-built agent could increasingly handle. Most leadership teams have never done this properly because the answer tends to be uncomfortable, yet no sensible automation decision can be made until you know what proportion of your revenue sits in each pile. Set aside a full day with your senior team, be ruthless about the classifications, and write the percentages down where the board will keep seeing them.

2. Automate the second pile aggressively, then redeploy the capacity

The trap I see everywhere is treating automation purely as a cost-out exercise, banking the margin and moving on. The businesses winning this transition take every hour freed up and push it deliberately up the value chain into deeper client strategy, business development or proprietary IP, and crucially they decide where that capacity goes before it evaporates into marginally lighter workloads. Write the redeployment plan at the same time as the automation plan, because one without the other amounts to only half a strategy.

3. Give every agent a named owner

Agents without a human accountable for their output degrade without exception, usually within weeks. If you are deploying them at any real scale, somebody senior needs to own the harness around them, meaning the review loops, the quality bar, the instruction files and the continuous tuning that keeps output sharp as models and business context shift underneath it. Budget for this as a genuine role with real hours attached rather than somebody’s enthusiasm project, because the maintenance load surprises every leadership team the first time around.

4. Retrain your best people as framers

The skills commanding a premium now are problem selection, scoping, quality judgement and taste, and remarkably few businesses train for any of them explicitly. Your strongest operators should spend the majority of their week deciding what is worth doing and reviewing whether it was done well, with agents handling the substantial middle between those two points. Build this into role descriptions, progression frameworks and hiring criteria now, before your competitors work out that the market for genuine framers is about to become very tight indeed.

5. Weaponise the slop problem commercially

Your prospects are drowning in competent-but-generic output and they know it, which makes specificity your sharpest available differentiator. Every proposal, piece of content and deliverable leaving your business should demonstrate unmistakably that it was made for this particular client by people who deeply understand their situation, because anything a competitor with the same tools could plausibly have generated is already invisible to the buyer. Audit your outbound material against that standard this month and quietly retire whatever fails.

6. Reprice around outcomes rather than hours

If agents are collapsing your delivery time while clients continue buying your time, you are handing the entire efficiency gain straight back to them without anyone consciously making that decision. The commercial model has to change alongside the delivery model, which in practice means packaging the judgement, productising the IP and pricing the outcome, piloted with one willing client before rolling it across the book. This single move is frequently the difference between trading at a services multiple and trading at something considerably more interesting.

7. Position for the coming demand explosion in your niche

When execution becomes cheap, people attempt things they would never previously have attempted, and most of those attempts go wrong in ways that require an expert to repair or, far better, to prevent. Whatever your field, there is about to be dramatically more of your kind of work in the world, much of it half-done and quietly on fire, and the sensible response is to position yourself now as the framer people call, supported by diagnostic tools that let you hold that conversation at scale rather than one anxious founder at a time.

The honest conclusion

I will not pretend the transition is painless, because roles are changing shape quickly and businesses reorganising around this technology will make hard calls along the way. After a couple of years watching the heaviest adopters at close range, though, I am confident about the direction of travel: the more a business automates, the more its humans matter, provided those humans move up a level rather than clinging to the layer being commoditised beneath them.

There is also an ownership dimension that boards should sit with. Businesses whose value is anchored in judgement, IP and productised expertise trade on fundamentally better multiples than businesses whose value walks out of the door at half past five, which makes every action above an enterprise value decision as much as an operational one, and the market is already pricing this divide more aggressively with every passing year.

The threat was never really the model but rather being the business still selling yesterday’s competence while the market reprices it towards zero, and the opportunity is being the one who decides what tomorrow’s competence should be pointed at. We built Scaled, and more recently ScaledOS, for exactly this moment, offering consultant-led sales, marketing, GTM and demand generation support alongside the diagnostic tooling to work out honestly where your business sits.

If any of the above landed a little too close to home, come and have that conversation, because it costs nothing to find out where you stand and a great deal to find out three years too late.