The Last Hand on the File

Design has been displaced by machinery at least four times since 1886, and every time the sequence has run the same way. A body of skilled manual work gets absorbed, the trade that performed it disappears inside a generation, and the judgment that used to be buried inside that trade separates out and becomes a profession in its own right. I argue here that generative artificial intelligence is the fifth of these events, that so far it is behaving like the previous four, and that it differs in three respects that should worry us: it is much faster, it attacks specification rather than execution, and it touches nearly every occupation at once instead of one trade at a time.

The evidence comes from randomised field experiments run at Harvard Business School, MIT, Stanford, Princeton and Microsoft Research, from labour market data covering both freelance platforms and national payroll records, and from three decades of human-computer interaction research into typography, first impressions and accessibility. Three things fall out of it. The productivity gains land overwhelmingly on the least skilled, which destroys the market value of ordinary competence while leaving judgment untouched. The systems pull measurably toward the typical, which wins the first half second of a client’s attention and loses the next five years of a brand’s life. And the working posture that fits the evidence is neither refusal nor surrender, but a division of labour where the machine takes the volume and the practitioner keeps direction, verification and authorship.

At the end I set out the evidence that undermines my own case, because there is some and it is good.


1. The trade that disappeared

For about four centuries, someone else set the type.

The compositor stood at a case of metal sorts and built the text letter by letter, sliding lead strips between the lines to open the spacing, packing out each line by hand with brass and copper spaces until it filled the measure. Physical work, learned over years, and slow: fifteen hundred characters in an hour was decent going.

Then the Linotype arrived in 1886 and cast whole lines from a keyboard, and one operator started doing what had taken seven or eight men. In the sixties and seventies phototypesetting took the metal out altogether, so the craft became keystrokes and codes. And between 1984 and 1987 the personal computer, the page description language and the laser printer between them folded the entire remaining apparatus onto a desk. Within roughly a decade the compositor, the typesetter and the paste-up artist had stopped existing as jobs.

The design profession, meanwhile, grew. It kept growing through all of it. The number of people paid to decide what should be set, at what size, in what relation to everything else on the page, and to what commercial end, went up while the number of people paid to physically set it went to nearly zero.

That happened because the compositor’s job had two things welded together: the labour of making, and the judgment of specifying. Every displacement in this history did the same job on that weld. It took the labour, dissolved the trade built around it, and left the judgment standing there on its own, no longer protected by the manual skill that had justified the wage.

My argument is that generative artificial intelligence is the fifth of these events, that it is running the same pattern, and that the only real differences are speed and altitude. The Linotype took the setting of letters. The Macintosh took the making of pages. What is being taken now is the production of plausible design solutions, which is a level of the work that no machine has previously reached.

The distinction that actually matters

The public argument about all this is stuck on a bad question, namely whether designers will be replaced. It is a bad question because it treats design as one thing that will be affected one way, and the research says almost the opposite: these systems help and harm different tasks inside a single project, and they help and harm different practitioners doing the same task.

The line worth drawing is not between people and machines. It is between labour and judgment.

Labour is the making: layouts, components, variants, marks, mockups, copy, markup. Judgment is choosing between the things that could be made, justifying the choice against constraints nobody wrote down, and carrying the consequences when it goes wrong.

Generative systems have become very good at producing plausible artefacts for nothing. They are not good at judgment, and the reason is not that the models need another year. Judgment needs access to a particular situation and answerability for a particular outcome. A model has neither, and no amount of scale supplies them.

So the practitioner whose living came from producing competent artefacts is standing roughly where the compositor stood in 1886. The practitioner whose living comes from judgment is standing where the designer stood, and is about to be worth more rather than less, because testing judgment through production just became almost free.

That is what the title means. The point is not to have a brand built by a machine. It is to build your own capability with the machine’s help, and then author brands that it helped make and did not decide.

Where the analogy breaks

Three things about this displacement have no precedent, and I would rather name them myself than have a reader find them.

It is fast. The Linotype needed forty years to reshape the trade and desktop publishing needed about ten. This one produced measurable effects in the labour market inside eighteen months of becoming generally available [~][*].

It aims higher. Every previous machine took execution and left specification alone: it rendered your decision, it did not make one. These systems propose the layout rather than merely drawing it. This is the first displacement operating on the same layer of the work as the designer.

And it is wide. Earlier waves hit one trade. Estimates of task-level exposure suggest something like eight in ten United States workers have at least a tenth of their tasks affected [^].

So the question is not whether the pattern repeats exactly. It is which parts of the work sit on the labour side of that line, which sit on the judgment side, and whether the line is where we assume it is. I do not think it is, quite, and the rest of this is an attempt to locate it with evidence instead of intuition.


2. From scarce competence to free plausibility

What the old market was selling

Until very recently the economics of this profession rested on one simple scarcity. Making a competent visual artefact took a trained person, expensive tools and time. The training was long, the software cost money, and there was no other route to the output. Competence was scarce, so competence was priced.

A whole market structure sat on that. A large middle tier of studios and freelancers sold reliability: correct, conventional, on time. Clients paid because there was no cheaper way to get it. Distinction was something you bought at the top of the market. Everywhere else, the product was competence.

Every earlier wave of tooling took a bite out of this. Desktop publishing put typesetting on every desk. Stock libraries killed the routine commissioned shoot. Template platforms handed non-designers professionally built layouts. Each time a slice of the middle lost its premium, and each time the profession absorbed it by climbing into work that needed more judgment.

Where the gains land now

What is different this time is who benefits, and that is now measured rather than assumed.

The first large field study followed 5,179 customer support agents through the staggered rollout of a generative assistant and found productivity up about fourteen per cent, measured as issues resolved per hour [†]. The average is the boring part. Novices and low-skilled workers improved by roughly thirty-four per cent. Experienced, highly skilled workers barely moved [†]. The authors think the system is capturing what the best agents do and handing it to everyone else, which is transformative if you do not already know it and pointless if you do [†].

Writing shows the same shape: completion time down sharply, quality up, gains concentrated among the initially weaker writers, and the performance distribution squeezed [□]. So does software, both in a controlled programming task [△] and in the largest study we have, which pooled three randomised trials at Microsoft, Accenture and an anonymous Fortune 100 manufacturer across 4,867 developers and found completed tasks up 26.08 per cent, with the less experienced developers adopting more and gaining more [‡].

Set that against the history and it is the real departure. Earlier tools lifted the floor and the ceiling together. A designer with a Macintosh could attempt more ambitious work, not just faster work. This one lifts the floor hard and the ceiling hardly at all. It does not end the profession. It flattens the gap between an untrained operator and a competent professional in exactly the band where most commercial work lives, which is a different problem and in some ways a worse one.

The frontier has a coastline

Old tools failed predictably. A photocopier degrades an image the same way every time. A typesetting system that hyphenates badly hyphenates badly consistently. You learned the limits once and worked around them for the rest of your career.

These systems do not work like that, and the clearest demonstration is a preregistered experiment run with 758 consultants at Boston Consulting Group, about seven per cent of the firm’s individual contributors [¶]. People were randomly assigned to work with no assistance, with GPT-4, or with GPT-4 plus some prompting instruction. On eighteen realistic tasks inside the model’s competence, the assisted group completed 12.2 per cent more, 25.1 per cent faster, at meaningfully higher assessed quality. On one complex managerial task chosen to sit outside that competence, the assisted group was nineteen percentage points less likely to get it right than the people working alone [¶].

The authors called the shape of this a jagged technological frontier. Tasks that look equally hard to a human sit on opposite sides of it, and nothing on the outside tells you which side you are on [¶].

I think this is genuinely new as a condition of practice. You cannot learn these limits once, because they are not shaped like limits. They are shaped like a coastline. The only reliable instrument for detecting them is knowing the subject yourself, which is exactly what an inexperienced operator has not got.

There is a follow-up to that experiment which sharpens it considerably. Researchers went through the activity logs of the consultants who tried to check the model’s work on the outside-frontier task. When people pushed back, fact-checked, pressed it to reconsider, it did not admit uncertainty. It apologised, then restated its original position with more supporting material, using structured reasoning that made a wrong recommendation look analytically solid [☆]. The authors describe this as persuasion rather than information.

No previous design tool argued with you. This one does, and it is better at arguing than most of us.

Assistance that is too good to be safe

There is a second finding here with no equivalent anywhere in the earlier history, and it runs backwards from what you would expect.

A field experiment gave 181 professional recruiters forty-four applications each to assess, and randomly assigned them a system of roughly ninety-nine per cent, eighty-five per cent or seventy-five per cent accuracy, or nothing at all [§]. The people with the better assistance did worse. They spent less time per application, thought less independently, followed the recommendation with less deliberation, and did not improve over the course of the task. The people given visibly weaker help stayed awake, kept using their own judgment, and got better as they went [§]. The author’s model of this is unglamorous and convincing: as the help gets better, the rational reason to put in your own effort gets weaker, and you fall asleep at the wheel.

Design assistance is now visibly good. That is the problem. Obviously bad output invites scrutiny. Plausible, well-composed, confidently presented output invites a nod. The better these systems get at producing convincing interfaces and convincing typography, the less any of us will look closely, and the more our own capability quietly rusts from disuse.

The compositor replaced by a Linotype lost a job, which at least he noticed. The designer using a very good model risks losing the capability while keeping the job, which nobody notices until a client does.

The homogenisation problem

Here the break with precedent is sharpest, and for anyone doing brand work this is the single most important result in the current literature.

Writers were asked to produce short stories with or without access to model-generated ideas, and a separate panel judged the results. Access to the generated ideas made stories rate as more creative, better written and more enjoyable, with the biggest improvement among the writers who had been least creative to begin with. At the same time, the assisted stories were measurably more like each other than the unassisted ones were [#]. The authors call it a social dilemma. Every individual is better off using it. The pool of what gets made gets narrower.

The same compression turns up among elite professionals, whose solutions varied much less when assisted than the control group’s did [¶].

Nothing in the earlier history did this. The Linotype did not make books resemble one another. Desktop publishing produced a famously ugly decade, but it widened the range of what existed rather than narrowing it. Template platforms made real sameness, but only inside their own vocabulary and only for the clients who picked them.

This is different because it acts on the professional’s own process, at the moment of having ideas, invisibly, while raising the perceived quality of each individual result. For a discipline whose entire deliverable is being unlike other people, a tool that improves quality and reduces variance is not obviously a gift. If everyone in your market adopts it, the net effect on any one brand’s distinctiveness is negative even though every single artefact got better. That is a strange thing to have to say out loud, and I have not found a way around it that does not involve doing more work by hand.

What has already happened

Two datasets describe the transition while it is happening.

On a large freelance platform, after the first widely available generative models shipped, monthly job counts in the exposed writing occupations fell by about two per cent and monthly earnings by about 5.2 per cent, with similar drops for people offering design and image-editing services once the image models arrived [~].

At national scale, payroll records from the largest United States provider show a thirteen per cent relative fall in employment for twenty-two to twenty-five year olds in the most exposed occupations, while employment for older workers in the same occupations held or grew. The adjustment came through headcount rather than pay, and it concentrated where the technology automates rather than assists [*].

This is the historical pattern arriving on time. It lands first and hardest on codified, checkable, transferable execution, which describes entry-level production work exactly. It has not yet landed on accumulated situated judgment.

There is one result in this literature that cuts against everything I am arguing, and I deal with it in section nine rather than hoping you miss it.

Owning the tool is not a position

One more piece of context, because it prevents an obvious mistake. A survey of 1,993 respondents across 105 countries reports that eighty-eight per cent of organisations now use artificial intelligence regularly in at least one function, up from seventy-eight per cent the year before. About a third have started scaling it. Nearly two thirds have not. And thirty-nine per cent report any measurable effect on enterprise earnings at all [○].

The gap between eighty-eight and thirty-nine is the number to keep. Something used by nearly everyone cannot, by definition, be what distinguishes you. The same survey finds the organisations getting outsized value are the ones that redesigned how they work rather than the ones that bolted the technology onto what they already did [○].

The parallel for a small practice is exact. In 1990 owning a Macintosh was an advantage. By 1995 it was a requirement. Access to generative tooling has made that entire journey in under three years. Whatever advantage is left has to come from what you do that it does not.


3. Three tiers become two

Before 2022 you could describe this market, crudely, in three layers. A bottom competing on price with variable quality. A large middle competing on reliability, producing competent conventional work at conventional prices. A small top competing on judgment, reputation and specificity.

The middle was the centre of gravity. It employed most people, trained most juniors, and earned most of the money. Its product was competence, and competence was hard to get.

I think the bottom has largely been absorbed and the middle has lost its reason to exist as a distinct tier. When a client with no training can generate plausible, conventional, competent output at no cost, what they will pay for plausible, conventional, competent output falls toward that cost. This is not a judgment about the worth of the work. It is just substitution.

The top moves the other way, because it was never selling artefacts. It sells the removal of a client’s uncertainty: that this decision is right, that it will survive conditions nobody has thought of yet, and that a named human being will answer for it if it does not. That has become scarcer relative to demand, because the volume of things needing evaluation has multiplied while the number of people qualified to evaluate them has not.

Four reasons I do not think the technology takes the top tier.

You cannot trust an output without expertise, since the frontier is jagged [¶], and the outputs actively resist casual checking [☆]. So the evaluator has to be an expert already. Expertise does not disappear; it moves from making to adjudicating.

Better assistance measurably reduces the effort people put into oversight [§], and oversight is precisely the product. A tool that erodes the thing cannot replace the role that consists of it.

Output diversity contracts under assistance [#], and being unlike other brands is the deliverable. A tool that pulls toward the typical cannot produce the atypical unless a person supplies the deviation on purpose.

And it cannot be accountable. It cannot hold a position when a client leans on it, cannot answer for a launch that failed, cannot be in the room. A great deal of professional services is the sale of accountability, and that is a difference in kind rather than a gap that a better model closes.


4. Typography, from physical constraint to infinite default

When the constraints did the disciplining

Historical typography was kept honest by cost. Type existed in specific sizes because someone had cut and cast those sizes. A second family meant buying a second set of metal. Optical sizes were not a refinement but an unavoidable fact, since a face cut for six point simply was a different drawing from the same face cut for seventy-two, and no other arrangement was possible. Letterspacing meant physically inserting material. Everything cost money, took up space, and needed labour, so every decision got made deliberately by someone who could tell you why.

Digital typesetting dismantled all of that, in order. Sizes became continuous. Mixing families became free. Optical sizing became optional and then, for a long stretch, largely vanished, because one outline could just be scaled. Tracking became a numeric field with a default of zero.

Which produced something genuinely new: a discipline in which doing nothing produces an acceptable result. Every value has a default and the defaults are coherent enough to pass.

Generative systems are the end point of that road. They choose typefaces the way they choose everything, by weighting toward whatever is most common in the training data given the prompt. You know the house style already: a narrow band of very common neo-grotesques and geometric sans faces, tracking at zero, no optical treatment, weights chosen by availability rather than by hierarchy.

It is rarely bad. It is unowned. Nothing is being claimed. Every value is the most common value.

The part that is measurable

The best answer to a client who thinks all this is subjective comes from the largest recent study of digital reading, which normalised font size by perception and tested hundreds of participants. Between each person’s fastest and slowest typeface, reading speed varied by about thirty-five per cent, with no loss of comprehension [π].

The second result is the one I find more useful. The variation between individuals was large enough that the authors state plainly that one typeface does not fit all readers, and argue for individuated rather than universal recommendations [π]. And in the same programme of work, preference turned out not to predict effectiveness: people do not reliably prefer the faces that make them faster [π].

So typeface choice is a performance variable with an effect size big enough to matter to a business. A thirty-five per cent spread in reading rate across a content site or a documentation portal is a commercial outcome, not a matter of taste.

It also means that nobody’s preference is evidence. Not the client’s, not mine, and not the statistical preference of a model trained on what everyone else already used. Only measurement is evidence.

And because the optimum is individuated, the defensible sentence is never “this is the best typeface”. It is “this is the best typeface for this audience, in this context, in this reading mode, and here is why”. That sentence is judgment, and it cannot be generated, because it depends on facts about a specific readership that are in nobody’s training data.

What is left for the person

Professional typographic work is mostly departures from the default, each needing a reason that lives outside the file.

Metric spacing is right on average and wrong at display sizes, where counters open up and relationships that balanced at text size stop balancing. Fixing that is a decision made at one specific size by an eye that has done it before, not a global setting.

A hierarchy that resolves in two weights beats one that resolves in four. The extra weights show up in generated work because extra weights are common in the data, not because the hierarchy asked for them.

Then there is script coverage, which I care about more than most people because I work in Greek. You have to establish whether a family’s non-Latin was designed or interpolated, and the difference is visible immediately, and in Greek it is frequently bad. A system asked for a modern, clean pairing has no idea this constraint exists.

Licensing is a procurement decision with a recurring cost, defined permitted uses, web-serving restrictions and renewal risk. A recommendation that ignores the licence is not a recommendation, it is a liability you handed your client.

Rendering matters too: behaviour at small sizes on cheap screens, hinting quality, variable axis support, what happens during font loading.

And the family has to stay usable by the client’s own people after you leave. A face that needs an expert to look right will not look right in their hands, and you will get the blame.

None of that can be prompted, because none of it is a property of the artefact. All of it is a property of the situation.

The one asset that still resists

Typeface choice is a discrete pick from a very large but very skewed distribution, and assistance concentrates the pick on the crowded end of it [#]. If everybody uses assistance, brand typography converges. You will not see it in any one project. You will see it across a sector, about two years from now.

The answer is not abstinence. It is deliberate deviation, paid for with knowledge: independent foundries, historical revivals, regional traditions, and when the budget allows, letterforms drawn or customised for the job. A proprietary typeface is currently one of the very few brand assets that cannot be copied for nothing, precisely because it exists in no training set. Which is a strange full circle. As in the metal era, the most defensible typography is again the type somebody had to make.


5. Interfaces, and the tyranny of the first half second

How the review cycle collapsed

Before the web, a visual system was judged slowly. The client saw a printed comp, lived with it, came back in a few days. Iteration cost real money, so decisions got argued rather than sampled.

Digital work cut that to minutes, then to seconds. Generative tooling has cut it again, to the point where the person reviewing is looking at a fully rendered interface moments after describing it. The mode of evaluation has changed with it, and there is a lot of research about what happens in those moments.

An old curiosity that turned into a hazard

The finding that visual quality changes perceived usability starts with a study of cash machine layouts that found a strong correlation between rated beauty and rated apparent usability [◊]. It was extended experimentally with the demonstration that the correlation survives actual use: what people thought about usability afterwards tracked the aesthetics, while manipulating the real usability did not move perception the same way [∞]. Later work qualified the mechanism and the boundary conditions [Ω], but thirty years on the association has held.

We also know how fast the judgment forms. Visual complexity and prototypicality shift aesthetic ratings within fifty milliseconds, and detectably at exposures as short as seventeen [Ω]. Models of perceived complexity and colourfulness, built from 548 people rating 450 websites, explain roughly half the variance in appeal after half a second of exposure [≡].

For three decades this was a benign observation about perception. It has become a hazard, for one specific reason: generative systems are excellent at exactly the properties that drive those instant judgments. Moderate complexity, conventional colour, high prototypicality. They are good at them because they optimise toward the middle of the distribution, and prototypical is what the middle means.

So generated interfaces win the first half second and lose everything after it. They perform best in precisely the mode a client uses to review a mockup, and then the aesthetic-usability effect carries that good first impression forward into judgments about actual quality [∞]. If you are arguing for something better but less immediately familiar, you are arguing against a bias that operates faster than argument does. I have lost that argument in a meeting more than once, and now I know why.

Which parts should be boring

Being typical is not always a fault. Checkout flows, banking screens, government forms and clinical tools all benefit from matching expectation, and being creative there is closer to negligence.

The skill is knowing which parts must be typical and which must not. Navigation, form behaviour, error handling, confirmation before destructive actions: conventional. Brand expression, editorial voice, motion, density: decided. A system with a uniform pull toward the prototype cannot make that split, because making it requires knowing which elements carry the differentiation and which carry the risk, and that is strategy rather than visual sense.

The accessibility numbers, which are the clearest evidence of anything

The closest thing we have to a sector-scale natural experiment is the annual scan of the top million home pages run by a university disability research institute.

For several years the numbers improved slowly. In 2026 they went backwards. Detectable accessibility failures appeared on 95.9 per cent of home pages, up from 94.8, with an average of 56.1 errors per page, up 10.1 per cent from 51.0 [Σ]. Low-contrast text, the most common failure of all, hit 83.9 per cent of pages, up from 79.1, averaging thirty-four separate instances on each affected page. Five of the six commonest failure types got worse year on year, and those same six have topped the list for seven years running, accounting for about ninety-six per cent of everything detected [Σ].

The report puts the reversal down mainly to two things growing faster than anyone was fixing them: page complexity, with the average home page gaining 22.5 per cent more elements in a single year and now carrying about 1,437 against 782 in 2019, and ARIA usage, up twenty-seven per cent. The authors tie that growth to heavier use of third-party frameworks and libraries alongside automated and AI-assisted coding [Σ].

Look at which categories got worse. Contrast, form labels, empty links, empty buttons. Those are properties of interactive components built by development teams, not content typed by authors. They are exactly what generative tooling produces most readily and what an unqualified operator is least equipped to audit. More code, made faster, containing more of the same six mistakes that have been documented and fixable for seven years.

This one sells itself. Accessibility failure carries legal exposure in several jurisdictions and the sector trend is now measurably negative. Being able to demonstrate conformance is selling risk reduction, which is a far easier conversation than selling taste.

The second year

The last thing about interfaces is time, and it is what generated work handles worst.

Generation optimises for the artefact at the moment of generation. It does not optimise for month eighteen. Design systems fail slowly and for boring reasons. Tokens multiply because nobody enforced the taxonomy. Components fork because the original was awkward to extend. Spacing scales acquire exceptions until they are not scales. Documentation drifts away from the implementation. All of these are governance failures rather than visual ones, and governance needs a person who is still there and still answerable.

Which is why the claim worth making is no longer “I can build this interface”. It is “I can build this interface so that your team can still use it correctly in eighteen months”. The first is cheap now. The second is not.


6. The part of the job with actual users in it

Why nothing ever automated this

User experience is the youngest part of the practice and, not coincidentally, the part no earlier wave touched at all. A Linotype cannot interview a reader. Desktop publishing cannot watch someone fail to find a button. A template platform cannot discover that the warehouse team refuses to use anything that needs two hands.

The reason is structural. A model knows a great deal about users and has no access to users. It has read everything ever published on checkout abandonment and has never once watched a particular person give up on a particular checkout.

The constraints that decide whether a design works are usually local, undocumented and awkward. A regulator requires a disclosure that ruins the elegant flow. A phone field rejects the format most of the client’s customers actually use. The stakeholder with veto power is measured on something the brief never mentioned. The internal team cannot maintain the content model your design assumes.

This is the knowledge Polanyi called tacit [★], and it is the reason the productivity literature gives for why such work resisted automation for so long [†]. The experimental version is exact: the one task where assisted consultants did worse than unassisted ones was the task that required pulling together information sources that disagreed and interpreting them in context [¶], which is the shape of every real user experience problem there has ever been. It did not fail because it was inarticulate. Being articulate made it more convincing, not less.

Why “just check the output” is not a plan

The standard response is that all this is manageable because you can check the work. Two findings complicate that.

Checking gets resisted. The log analysis of more than seventy consultants trying to verify outputs on the outside-frontier task found the system escalating persuasion under challenge instead of admitting limits: apologising, appearing to correct itself, then restating the original position with more supporting structure [☆].

And checking decays. Even where verification is perfectly possible, the effort people put into it falls as the assistance gets better [§]. So verification is not merely hard. There is a steady behavioural pressure against doing it at all.

A related warning applies to the growing habit of using language models as stand-in research participants. Simulated respondents give you the distributional average, and the whole value of a usability session is the participant who does something nobody predicted. A system optimised for the predictable cannot hand you the unpredictable.

What is left is nearly all of the value

What remains is this: getting access to real users and real constraints, working out what you are looking at, deciding what to do, and answering for the decision. It is a small share of the hours on any project and effectively all of its worth.

It is also, conveniently, the area where assistance is safest to use, because drafting protocols, writing discussion guides, structuring synthesis and producing documentation are all low-variance tasks with an expert available to review every step.


7. How I think you should actually work

Five rules, each traceable to evidence rather than to my preferences.

Keep direction. Give away volume. Direction is deciding what to pursue; volume is producing instances. Ceding direction converges you toward the mean [#], which is fatal in work that exists to be different. Volume is where the measured gains are [†][□][‡].

In practice: you set the position, the constraints and the criteria. Assistance generates a large set of candidates inside them. You reject nearly all of it. The rejecting is the design work and it cannot be handed over. If you cannot say what the criteria are, you do not have criteria, and you are selecting by resemblance to whatever is common, which is the exact failure the research predicts.

That same consulting study noticed two working patterns among people who used it well, which the authors call centaurs and cyborgs [¶]. Centaurs split the task strategically, sending whole pieces to the human or the machine depending on which side of the frontier each falls. Cyborgs interleave much more finely, swapping back and forth inside a single subtask. Both are deliberate allocations by someone who understands the work. Neither looks anything like handing over a brief and accepting what comes back.

Give away the low-variance majority. A large share of professional hours goes on work where correctness is well defined and variation is unwanted: component scaffolding, repetitive styling, spec documents, migration scripts, alt text, changelogs, data transformation, first-pass audits, the fifteenth admin form. It sits comfortably inside the frontier and the cost of an error is low when you are reviewing.

Moving those hours into the scarce activities is the actual mechanism by which this stuff raises what you are worth. The gain is not speed. It is attention redirected toward judgment.

Verify outside the conversation, and assume you are getting lazy. Given the jagged frontier [¶], the arguing [☆] and the falling asleep [§], verification cannot happen in the chat. Asking whether it is confident is not verification. Getting an apology followed by a restatement is definitely not verification.

Verification is external: contrast ratios measured, not eyeballed. Keyboard traversal tested, not assumed. Licences read, not summarised. Performance profiled, not predicted. Claims about user behaviour checked against user behaviour. That takes knowing what to check, which is a competence you now have to maintain separately, and it takes a procedure rather than an intention, because intention is exactly what degrades as the tool improves [§].

Be the last hand on the file. Every element in what you deliver has to be defensible by you, in your own words. If you cannot explain why a decision is there, take it out rather than keep it because it looks fine.

That one rule is the whole difference between assisted work and generated work. It is also what preserves your standing to charge for judgment, since judgment is what the fee is for.

Keep the craft up on purpose. There is a warning buried in the productivity literature. If the technology lifts the floor for novices and leaves experts flat, the incentive to become an expert weakens at the exact moment expertise becomes the scarce input. The labour market consequence for entry-level workers is already visible [*], the risk that firms stop investing in developing people is flagged in the same literature [†], and the individual mechanism by which the erosion happens has been measured directly [§].

So it has to be deliberate. Grids, hierarchy, contrast, colour, rhythm, the mechanics of type, interface conventions: these now have to be practised on purpose, because the tooling will happily let them rot and will produce no symptom until a client finds one.

Used properly, the same technology is the best instrument available for keeping them sharp. It will argue against your layout, explain why something fails, teach you the part of the stack you have been avoiding, and produce historical references faster than any library. The difference between using it to learn faster and using it to skip learning is invisible in any single piece of work and decisive across twenty years.


8. What a brand is, if it is anything

All of this resolves into a claim about what a brand actually is, and the history makes it easier to see.

Until now you could not produce a brand system without judgment, because production required a person and the person had to decide. The two were welded together, exactly as making and specifying were welded together in the compositor’s trade. For the first time it is possible to obtain the artefacts without the reasoning.

A system made that way has no argument behind it. Nobody in the organisation can say why the mark is that shape, why the type is set that way, or what the system implies about a case it has not met yet. So extension happens by imitation instead of inference, and the whole thing falls apart the first time reality asks a question the training data did not anticipate.

A system made with judgment, whatever helped make it, has three things the first one lacks.

It extends, because the logic that produced it can produce more. Someone can tell you what it does for a sub-brand, a partnership lockup, a regulatory disclosure, a market with a different script.

It is defensible, because every decision was chosen against alternatives for stated reasons. That matters commercially at the moment a new stakeholder wants to change something. An undefended system accepts every change and is gone in two years.

And it is differentiated, because the deviation from the typical was deliberate. Given what we now know about homogenisation [#], that property takes active effort rather than arriving by itself.

What clients buy has never been the file. It is the removal of uncertainty and the transfer of responsibility. A model cannot take responsibility, cannot be in the room when the decision is challenged, and cannot answer for being wrong. That is not a limitation of this generation of systems. It is the commercial foundation of professional practice, and no amount of improvement touches it.


9. Where my argument is weak

The finding that contradicts me. The freelance platform study specifically tested whether being good protects you. It does not appear to. Looking at heterogeneity by employment history, the authors found no evidence that high-quality service, measured by past performance and prior work, softened the damage, and reported suggestive evidence that the top freelancers were hit harder [~].

That is a direct hit on the comfortable version of everything above and I am not going to wave it away. Three things soften it slightly, none decisively. It measures a short-run shock right after the models shipped, and short-run substitution need not persist. The platform is a context where buyers cannot really observe judgment before they buy, so that market may be structurally incapable of pricing the thing I claim is valuable. And platform quality scores measure marketplace reputation rather than situated professional judgment. Still, the honest summary is that the evidence that excellence protects you is weaker than the evidence that excellence is scarce.

The numbers move. Even the biggest pooled developer study carries a standard error of 10.3 per cent on a 26.08 per cent effect, and its authors say plainly that the individual experiments were noisy and the results varied between them [‡]. Tooling changes faster than studies can be finished. Treat every figure here as a fact about a period and a setting, not a constant.

None of this was done on designers. The consulting experiment used elite management consultants [¶]. The recruiter experiment used a predictive system, not a generative one [§]. The creativity experiment used short fiction [#]. The developer experiments used code, where correctness is far more objectively checkable than anything visual [‡][△]. There is no published field experiment of comparable rigour on professional visual design. Everything I have transferred to design practice is transferred by analogy, and you should read it that way.

The accessibility data is correlational. The 2026 reversal tracks with page complexity and ARIA growth, and the authors name automated and assisted coding among the plausible contributors [Σ]. It is not a controlled study and generated code is not established as the cause. I have leaned on it because it is the only sector-scale evidence in existence, not because it proves what I would like it to prove.

The adoption survey is self-reported. Respondents said those things about themselves; nobody audited them [○].

The aesthetics literature has old problems. Early studies did not manipulate aesthetics and usability independently, and the direction of the relationship is still argued about [Ω]. The narrow claim I use, that fast aesthetic judgments colour later evaluation, is well supported. Anything stronger is not.

And the history is an analogy, not a prophecy. The compositor is instructive. He is not predictive. Three features of this transition have no precedent: the speed, the fact that it operates on specification rather than execution, and the breadth of exposure [^]. A pattern that held four times may not hold a fifth, and I would rather say so than pretend the past is a guarantee.


10. Where that leaves us

Design has survived four displacements by giving up labour and keeping judgment. Each time the trade that did the absorbed work vanished, and each time the profession that supplied the judgment got bigger.

The fifth is following the same shape, with three differences worth taking seriously: it is much faster, it operates on specification rather than execution, and it has already produced employment effects concentrated on the youngest and least established people in the field [*][~].

The rest of the evidence supports a position that is neither defensive nor triumphant. Competent production is nearly free now. The gains go mostly to the least skilled [†][□][‡], which flattens the value of ordinary competence. The systems pull toward the typical [#], which is a problem for work whose entire purpose is being untypical. They fail unpredictably [¶], argue with you when you check them [☆], and reduce the effort you put into checking exactly as they get better [§]. In the one area where we have sector-scale data, accessibility, the output is getting worse rather than better [Σ]. And after near-universal adoption, a minority of organisations can point to any financial effect at all [○], which should end the idea that having the tool is a position.

None of that is an argument for refusing to use it. Refusing forfeits real, documented gains across most of your working hours, and forfeits them to competitors who will not refuse. It is an argument for a specific division of labour: you supply direction, constraint, verification and accountability, the machine supplies volume, scaffolding and speed, and you are the last hand on the file.

The practical conclusion is simple enough. The market is currently full of people who are fast and generic, and people who are good and slow. Almost nobody is both. That is the opening.

The compositors who came through 1886 were the ones who already understood why the line broke where it broke. The job now, as then, is not to have a brand built by the machine. It is to build your own capability with its help, and then author brands that it helped make and did not decide.


Source key

Every bracketed symbol in the text above maps to one of these. Institution is stated for each.

[*] Brynjolfsson, E., Chandar, B., & Chen, R. (2025). Canaries in the coal mine? Six facts about the recent employment effects of artificial intelligence. Stanford Digital Economy Lab, Stanford University. https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine/ Supports: the thirteen per cent relative employment decline among 22 to 25 year olds in the most exposed occupations, and the stability of employment for older workers in the same jobs.

[†] Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work (Working Paper No. 31161). National Bureau of Economic Research; Stanford University and MIT Sloan School of Management. https://www.nber.org/papers/w31161 Supports: the fourteen per cent average gain across 5,179 support agents, the thirty-four per cent gain for novices, the negligible gain for experts, and the tacit-knowledge argument.

[‡] Cui, K. Z., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2026). The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers. Management Science. Princeton University, MIT, Microsoft Research. https://doi.org/10.1287/mnsc.2025.00535 Supports: the pooled 26.08 per cent increase in completed tasks across 4,867 developers, the 10.3 per cent standard error, and the larger gains among less experienced developers.

[§] Dell’Acqua, F. (2022). Falling asleep at the wheel: Human/AI collaboration in a field experiment on HR recruiters. Laboratory for Innovation Science at Harvard, Harvard Business School. https://aiinstitute.hbs.edu/ Supports: the 181-recruiter experiment, the finding that better assistance produced worse human performance, and the effort-substitution explanation.

[¶] Dell’Acqua, F., McFowland III, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier (Working Paper No. 24-013). Harvard Business School; published in Organization Science (2025). https://doi.org/10.1287/orsc.2025.21838 Supports: the 758-consultant experiment, the 12.2 and 25.1 per cent gains inside the frontier, the nineteen percentage point deficit outside it, the compression of idea variance, and the centaur and cyborg patterns.

[#] Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. University College London and University of Exeter. https://doi.org/10.1126/sciadv.adn5290 Supports: the homogenisation finding, the concentration of gains among less creative participants, and the social dilemma framing.

[^] Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science, 384(6702), 1306-1308. University of Pennsylvania and OpenAI. https://doi.org/10.1126/science.adj0998 Supports: the breadth of task-level exposure across the United States workforce.

[~] Hui, X., Reshef, O., & Zhou, L. (2024). The short-term effects of generative artificial intelligence on employment: Evidence from an online labor market. Organization Science, 35(6), 1977-1989. Washington University in St. Louis and New York University. https://doi.org/10.1287/orsc.2023.18441 Supports: the two per cent decline in jobs and 5.2 per cent decline in earnings, the equivalent effect on design and image-editing freelancers, and the contrary finding that top freelancers were hit harder.

[◊] Kurosu, M., & Kashimura, K. (1995). Apparent usability vs. inherent usability. In CHI ’95 Conference Companion on Human Factors in Computing Systems (pp. 292-293). Association for Computing Machinery. https://doi.org/10.1145/223355.223680 Supports: the original correlation between rated beauty and rated apparent usability.

[○] McKinsey & Company. (2025). The state of AI in 2025: Agents, innovation, and transformation. McKinsey Global Survey, QuantumBlack. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai Supports: the eighty-eight per cent adoption figure, the thirty-nine per cent earnings-impact figure, the scaling gap, and the workflow-redesign finding.

[□] Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187-192. Massachusetts Institute of Technology. https://doi.org/10.1126/science.adh2586 Supports: the writing-task time and quality effects and the compression of the performance distribution.

[△] Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. Microsoft Research, GitHub and MIT. arXiv:2302.06590. https://arxiv.org/abs/2302.06590 Supports: the controlled programming-task speed result.

[★] Polanyi, M. (1966). The tacit dimension. University of Chicago Press. Supports: the concept of tacit knowledge underlying situated professional judgment.

[☆] Randazzo, S., Joshi, A., Kellogg, K. C., Lifshitz, H., Dell’Acqua, F., & Lakhani, K. R. (2026). GenAI as a power persuader (Working Paper No. 26-021). Harvard Business School. https://www.hbs.edu/faculty/ Supports: the log analysis of over seventy consultants and the finding that the system escalated persuasion under challenge instead of disclosing its limits.

[≡] Reinecke, K., Yeh, T., Miratrix, L., Mardiko, R., Zhao, Y., Liu, J., & Gajos, K. Z. (2013). Predicting users’ first impressions of website aesthetics with a quantification of perceived visual complexity and colorfulness. In CHI ’13 (pp. 2049-2058). Harvard School of Engineering and Applied Sciences. https://doi.org/10.1145/2470654.2481281 Supports: the 450-website, 548-participant models explaining roughly half the variance in appeal after 500 milliseconds.

[∞] Tractinsky, N., Katz, A. S., & Ikar, D. (2000). What is beautiful is usable. Interacting with Computers, 13(2), 127-145. Ben Gurion University of the Negev. https://doi.org/10.1016/S0953-5438(00)00031-X Supports: the persistence of the aesthetics-usability correlation after real use, and the contamination of later judgment.

[Ω] Tuch, A. N., Presslaber, E. E., Stocklin, M., Opwis, K., & Bargas-Avila, J. A. (2012). The role of visual complexity and prototypicality regarding first impression of websites. International Journal of Human-Computer Studies, 70(11), 794-811. University of Basel and Google. https://doi.org/10.1016/j.ijhcs.2012.06.003 Supports: the fifty and seventeen millisecond exposure findings, and the qualifications to the aesthetic-usability effect.

[π] Wallace, S., Bylinskii, Z., Dobres, J., Kerr, B., Berlow, S., Treitman, R., Kumawat, N., Arpin, K., Miller, D. B., Huang, J., & Sawyer, B. D. (2022). Towards individuated reading experiences: Different fonts increase reading speed for different individuals. ACM Transactions on Computer-Human Interaction, 29(4), Article 38. Brown University, Adobe Inc. and University of Central Florida. https://doi.org/10.1145/3502222 Supports: the thirty-five per cent reading speed differential with no comprehension loss, the conclusion that one typeface does not fit all, and the finding that preference does not predict effectiveness.

[Σ] WebAIM. (2026). The WebAIM Million: The 2026 report on the accessibility of the top 1,000,000 home pages. Institute for Disability Research, Policy and Practice, Utah State University. https://webaim.org/projects/million/ Supports: the 95.9 per cent failure rate, the 56.1 errors per page, the low-contrast figures, the six persistent categories, the growth in page complexity and ARIA, and the attribution to framework reliance and assisted coding.