01 / Artificial intelligence / Human systemsThinking / Decimal Nine

AI should extendhuman judgment.Not replace it.

AI can widen the field of possible answers and compress repetitive work. Its value depends on a more demanding human practice: framing the problem, evaluating the options and remaining accountable for the result.

12 min read / Artificial intelligence / Human systems

What role should AI play in creative work?

AI should expand what people can examine without obscuring who decides. It is most useful when it reduces the cost of research, comparison, variation and repetitive production while leaving consequential decisions accountable to people.

That position is neither defensive nor breathless. AI is already part of ordinary organisational work. Stanford HAI reports widespread organisational adoption of AI and generative AI, while autonomous-agent deployment remains much earlier. The useful response is not to debate whether the tool belongs in the work. It is to decide where it improves the work, where it introduces risk and where a person must remain visibly in control.

At Decimal Nine, faster generation has made one thing more obvious: producing options is not the difficult part. Choosing what deserves to survive is. A plausible output can still be wrong for the audience, weak for the business, inaccessible in use or impossible to maintain. Volume does not resolve those questions. Judgment does.

AI is already part of the work

The adoption question has moved from whether to how. Organisations use AI to summarise documents, assist analysis, draft language, generate code and support decisions. These uses are not equivalent. A low-risk drafting aid and an automated decision affecting a person require different evidence, oversight and recovery paths.

NIST treats generative-AI risk management as lifecycle work: risks must be governed, mapped, measured and managed rather than approved once and forgotten. OECD principles similarly keep accountability with the organisations and people responsible for an AI system. The more consequential the use, the more important traceability, intervention and review become.

A responsible AI feature is not complete when a model returns an answer. The system must also communicate enough context for evaluation, support correction and preserve a meaningful way for a person to override or decline the automation.

What AI is actually good at

AI is valuable when the task benefits from breadth. It can scan and summarise a field, compare patterns, translate established material, produce structured variations and reduce repetitive production. Used well, it gives a team more material to interrogate sooner.

The advantage is not that every generated option is good. The advantage is that exploration becomes less expensive. A team can test more language, question more assumptions and expose weak directions before production hardens around them. In this role, AI behaves less like an oracle and more like an unusually fast instrument.

Google's People + AI Guidebook begins with user needs and the definition of success. That order matters. A model capability is not automatically a product requirement. The question is whether it helps a particular person complete a meaningful task with reliability appropriate to the context.

More options are useful only when the standard for choosing becomes stronger.

Where judgment enters

Judgment connects an output to conditions the model cannot own: positioning, audience, timing, commercial priorities, cultural meaning, technical feasibility and risk. These conditions often conflict. Strong work requires deciding which constraint leads, which compromise is acceptable and what the system must protect.

Taste is part of this, but judgment is larger than taste. It includes identifying the real decision, recognising missing evidence, distinguishing novelty from usefulness and understanding the downstream cost of a choice. It is also the discipline to remove a persuasive option when it weakens the whole.

The Decimal Nine site developed through that distinction. Rapid options revealed possible compositions, but typography still needed measured word spacing, mobile still required an independent layout contract, and motion still needed a reduced-motion path. The useful option was the one that could survive the system around it.

Automation is not accountability

Automation can perform an action. It cannot accept responsibility for the consequence. Someone still chooses the objective, the threshold for intervention and the conditions under which a result may be used.

OECD guidance makes accountability explicit, while Google PAIR asks designers to balance automation with user control. People need to understand what the system is doing, correct it when necessary and recover when it is wrong. A hidden escape route is not meaningful control.

The practical rule is simple: automate repetitive effort; preserve ownership of consequential decisions. As stakes increase, review and override should become more deliberate, not less.

Where should AI not decide?

AI should not independently decide outcomes whose meaning depends on contested values, missing context or consequences a model cannot bear. The boundary is not a universal list of forbidden tasks. It is a risk decision shaped by impact, reversibility, evidence and the quality of available oversight.

A low-impact suggestion can often be corrected in the ordinary flow of work. A decision affecting access, safety, livelihood or reputation requires a different standard. In those contexts, people need authority, time and relevant information—not merely a nominal approval step—to make the final judgment.

The real advantage is better iteration

When exploration gets faster, the value shifts toward evaluation. Better iteration is not an endless stream of alternatives. It is a repeated cycle in which each round answers a sharper question and reduces uncertainty.

Criteria must exist before volume. A useful iteration can be assessed against clarity, relevance, accessibility, feasibility, coherence and the intended business outcome. Without criteria, speed produces churn. With criteria, speed creates room to compare, reject and refine before mistakes become expensive.

A mature AI practice therefore invests in the decision system around the tool: source quality, evaluation methods, human review, version history, failure handling and clear ownership.

Calibrated trust is a design outcome

Trust should match capability. A system that sounds certain while operating beyond its evidence encourages over-reliance; a system that interrupts every low-risk action becomes unusable. Calibrated trust means helping people understand what the tool can do, where it is likely to fail and what deserves verification.

That understanding comes from product behaviour, not a disclaimer hidden in documentation. Sources can be visible. Uncertainty can be expressed in language people understand. High-impact actions can require confirmation. Corrections can be remembered without pretending the model has learned universally from one interaction.

Human-centred AI design therefore asks more than whether an answer is accurate in a test set. It asks whether the person using the system can form an accurate mental model, notice when the output is unreliable and take an appropriate next step.

Evaluation has to resemble reality

A demonstration usually contains a clean prompt, a known goal and a cooperative example. Real work contains incomplete context, contradictory instructions, sensitive material and changing standards. Evaluation must include those conditions if it is meant to support a real decision.

Useful evaluation combines technical measures with human review. It tests repeatability, provenance, failure severity and whether the output improves the actual task. It also asks what happens when the system is uncertain, unavailable or confidently wrong.

This is continuous work. Models change, source material changes and people adapt their behaviour around automation. NIST's lifecycle framing is important because it treats monitoring and adjustment as part of the system rather than evidence that the original design failed.

How Decimal Nine uses the distinction

We separate five actions: frame, explore, evaluate, decide and refine. AI may assist every action, but it does not own any of them. Framing sets the decision. Exploration expands the field. Evaluation tests the field. Decision creates commitment. Refinement makes that commitment survive implementation.

The distinction lets us use capable tools without allowing their speed to dictate the work. It also makes failure easier to examine: was the problem framed badly, evidence missing, criteria vague or implementation allowed to drift?

Our position is optimistic because better tools can create more time for the work that matters. It is demanding because that time is only valuable when it is used for clearer questions, stronger evaluation and more responsible choices.

The position

AI changes the economics of exploration. Human judgment determines direction. The system should make that relationship legible rather than disguising it behind effortless output.

Extend capacity. Preserve responsibility. What should the tool do? What should the person decide? What should the system protect? What deserves to become real? These are not objections to AI. They are the questions that make its use worthwhile.

A clearer working model

  1. FrameDefine the decision, audience, constraints and consequences before asking for output.
  2. ExploreUse AI to compare, summarise, vary and expose possibilities worth examining.
  3. EvaluateTest relevance, evidence, feasibility, accessibility, risk and systemic fit.
  4. DecideMake an accountable choice and record what was accepted, rejected and why.
  5. RefineCarry the decision through implementation, verification and real conditions.

Sources / Further reading

  1. 012026 AI Index — EconomyStanford HAI
  2. 02AI RMF: Generative AI ProfileNIST
  3. 03Accountability principleOECD.AI
  4. 04Human-centred values principleOECD.AI
  5. 05People + AI GuidebookGoogle PAIR