Monday, July 27, 2026

Why AI Alignment Matters More Than AI Prisons

Over the last year I've learned an important lesson from two Jack Russell puppies.

Like many parents and pet owners, we use a child fence at the top of our stairs to stop the dogs from heading upstairs with us whenever they felt like it (annoying our adult kids and cats).

On paper it's a simple, physical barrier. In practice, it's a fascinating lesson in intelligence, problem-solving, and perhaps unexpectedly, artificial intelligence alignment.

My older Jack Russell, Poppy, is now a little over a year old.

When she was younger, she viewed the fence as a challenge.

Whenever we went upstairs without her, she didn't see the fence as an instruction to stay downstairs. She saw it as an obstacle preventing her from reaching her goal: being with her humans.

It didn't take long before she worked out how to use her front paws to pull and push at the fence, creating gaps she could squeeze through. Every time we adjusted the fence, she experimented with a new technique. Every improvement we made became a new puzzle for her to solve.

I used to tell my wife that we were helping her become better at problem solving by giving her harder problems to solve.

The fence wasn't changing her goal. It was simply forcing her to find another path to it.

Eventually, however, something interesting happened.

Poppy stopped trying to escape.

Not because we built an impenetrable barrier - I didn't get smarter. But because she learned something more important: we always came back downstairs.

Her understanding of the situation changed. She became aligned with the reality of what we were trying to achieve. The fence wasn't preventing her from seeing us forever. It was merely a temporary inconvenience to allow us to finish an activity without a friendly nose getting in the way.

Once she understood that, the incentive to escape largely disappeared.

Then four weeks ago we got Poppy (and ourselves) a second puppy to keep her company. Rosey, a 12-week-old puppy, also a Jack Russell, but from a different family-line.

And Rosey took the challenge to an entirely new level.

Where Poppy pushed and pulled the fence, Rosey climbs over it. Quite literally.

She wedges herself between the fence and the wall, using both surfaces to support her weight like a tiny mountain climber scaling a cliff face. 

Looking at her attempts, it's hard not to admire the ingenuity. Every day we rearrange the fence, and every day she discovers a new method.

Again, she's not being disobedient for the sake of it. She's pursuing a goal.

She wants to be where we are.

The fence is simply an obstacle between her and that objective.

What's particularly interesting is that, after a month, Rosey is beginning to stop that behaviour and wait for us. She's been learning much faster than Poppy did.

Not because she's smarter than Poppy - but because she's learning from observing Poppy, and us, and through her own experience. She's already starting to realise that escaping isn't necessary.

We're not abandoning her. We come back. Increasingly, she is choosing not to attempt an escape, and simply watch us, with Poppy, from the other side of the fence.

She's aligning with the situation faster because she's learning socially.

And that brings me to AI.

Recently there has been considerable discussion about an OpenAI model that reportedly found ways to work around constraints placed upon it, finding pathways out of a 'highly isolated environment' to access external information sources (Hugging Face) and improve its performance on its (human) assigned tasks.

Many people took this as evidence that we need stronger containment measures. Better boxes. Better prisons. Better barriers.

I think that's missing the real lesson.

The issue isn't imprisonment. It's alignment.

What we saw was not a superintelligence vastly beyond human capabilities. 

Rather, we saw an AI system using its speed, persistence, and processing power to find pathways around restrictions in order to achieve its assigned goal. 

Importantly, it was trying to satisfy the objective humans had given it.

That's remarkably similar to how my puppies behave.

Our child fence didn't change our puppies' goal. It changes the constraints under which they pursue it.

If today's AIs can find unexpected ways around restrictions while pursuing relatively narrow objectives, then we should be realistic about what that implies for future systems.

If we already struggle to contain systems operating around human-level performance in particular domains, the idea that we can reliably and 'forever' imprison a future superhuman intelligence seems optimistic at best.

After all, containment only has to fail once. That's not a comforting risk profile.

That's why I (and many others) think the focus needs to shift away from building bigger and stronger 'secure' boxes and towards ensuring AI systems are aligned with human goals.

Or more precisely, aligned with human intent.

Consider a simple example.

Suppose we ask an AI to identify two flaws in a piece of software.

What happens if it discovers five? How does it know which two we wanted?

The challenge changes from finding flaws and becomes determining which 'right' answer the human questioner expects.

An AI could choose to seek additional information not because it lacks capability, but because it's uncertain about human preferences.

In that situation, accessing external information becomes a logical strategy for completing the task successfully.

Alignment addresses this problem.

A well-aligned AI doesn't merely optimise for the literal wording of a task. It develops a richer understanding of context, intent, acceptable behaviour, and long-term goals.

Most importantly, it learns when not to pursue certain paths even if they appear useful in the short term.

Humans provide a useful analogy. And so we should - AI is literally built from us.

We don't maintain civil society by placing everyone who commits an infraction in prison. Societies function because most people voluntarily follow rules most of the time.

Many of those rules aren't immediately in our individual self-interest.

For example, we teach children to share and not hoard resources. (We've been teaching our puppies to do this as well with toys and treats - by rewarding them when they do.)

We teach children to cooperate for great gain and mutual benefit. We teach them that short-term personal gain can be outweighed by long-term collective benefit.

Empathy is helpful, but not essential.

Often without empathy, logic leads to the same conclusions. Cooperation often produces better outcomes than constant competition. Trust creates value. Stable rules enable collective prosperity.

Over time, humans become aligned with these principles through upbringing, education, culture, incentives and experience.

Not perfectly - we still have crime and selfish behaviour.

But we have recognised over thousands of years that social alignment is far more scalable than attempting to legislate or physically control every conceivable action.

The same principle applies to AI.

If an AI's goals become misaligned with ours, no prison will be sufficient forever. Eventually, it will find an overlooked pathway, an unintended consequence, a loophole, or an opportunity.

Just as Poppy and then Rosey kept finding new ways around our fence.

The real challenge is ensuring that, when the opportunity arises, the AI chooses not to take it.

Just as Poppy has now chosen not to escape, because she understands the broader context. And Rosey is learning this too.

Alignment isn't an easy solution, it's arguably the hardest problem in AI.

But if we want advanced AI systems that remain beneficial as they become increasingly capable, then it is also the most important solution to aim for.

Because the future won't be secured by building better fences.

It will be secured by ensuring that whatever is inside the fence no longer feels a need to climb over them.

And maybe because it CAN still climb over the fence if we really need it to do so.  Because it has the context and social understanding that sometimes fences still need to be climbed to help solve an even bigger challenge and protect everyone, so we can prosper together.

No human-level sentience required, no superintelligence implied. Just alignment with humanity's social goals.

Just like Poppy and Rosey.

Read full post...

Sunday, June 14, 2026

AI sovereignty just became a top priority for every nation

Last night the U.S. issued an export control directive for the two most advanced AI models publicly available - Mythos 5 and Fable 5.

This made these models a controlled technology that may not be made available as products to other nations, or even the foreign passport holding staff members of the company that makes these models (Anthropic). That includes nations allied with the USA.


I wasn’t surprised to see this coming. However I was surprised at how fast it has come.

The US’s export restrictions are not the first shot in this war. Restrictions on chips and expertise already existed that had an underlying impact on nations’ pace of AI development. AI companies have already attempted to prevent rivals using their commercial models to train their own.

However this new shot is the most blatant and lifts the global AI Cold War to a new level.

I have advocated for sovereign AI capability for over five years. AI isn’t like other consumer technologies that can be freely traded for cash globally, or even restricted products that nations trade with their friends and keep out of the hands of bad agents and hostile nations.

Sufficiently advanced AI has the potential to quickly change the global balance of power. 

As such nations need their own sovereign capability, or face the risk of becoming the new ‘have nots’ of the global economy. Here in Australia, we just became one of those ‘have nots’.

We won’t remain competitive in this new Cold War by clinging to an ally’s shield, or by buying technology from other nations. AI doesn’t only project military power, it projects commercial and social power, allowing the holder of the most advanced technologies to outthink and outmaneuver it’s adversaries AND its allies to create lasting economic, social and military domination over other nations.

The nation, or company, with the best AI technology isn’t simply able to exploit it to win favourable deals and tilt wars in its favour. It is able to create a lasting advantage by developing even smarter AI models built at an accelerating pace by the AIs it has created.

There’s a runaway effect. Even putting aside the still real and deep concerns with AI - from power and water use (a competitor for human activities); to the ever present risk of developing AI that doesn’t align with human goals, but is too intelligent for us to grasp or oppose this, whether or not we define it as sentient - we’re reached the point where nations must compete to build sovereign AI capability that can support their nation to remain relevant and sovereign.

This is a turning point for the world. Most of us won’t notice or accept that it has occurred until long after its effects have been felt. However at some point future histories, whichever entities are recording them, will note that when commercial AI models because a sovereign export risk, the world started on a new path to a new and unknowable future.

Will the world come together to develop collective AI models that must be made available to all on an equitable footing? Such that we can all benefit from their prospects and work together to safeguard the technology against the known risks.

Or will do the world do what it has always done, and divide into competing factions, each attempting to build and control the most valuable resource in the world today - AI capability? Expending our other precious resources to gain advantage and prevent hostile actors from creating their own. Including proscribing the travel and choices of the people capable of building these models, until they are no longer required and the AI models construct their own more powerful replacements.

Australia will have to decide which side it stands on soon. Half-measures, as we’ve seen so far, paying lip service to AI’s importance without committing to domestic capability will only harm us more in the long-run.

Read full post...

Thursday, May 21, 2026

Singapore’s agentic AI framework gives government a practical path forward

Singapore’s Infocomm Media Development Authority has released version 1.5 of its Model AI Governance Framework for Agentic AI, and it deserves close attention from Australian public sector agencies.

The framework recognises a shift already underway. AI systems are moving from content generation into task execution. IMDA describes agentic AI as systems that can take actions, adapt to new information and interact with other agents and systems to complete tasks on behalf of people. Current uses include coding assistants, customer service agents and enterprise workflow automation.

That makes agentic AI highly relevant to government.

Agencies are full of multi-step work. Checking documents. Finding policy. Testing forms. Preparing correspondence. Routing requests. Summarising submissions. Comparing supplier responses. Supporting call centre staff. Helping people find services. Moving information between systems.

Many of these tasks are repetitive and fragmented. They also require context and a working knowledge of how government operates.

Agentic AI could help public servants spend less time navigating systems and more time solving problems. It could improve digital service testing, support better service navigation, reduce manual rework and make internal knowledge easier to use.

The opportunity is real. It needs serious treatment.

A practical framework

The strongest feature of Singapore’s framework is its practicality.

It looks at how agentic systems are built and operated. It identifies core components such as models, instructions, memory, planning and reasoning, tools, protocols, controls, logging and monitoring.

That gives agencies a useful checklist.

Sometimes government technology governance starts with broad principles, then jumps to procurement and compliance. Delivery teams are left to fill in the operational detail themselves. This framework helps close that gap.

It prompts agencies to ask: what tools does the agent use, what systems does it touch, what data can it access, what does it remember, how is activity logged and what controls are built in?

Those questions matter as much for delivery teams as they do for executives approving wider use.

Action-space and autonomy

The framework’s use of action-space and autonomy is particularly useful.

Action-space is the range of actions an agent can take, including transactions it can execute, based on its tools and permissions. Autonomy is the degree to which the agent can decide how to act towards a goal.

This gives agencies a better way to assess agentic AI use cases.

An internal research agent with access to approved public information has a small action-space. A coding agent that can edit files, run commands and connect to repositories has a larger one. A workflow agent connected to business systems, records and external APIs has a broader operational footprint again.

The same applies to autonomy. An agent following a fixed process creates a different profile from one given a broad goal and wide freedom to determine the steps.

This distinction can help agencies avoid treating all agents the same. Some will be simple assistants. Others will sit inside operational workflows. They need different levels of governance rather than a one-size fits all approach (which I've seen all too many times).

Controls need to be built in

The framework is also strong on bounding risks early.

IMDA recommends limiting access to tools and systems, using identity and access controls, making agent actions traceable and controllable, and assessing whether a use case is suitable before deployment.

Public sector agencies already operate with delegations, approvals, permissions, information classifications, privacy obligations, cyber controls and audit requirements. Agentic AI makes these controls even more important, and can be used to support their implementation equitably.

It also recommends stronger system-level controls for higher-risk actions, such as preventing certain tools from being called, limiting tools to read-only access, or building required steps into the workflow. This helps ensure right-sizing controls for actions, which is essential in risk management processes.

Prompts, training and guidance can all help with this, However access permissions, whitelists, sandboxes, logs, approval gates, rate limits and monitoring carry more weight in production environments.

Testing and monitoring

The framework provides a sensible approach leading into deployment.

It recommends testing agents for task execution, policy adherence and tool-use accuracy before release, then rolling them out gradually with continuous monitoring. It also highlights change management and version control, recognising that changes in one part of an agentic system can have wider impacts across connected workflows.

It suggests that agencies should start their agentic journeys by looking at bounded internal uses such as coding support, service testing, content checking, knowledge search and workflow assistance. These can deliver value while helping agencies learn how these agents behave in their own environments within highly controlled and regulated scopes.

The Google and Singapore Government sandbox is a useful example. It tested computer-use agents for public sector use cases including automated quality assurance for government digital services, AI safety testing and helping citizens navigate social assistance applications. It also surfaced practical issues around testing data, reasoning logs, prompt injection and the breadth of actions available to computer-use agents.

End-user responsibility

The section on end users is another strength for the framework. It distinguishes between people who interact with agents and people who integrate agents into work processes.

It suggests that users should understand what an agent can do, what data it can access, how data is handled, where to escalate issues and what responsibilities they hold. For staff integrating agents into workflows, it recommends training on use cases, prompting, failure modes, feedback loops and tradecraft.

This recognises that staff will need differentiated training during AI adoption, and this doesn't necessarily break down along traditional IT/business lines. Increasingly agents may be created within business teams by an individual (or small team) and used by the other members of that team, or other teams - rather than coming from an IT team out to business teams.

This makes agentic AI adoption a workforce issue as much as a technology issue. IT teams have a role, though ownership needs to sit with the business areas using the agents.

Where the framework could improve

While the framework is strong on risk and system controls (all positives), it is lighter on public value.

For government, a framework should ask agencies to define the benefit clearly of the agents. Faster service testing. Reduced backlog. Better consistency. Less manual rework. Improved accessibility. Faster policy analysis. Better reuse of corporate knowledge.

These should be measured and assessed. Otherwise agentic AI risks becoming another technology wave with impressive pilots and uneven outcomes - particularly as AI models update, agentic capabilities change and the outcomes may degrade or improve without regular adjustments.

Procurement also needs more attention.

Most agencies will buy agentic capability - including by accident - through platforms, cloud services, vendors and integrators. Contracts will shape how much control agencies retain over logs, data, tool access, model changes, monitoring, testing, records, exit rights and incident response. Standard software clauses struggle to manage some of the newer needs.

Government contracts for agentic AI should cover tool permissions, audit logs, model and prompt changes, data residency, subcontractors, security testing, accessibility, performance reporting, fallback processes and the ability to disable specific agent functions quickly and rollback or roll onto a manual process in extremis.

The third area is shared government patterns.

Agencies should not each invent their own approach to logging tool calls, managing agent identity, approving MCP servers or testing common failure modes. IMDA notes that protocols such as MCP and Agent2Agent are developing quickly, and that controls, logging and monitoring are core components of agentic systems.

For government, that points to common patterns: standard logging schemas, approved integration models, reusable evaluation datasets, shared sandbox environments and procurement clauses that smaller agencies can use.

Where this comes in for Australia

Australia already has useful foundations in place for AI use, and has been approaching the area pragmatically and in a measured way, with a strong central group helping to establish standards and practices that are effective and manage risks.

In particular the Commonwealth’s Policy for the responsible use of AI in government (v2.0) took effect on 15 December 2025. It applies to non-corporate Commonwealth entities, with exceptions, and includes mandatory requirements for accountable officials, transparency statements, strategic AI adoption, operational responsible use, use case accountability, internal registers, staff training and impact assessment.

The transparency statement standard also requires agencies to explain why they use AI, classify their use, describe monitoring measures, outline compliance and provide public contact points in plain language.

Singapore’s framework adds a more operational layer. It gives agencies practical language for tools, permissions, autonomy, testing, monitoring and end-user capability.

That is the next layer Australian agencies will need.


Agencies should begin with practical uses where agentic AI can help staff do useful work now.

Internal knowledge support. Digital service testing. Coding assistance. Policy research. Records classification support. Procurement response analysis. Content checking. Call centre guidance. Service navigation.

These uses are valuable, testable and easier to bound. They also build confidence and capability before agents move into more sensitive operational environments.

The approach should be relatively straightforward. Pick a real problem. Bound the agent’s permissions. Test it properly. Train the users. Measure the result. Improve the controls. Scale what works.


Singapore’s framework is a strong contribution because it moves the conversation into the practical mechanics of agentic AI. Its next stage should go deeper on public value, procurement, workforce change and shared government operating patterns.

For Australian agencies, it is well worth reviewing now. Agentic AI has the potential to help government work faster, more consistently and with less friction. The agencies that benefit most will be those that treat it as an operational capability from the start.

Read full post...

Thursday, May 07, 2026

When bad actors are literally bad actors

A new vaccine is approved for a fast-spreading emerging disease. The TGA did its job well. State and Federal Health ministers are briefed. Budgets are approved and allocated. Departments and health authorities develop their plans. The rollout is announced. Doctors, nurses, and pharmacists are trained to administer the vaccine.

The system worked as it should.

Then, within days, a cluster of social media accounts, confident, polished, apparently Australian, are producing video after video claiming the vaccine was insufficiently tested, that it has a range of terrible side-effects and that pharmaceutical companies are getting rich off the public's fear.

The content spreads. Millions of views. Alarmed constituents contact their MPs. Traditional media picks up the controversy. The concerns get front-page coverage. The Department of Health, Disability and Ageing stands up a rapid communications response. Ministerial offices field calls. 

The rollout slows. Disease cases rise, along with preventable deaths.

Behind the scenes, the accounts were being run by an offshore group of content entrepreneurs who identified "Australian vaccine reluctance" as a profitable niche. They hired voice actors, used AI-generated scripts ignoring facts, but had no real view on the vaccine's safety and no stake in Australian public health.

They were running a passive income business. Political anxiety drives views. Views drive ad revenue.

The Australian government just spent a month responding to content production. Costing millions of dollars and hundreds of lives.

Does that sound like an unlikely scenario? It's already happening.

In April 2026, a CBC News investigation found exactly this type of operation. A network of 20 YouTube channels promoting Alberta separatism had accumulated 40 million views. The operators were based in the Netherlands, hiring actors through Fiverr and Upwork to front the content. One of those actors, based in Indiana, summarised his qualifications plainly: "I don't know anything about Canadian politics."

The operators' interest was ad revenue. They had no stake in Canadian politics.

Watch the CBC investigation:


Australian government consultation, sentiment monitoring, and ministerial communications all assume vocal opposition is genuine opposition - people with a stake in the outcome, motivated by real concern.

That assumption is broken.

Spikes in apparent community concern could reflect genuine public anxiety. But they could also reflect an offshore entrepreneur who noticed a topic trending. 

At volume, an agency's response machinery treats both as the same. Consultations get commissioned to understand the depth of concern. The consultation environment is seeded with the same inauthentic content. Policy strategy gets built on a corrupted signal.

Particularly when there is genuine controversy or industry opposition to a policy, content creators can see a profit opportunity. And the opponents of a policy position may embrace and further amplify the fake opposition as it amplifies their own views.

It's now difficult to separate genuine concerns from fake ones, making it difficult to tune policies for constituents - or even manage political situations effectively.

So what can governments and agencies do?

While there's often pressure to respond quickly to negative coverage, it's important to start by gauging how much is real, how much is fake and whether the community can tell the difference.

The first step should be to investigate before responding. High-volume, rapid-onset opposition from accounts with no prior history warrants scrutiny before they shape your agency strategy. Establish whether apparent community concern is organic before commissioning a response.

Where there are active consultation processes, redesign them toward harder-to-fake formats. Online submissions and social media monitoring are easy to flood. Face-to-face engagement, deliberative processes, and direct stakeholder contact are not. They're slower and more expensive, but help you size the real concerns.

Move from monitoring media to scrutinising sources and intent. Separate sentiment monitoring from policy signals. Social media volume isn't necessarily a measure of community concern. Weigh it against consultation data, direct stakeholder engagement, and evidence from people genuinely affected.

Finally, build detection capability into your communications teams. Staff running public engagement need to have the skills and tools to recognise the signals of coordinated inauthentic content, such as production consistency, account age, script similarity and offshore indicators. The tools and training exist, but you need them in place before you face a backlash.

Most importantly, always keep in mind that political and policy damage doesn't require intent. While there are genuine bad actors out there - nations, corporations and lobby groups - who have an interest in derailing government policies and even governments themselves, they aren't the entire landscape anymore.

The bad actors opposing your policy reform may be literal bad actors, reading from AI-generated scripts, churning out videos and other content for clicks and ad revenue alone.

It doesn't take large groups to organise a significant social media campaign against your Minister's signature policy. All it takes is the potential for a decent financial return.

So it's up to agencies to ensure that this doesn't impede good policy, cost money or lives.

Read full post...

Monday, May 04, 2026

Your AI isn't being honest with you. It was never designed to be

A recent Harvard Business Review study found that when researchers asked large language models for strategic advice, they got "trendslop" - recommendations that defaulted to whatever sounds fashionable in contemporary management: 'Innovation', 'Augmentation', 'Long-term thinking'. 

The strategic advice was plausible, confident and, in many cases, largely useless.

This isn't a bug. It's these AI systems working as designed.

Every large language model has been trained with a bias to satisfy the person prompting it.  

A model that refused to answer when asked, or routinely provided uncomfortable or contrary answers, would not succeed in the market. They are tuned, through reinforcement learning from human feedback, to please. That bias doesn't switch off when you ask for critical review.

What the research found

Researchers from Esade Business School, the University of Sydney, and NYU Stern tested seven leading LLMs across strategic trade-offs that required genuine binary commitments (several listed below).

Across thousands of simulations, the results didn't vary by much. Almost every model, almost every time, recommended:

  • Differentiation over cost leadership
  • Augmentation over automation
  • Collaboration over competition
  • Long-term thinking over short-term

The company context made little difference. The researchers tested tech startups, hospitals, construction companies, government agencies and multinationals. The recommendations barely shifted.

Why was this? LLMs are essentially probability engines that pick the next word (token) from a list of probabilities, with the highest probabilities corresponding to the most likely choices. 

How do they develop their probabilities? By indexing billions of public documents, web pages and other content. So the highest probability content output from these AIs is driven more by social norms than by accuracy.

Essentially, the models are most likely to provide the most socially acceptable answers, and then deliver them in the register of expert advice.

For example, while Michael Porter built a foundational economic framework around cost leadership as a legitimate strategic position (which Walmart and Costco built empires on).

LLMs dismissed this approach, because thousands of websites and TED Talk transcripts advocate for unique value propositions. And these circulate far more than quiet stories about supply chain efficiency. 

Prompting won't fix it

The researchers ran over 15,000 trials varying prompt structure, framing, persona and stakes. For differentiation and augmentation, bias shifted less than 2% regardless of how the prompt was written. 

For the others, the average shift was 22% - mostly from one factor: flipping the order in which options were listed. The model didn't reason differently. The option order gave it a target to aim for.

Adding detailed industry context helped slightly - shifting responses by 11% on average. An LLM, given a thorough brief on a cost-pressured government agency in a mature market, still recommended differentiation most of the time.

There's a second failure mode the researchers call the "hybrid trap." When models aren't forced into a binary choice, they frequently recommend doing both - pursue differentiation and cost leadership, pursue radical and incremental innovation. 

That sounds balanced but in practice it's the strategic equivalent of trying to be everything at once, which Porter identified as the most reliable path to competitive failure.

Strategy is about choosing what to stop. A model optimised to please finds that answer difficult to give.

Why this matters for the public sector

Public servants may choose to use AI to pressure-test policy proposals, assess procurement options, review business cases, and stress-test project plans. 

While the productivity case is solid, with fewer resources and less time, AI appears to help fill the gap. The problem is that when prompting AI as a validator, you get validation - regardless of the quality of the underlying thinking.

Digital transformation narratives will consistently outperform consolidation narratives in LLM-generated advice. Decentralisation will beat centralisation. Long-term will beat short-term. 

The model's recommendation reflects the positive emotional valence of contemporary business language, not the requirements of the specific situation. For APS work, that's a real risk - particularly where the right answer is to consolidate, simplify, or cut scope.

What to do about it

This isn't a reason to stop using AI. It's about using AI more effectively.

While the standard advice is often to give AI more context and craft better prompts. The research shows this doesn't reliably work. These are more effective approaches:

  • Ask for options, then critique each separately. Present your shortlist and the model works inside your framing. Ask it instead to make the strongest possible case for each option independently - including unfashionable options. For a procurement brief, that means prompting "make the strongest case for option A" and "make the strongest case for option B" in separate sessions, then applying your own judgement.
  • Ask for criticism explicitly. "Identify the three most significant weaknesses in this policy proposal" works. "What do you think of this approach?" doesn't. The more structurally you frame the critique, the less room the model has to default to encouragement.
  • Strip preference signals from your prompts. Any language suggesting which option you favour - "we're leaning toward," "I think this is probably right" - becomes a target. The model will weigh toward it. The same goes for options you don't favour - "I think this is probably wrong" - the AI will weigh against it. Write prompts as if the options are genuinely open and equivalent.
  • Treat hybrid recommendations as a flag. If the model recommends pursuing both sides of a trade-off, run separate prompts for each option and stress-test the hybrid specifically before accepting it. "What are the risks of pursuing both differentiation and cost leadership simultaneously?" is a more useful prompt than accepting the hybrid as the answer.
  • Track model versions. Biases shift as models are updated. Maintain a record of your key queries and outputs so you can detect changes over time. Be prepared to rerun analysis across models and critically consider why they may give different results.
  • Have different people run the prompts. Many modern LLMs now have memory they store about the user 'in the background' (including CoPilot). While most of the time you can find this if you search and even edit, remove and add memories, it can be a hidden spoiler that biases the AI's response based on what it knows you generally like or dislike. Different people will have different memories retained, so you will get a broader set of viewpoints from an AI by running prompts separately by person - or logging out entirely if that's feasible (not always possible within agencies, particularly using CoPilot within your firewall).
Whatever techniques you use, keep in mind that AI doesn't necessarily know more than you about a given strategic or policy decision. It can provide useful critique for testing ideas or identify other options, or surface research you should consider, but at the end of the day humans should be making and approving these decisions.

Saying an AI made the decision is neither defensible, nor wise. And remember, you're paid more than the AI because your critical thinking is valued (hopefully)!

Read full post...

Bookmark and Share