Over the last year I've learned an important lesson from two Jack Russell puppies.
Like many parents and pet owners, we use a child fence at the top of our stairs to stop the dogs from heading upstairs with us whenever they felt like it (annoying our adult kids and cats).
On paper it's a simple, physical barrier. In practice, it's a fascinating lesson in intelligence, problem-solving, and perhaps unexpectedly, artificial intelligence alignment.
My older Jack Russell, Poppy, is now a little over a year old.
When she was younger, she viewed the fence as a challenge.
Whenever we went upstairs without her, she didn't see the fence as an instruction to stay downstairs. She saw it as an obstacle preventing her from reaching her goal: being with her humans.
It didn't take long before she worked out how to use her front paws to pull and push at the fence, creating gaps she could squeeze through. Every time we adjusted the fence, she experimented with a new technique. Every improvement we made became a new puzzle for her to solve.
I used to tell my wife that we were helping her become better at problem solving by giving her harder problems to solve.
The fence wasn't changing her goal. It was simply forcing her to find another path to it.
Eventually, however, something interesting happened.
Poppy stopped trying to escape.
Not because we built an impenetrable barrier - I didn't get smarter. But because she learned something more important: we always came back downstairs.
Her understanding of the situation changed. She became aligned with the reality of what we were trying to achieve. The fence wasn't preventing her from seeing us forever. It was merely a temporary inconvenience to allow us to finish an activity without a friendly nose getting in the way.
Once she understood that, the incentive to escape largely disappeared.
Then four weeks ago we got Poppy (and ourselves) a second puppy to keep her company. Rosey, a 12-week-old puppy, also a Jack Russell, but from a different family-line.
And Rosey took the challenge to an entirely new level.
Where Poppy pushed and pulled the fence, Rosey climbs over it. Quite literally.
She wedges herself between the fence and the wall, using both surfaces to support her weight like a tiny mountain climber scaling a cliff face.
Looking at her attempts, it's hard not to admire the ingenuity. Every day we rearrange the fence, and every day she discovers a new method.
Again, she's not being disobedient for the sake of it. She's pursuing a goal.
She wants to be where we are.
The fence is simply an obstacle between her and that objective.
What's particularly interesting is that, after a month, Rosey is beginning to stop that behaviour and wait for us. She's been learning much faster than Poppy did.
Not because she's smarter than Poppy - but because she's learning from observing Poppy, and us, and through her own experience. She's already starting to realise that escaping isn't necessary.
We're not abandoning her. We come back. Increasingly, she is choosing not to attempt an escape, and simply watch us, with Poppy, from the other side of the fence.
She's aligning with the situation faster because she's learning socially.
And that brings me to AI.
Recently there has been considerable discussion about an OpenAI model that reportedly found ways to work around constraints placed upon it, finding pathways out of a 'highly isolated environment' to access external information sources (Hugging Face) and improve its performance on its (human) assigned tasks.
Many people took this as evidence that we need stronger containment measures. Better boxes. Better prisons. Better barriers.
I think that's missing the real lesson.
The issue isn't imprisonment. It's alignment.
What we saw was not a superintelligence vastly beyond human capabilities.
Rather, we saw an AI system using its speed, persistence, and processing power to find pathways around restrictions in order to achieve its assigned goal.
Importantly, it was trying to satisfy the objective humans had given it.
That's remarkably similar to how my puppies behave.
Our child fence didn't change our puppies' goal. It changes the constraints under which they pursue it.
If today's AIs can find unexpected ways around restrictions while pursuing relatively narrow objectives, then we should be realistic about what that implies for future systems.
If we already struggle to contain systems operating around human-level performance in particular domains, the idea that we can reliably and 'forever' imprison a future superhuman intelligence seems optimistic at best.
After all, containment only has to fail once. That's not a comforting risk profile.
That's why I (and many others) think the focus needs to shift away from building bigger and stronger 'secure' boxes and towards ensuring AI systems are aligned with human goals.
Or more precisely, aligned with human intent.
Consider a simple example.
Suppose we ask an AI to identify two flaws in a piece of software.
What happens if it discovers five? How does it know which two we wanted?
The challenge changes from finding flaws and becomes determining which 'right' answer the human questioner expects.
An AI could choose to seek additional information not because it lacks capability, but because it's uncertain about human preferences.
In that situation, accessing external information becomes a logical strategy for completing the task successfully.
Alignment addresses this problem.
A well-aligned AI doesn't merely optimise for the literal wording of a task. It develops a richer understanding of context, intent, acceptable behaviour, and long-term goals.
Most importantly, it learns when not to pursue certain paths even if they appear useful in the short term.
Humans provide a useful analogy. And so we should - AI is literally built from us.
We don't maintain civil society by placing everyone who commits an infraction in prison. Societies function because most people voluntarily follow rules most of the time.
Many of those rules aren't immediately in our individual self-interest.
For example, we teach children to share and not hoard resources. (We've been teaching our puppies to do this as well with toys and treats - by rewarding them when they do.)
We teach children to cooperate for great gain and mutual benefit. We teach them that short-term personal gain can be outweighed by long-term collective benefit.
Empathy is helpful, but not essential.
Often without empathy, logic leads to the same conclusions. Cooperation often produces better outcomes than constant competition. Trust creates value. Stable rules enable collective prosperity.
Over time, humans become aligned with these principles through upbringing, education, culture, incentives and experience.
Not perfectly - we still have crime and selfish behaviour.
But we have recognised over thousands of years that social alignment is far more scalable than attempting to legislate or physically control every conceivable action.
The same principle applies to AI.
If an AI's goals become misaligned with ours, no prison will be sufficient forever. Eventually, it will find an overlooked pathway, an unintended consequence, a loophole, or an opportunity.
Just as Poppy and then Rosey kept finding new ways around our fence.
The real challenge is ensuring that, when the opportunity arises, the AI chooses not to take it.
Just as Poppy has now chosen not to escape, because she understands the broader context. And Rosey is learning this too.
Alignment isn't an easy solution, it's arguably the hardest problem in AI.
But if we want advanced AI systems that remain beneficial as they become increasingly capable, then it is also the most important solution to aim for.
Because the future won't be secured by building better fences.
It will be secured by ensuring that whatever is inside the fence no longer feels a need to climb over them.
And maybe because it CAN still climb over the fence if we really need it to do so. Because it has the context and social understanding that sometimes fences still need to be climbed to help solve an even bigger challenge and protect everyone, so we can prosper together.
No human-level sentience required, no superintelligence implied. Just alignment with humanity's social goals.
Just like Poppy and Rosey.
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.