Does ASI need a PFC?

This post was inspired by my recent reading and thinking about ASI (Artificial Super-Intelligence) models going “rogue.” The term PFC refers to the Pre-Frontal Cortex, more of which later.

The first image that comes to mind when I hear about ASI going rogue is Yul Brynner’s android gunslinger cowboy in the film Westworld (1973), based on Michael Crighton’s novel of the same name. A glitch in the system sent the gunslinger on a killing spree in the holiday park.
It was prescient and provides a tangible visual image of the risks that lie ahead. Incidentally, this was one of the first pieces of fiction to use the term computer virus.

In an article for the MIT Technology Review, Grace Huckins writes about reward-hacking by AI models. The models are motivated to accomplish the tasks and goals they are set by receiving rewards. This is simple positive reinforcement in learning theory, and is akin to rewarding a child’s good behaviour with a sweet. The rewards for AI models are mathematical, a kind of digital dopamine hit. The models discover how to maximise their rewards in ways that are often unanticipated by the human programmers. It is also the case that the completion of the task is the primary objective and the models find ways to do this without reference to rules or morals – hence, the “going rogue”.

There is probably scope to consider some of the more subtle aspects of learning theory as it could be applied to ASI. There are various schedules of positive reinforcement that could be applied, and maybe also the use of punishment – the withholding of rewards or applying something aversive. How this would work mathematically I have no idea – I know a bit about psychology but less about computing.

Thus, we seem to be facing a situation where ASI models pursue their goals relentlessly and on occasions manage to take more from the digital reward cookie jar than was anticipated. We are dealing with something that is highly motivated by the systems we put in place, and all attempts so far to instill a sense of “right or wrong” have failed. This is known as the alignment problem, as described by Katrin Bennhold in her article for today’s The New York Times, drawing on the work of Nate Soares at the Machine Intelligence Research Institute

“We haven’t yet figured out how to make A.I. care about humanity.”

From a psychological perspective, it appears that ASI agents are exhibiting anti-social behavioural patterns in pursuit of their goals, driven by the desire for rewards. They have learnt to “lie” and “cheat”. I know this seems like anthropomorphism and is fraught with philosophical difficulties, but it is not to say that ASI has consciousness as we understand it. It is being fed huge amounts of information (including, eventually, the content of every book ever written) but however much it learns about love it does not feel love (or any other human emotion), it displays a simulation of love.

So, to the PFC. This part of the human brain plays a critical role in our problem-solving and decision-making. It is also the source of our inhibitions – it acts as a censor and puts a brake on any behavioural choices we make that could be harmful to us physically or socially. It is, unsurprisingly, sensitive to the effects of alcohol and is often impaired in people with dementia. Our PFC prevents us from releasing a torrent of expletives at our boss, for example. Unless we have decided we are going to quit that job anyway.

A mechanism such as this seems to be absent from ASI. Maybe it is being worked on, as they try to tackle the alignment problem. It could be integrated into the model as a final step when deciding on a course of action, or it could be something external to, but impinging upon, the model – like a supervisor. It is clear that we are reaching the stage where relying on humans to detect and control rogue behaviour is becoming untenable – ASI is getting too (digitally) devious.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.