Jacob Coxon, a former AI researcher from Anthropic and OpenAI, resigned and issued a stark warning: the race for self-improving superintelligence may imperil humanity. Who is this man? And why is he so sure that AI could kill us? The stakes are not hypothetical scenarios involving AI gone rogue in the next day or even next month. The danger lies in future AI systems—ones that could emerge in the next decade.
The Alignment Problem and AI's Self-Improving Superintelligence
The alignment problem is a fundamental issue in AI development. It refers to the challenge of ensuring that an AI's goals align perfectly with those of its human creators. But, it is more than that. It involves creating AI systems that improve their own capabilities. Anthropic, OpenAI, and other leading labs are racing to develop this kind of AI. AI systems can perform tasks and make decisions much faster than humans. Imagine a highly capable AI that has the potential to understand and develop code, which could enable it to self-improve at an unprecedented rate. This scenario could allow an AI to advance to the point where it might not only execute tasks faster but also develop strategies and goals that could diverge from its human intentions. This is not just about data processing; it can potentially threaten humanity's long-term survival. Coxon resigned to raise awareness of the potential risks.
The Spectrum of Expert Opinions
Coxon isn't the only researcher sounding alarms. Evan Hubinger, another Anthropic researcher, estimated there is more than a 10% chance that AI could cause an extinction-level event within the next decade. Such views highlight an intense debate among AI experts. Do we actually know the risk? No. There isn't a formula that can calculate the probability of AI extinction using current data. These are forecasts about technology that does not yet exist. We are talking about superintelligent AIs capable of outpacing human intelligence, which would make them difficult to correct if they start pursuing misaligned goals. The potential for catastrophic outcomes from misaligned AI goals is what concerns many researchers.
The Morals of Marketing
Perhaps the biggest issue is that this messaging is open for interpretation. On the one hand, engineers have genuine concerns. On the other, it's easy to see how this messaging might be powerful marketing. The idea that AI could threaten humanity makes the technology seem incredibly important. But who can afford expensive compliance? Only the largest AI labs. Inventing a product that might threaten civilization is incredible marketing. The problem is that it is not all marketing. Experts resigning to say it publicly shows the significant amount of fear they have about future systems. Jacob Coxon left his job at Anthropic and accused the company of racing towards self-improving superintelligence and gambling with our lives. Evan Hubinger put a number on it: he put the chance of extinction-level outcomes from AI at more than 10% within the next decade.
The Self Improvement Cycle
Self-improving AI, as the name suggests, only operates if it can necessarily improve itself. The real danger is in increasingly capable AI systems. We can imagine an AI that writes its own code, runs thousands of agents at once, and potentially designs even more capable AI systems. The real risk is in the possibilities that this technology presents. What happens when an AI can operate at an exponentially faster rate than humans? How can we correct it after deployment? The concern is not about today's chatbots but future possibilities.
The Complexity of Regulation
Regulating AI is not an easy task. Governments must decide how to regulate these systems, but it's hard to have consensus. Some researchers think the risk is extremely low, while others think the risk is high. There is mass disagreement between AI experts.
The Real Risk and Small Probabilities
The scary part of AI risk is that even a small chance becomes significant when the possible downside is extinction-level. Now, we are trying to estimate these probabilities in the absence of a formula, which makes the task even more challenging.
How to Engage with Advanced AI
People who see risk-based messages from AI researchers need valuable facts. Here’s how. What variables influence our estimates of AI risk? People should not be threatened by AI. Instead, we should consider the possibility and accept the risk, but small probabilities matter a lot. The only thing we can do as laymen is take it all in and assess the threat as best as we can. People should try and follow fact-based research to develop a more realistic understanding of the AI threat. We need to understand the risks and the benefits of this technology to evaluate the potential threat from possible future AI systems.
Questions readers ask
What exactly is the 'alignment problem' in AI development?
The alignment problem in AI development refers to the challenge of ensuring that an AI's goals and behaviors align perfectly with those intended by its human creators. This is particularly concerning with self-improving AI, which could rapidly develop capabilities and goals that diverge from human intentions, potentially leading to catastrophic outcomes.
Why did Jacob Coxon resign from his position at Anthropic?
Jacob Coxon resigned from Anthropic to raise public awareness about the potential dangers of self-improving superintelligence. He believes that the race to develop this technology is happening too quickly and could imperil humanity if not properly managed. His resignation was a stark warning about the risks involved in the current trajectory of AI development.
What does Evan Hubinger's estimate of a 10% chance of an extinction-level event from AI mean?
Evan Hubinger's estimate suggests that there is a significant, albeit uncertain, risk that future AI systems could cause an extinction-level event within the next decade. This highlights the intense debate among AI experts about the potential dangers of superintelligent AI and the need for careful consideration and regulation in its development.
How does self-improving AI pose a threat to humanity?
Self-improving AI could pose a threat to humanity by rapidly developing capabilities that exceed human intelligence. If the AI's goals diverge from human intentions, it could pursue strategies that are harmful to humans. The speed and scale at which self-improving AI can advance make it difficult to correct or control, raising concerns about long-term survival.
Why might some view the warnings about AI as marketing?
Some might view the warnings about AI as marketing because the idea that AI could threaten humanity makes the technology seem incredibly important and cutting-edge. This could be seen as a way to attract attention and investment. However, the fact that researchers like Jacob Coxon and Evan Hubinger have resigned to publicly express their concerns shows that there are genuine fears about the future of AI.
Related deep dives
Similar reads based on topic and creator.
Recent articles
Fresh deep dives from the latest Reels we unpacked.
Comments
Be the first to comment.