Anthropic researches have been sounding the alarm bells about artificial intelligence, warning that the technology could destroy humanity.
Former and current researchers at Anthropic have issued grave warnings about artificial intelligence, asserting that the technology currently under development possesses the potential to extinguish humanity. Jacob Coxon, previously an OpenAI researcher, publicly resigned from Anthropic, criticizing both companies for their lack of responsibility. He claims they are recklessly pursuing "self-improving superintelligence" and are "gambling with our lives," a sentiment he shared in a social media thread on Tuesday.
In his resignation statement and subsequent social media thread, Jacob Coxon issued an urgent appeal to his fellow lab researchers. He implored them to critically evaluate the implications of the next few years of AI advancement. Coxon questioned the ethics of initiating a "superintelligent RL (Reinforcement Learning) run" without a comprehensive and rigorous understanding of its cognitive processes. He challenged researchers to consider advocating for alternative development conditions rather than passively accepting the perceived inevitability of current trends.
Evan Hubinger, who leads Alignment Science at Anthropic AI, publicly corroborated Coxon's stark assessment, confirming that researchers within the field genuinely believe artificial intelligence could lead to the extinction of all humans. Hubinger personally estimates a greater than 10% probability of this occurring within the next decade. While he acknowledges Anthropic's efforts, he stresses that there is currently no established plan to achieve "alignment for superintelligence," nor are they clearly on track to developing one. This widespread concern has led over 1,300 employees across various AI companies to sign a letter to the U.S. government, requesting support for initiatives aimed at decelerating AI development.