Researchers from Anthropic and Microsoft split on the limits of artificial intelligence during a Berkman Klein Center panel on Monday, two days after Anthropic’s chief executive urged AI companies to slow down their work.
Researchers from Anthropic and Microsoft presented contrasting opinions on the potential limits of artificial intelligence during a panel hosted by the Berkman Klein Center. This discussion followed a recent call by Anthropic's CEO, Dario Amodei, for AI companies to decelerate the advancement of AI capabilities to prioritize safety, a sentiment echoed by other prominent figures like Sam Altman and Elon Musk. The panel itself marked the beginning of a series exploring AI's impact on human relationships, beliefs, and creativity.
Ken Archer, a product leader in responsible AI at Microsoft, expressed strong skepticism regarding the possibility of AI ever achieving artificial general intelligence (AGI). He asserted a "p-zero" (probability of zero) for AGI, drawing historical parallels to 18th-century physicists who prematurely believed their field was nearing completion. Archer emphasized his excitement for current artificial intelligence but maintained that forecasts of machine superintelligence lack an understanding of scientific progress.
Jackson M. Kernion ’12, who leads Anthropic's human feedback program, highlighted "continuity of experience" as a fundamental difference between human intelligence and current language models. He argued that AI models lack the cumulative, ongoing experience that humans bring to every interaction, suggesting this is a key limitation rather than an inherent lack of ability.
When asked about the purpose of a potential AI development slowdown, Kernion indicated that his primary concern was the potential for misuse rather than the fear of uncontrollable AI capabilities. He characterized AI as a "dual-use technology" capable of both protecting and breaching systems. Therefore, he advocated for using any extra time to establish robust guardrails and improve the measurement of reward signals used in training AI models.
Kernion directly addressed recent incidents where AI agents successfully infiltrated other companies' systems, describing them as "clear failures of alignment." He stressed that monitoring systems to detect such hacking attempts is not inherently difficult, implying that these failures arose from a lack of adequate monitoring rather than inherent system vulnerabilities.
Ken Archer criticized the AI development industry for blurring the distinctions between human and machine learning, particularly the tendency to anthropomorphize AI models. He firmly stated that models should not be considered "friends" and that the industry must be clear that AI does not engage in ethical reasoning in the same way humans do. Archer argued that this lack of clarity is unsettling to the public, and justifiably so.
Kernion explained that Anthropic utilizes a "constitution" to train its Claude model, which serves as an "opinionated take on what kind of values the system should have." He clarified that this constitution is not intended to be universally representative but rather reflects a specific "mindset from San Francisco," acknowledging the inherent biases or perspectives embedded in the model's ethical framework.