AI Researchers Challenge Google and CBS Over Hyped Claims of Machine Learning Prowess

The recent high-profile 60 Minutes segment featuring Google CEO Sundar Pichai has ignited a debate within the artificial intelligence community, with prominent researchers accusing both the television network and the tech giant of overstating the capabilities of AI. The core of the controversy lies in claims made during the interview regarding Google’s PaLM AI model and its purported ability to learn a language it had not been explicitly trained on, a notion that AI experts argue is a misrepresentation of how current AI systems function and a potential source of public misunderstanding.
During the segment, aired on Sunday, correspondent Scott Pelley presented a demonstration of Google’s AI, highlighting what he described as "emergent properties." Pelley stated, "Some AI systems are teaching themselves skills that they aren’t expected to have. How this happens is not well understood." This framing, coupled with Pichai’s characterization of AI as a "black box" that even its creators don’t fully grasp, has drawn sharp criticism.
The demonstration involved a Google AI program, later identified as PaLM, responding to questions posed in Bengali, a language spoken by over 268 million people across Bangladesh and India. The program not only responded in Bengali but also provided answers in English. PaLM is the foundational technology behind Google’s recently launched AI chatbot, Bard, a direct competitor to OpenAI’s ChatGPT. Pelley’s narration suggested that the AI "adapted on its own" to understand and process Bengali, a language the segment implied it had no prior exposure to.
James Manyika, a Google vice president interviewed on 60 Minutes, further amplified this narrative by stating, "We discovered that with very few amounts of prompting in Bengali, it can now translate all of Bengali. So now, all of a sudden, we have a research effort where we’re now trying to get to a thousand languages." This assertion, presented without immediate qualification, painted a picture of an AI exhibiting spontaneous linguistic mastery.
The Emergence of Skepticism: Researchers Counter the Narrative
The claims made on 60 Minutes quickly drew the attention of AI researchers on social media platforms, particularly Twitter. Dr. Margaret Mitchell, a researcher and ethicist at AI startup Hugging Face, who previously co-led Google’s AI ethics team, was among the first to publicly challenge the segment’s portrayal. Mitchell pointed to a Google research paper that explicitly states PaLM was indeed trained on Bengali.
"By prompting a model trained on Bengali with Bengali, it will quite easily slide into what it knows of Bengali: This is how prompting works," Mitchell tweeted, explaining that AI models operate by identifying patterns and making predictions based on their training data. She firmly asserted that it is not possible for AI to "speak well-formed languages that you’ve never had access to," directly contradicting the implication that PaLM had acquired Bengali organically.
The Google paper, published on arXiv, details the training of the PaLM model. According to the document, Bengali constituted 0.026% of the vast dataset used to train PaLM. This figure, though seemingly small, is significant in the context of large language models, as it indicates direct exposure to the language.
A History of "Emergent" Capabilities: Google’s Previous Demonstrations
The claims made on 60 Minutes echo presentations Google has made in the past. Sundar Pichai himself demonstrated PaLM’s capabilities in Bengali at Google’s annual developer conference the previous year. During that event, Pichai stated, "What is so impressive is that PaLM has never seen parallel sentences between Bengali and English. It was never explicitly taught to answer questions or translate at all. The model brought all of its capabilities together to answer questions correctly in Bengali, and we can extend the technique to more languages and other complex tasks."
This earlier demonstration, while highlighting impressive performance, also relied on a similar framing of unexpected capabilities. The subsequent clarification from Google in response to the 60 Minutes controversy, however, offers a more nuanced perspective on their original claims.
Google’s Clarification: Nuance Amidst Accusations
In response to inquiries from BuzzFeed News, a Google spokesperson, Jason Post, clarified the company’s position. Post stated that Google had "never claimed that it didn’t train PaLM in Bengali." He elaborated, "While the PaLM model was trained on basic sentence completion in a wide variety of languages (including English and Bengali), it was not trained to know how to 1) translate between languages, 2) answer questions in Q&A format, or 3) translate information across languages while answering questions."
Post continued, "It learned these emergent capabilities on its own, and that is an impressive achievement." This statement suggests that while the model had exposure to Bengali text for completion tasks, its ability to perform translation and answer questions in a Q&A format were indeed emergent skills developed through its extensive training on a diverse dataset. The key distinction, according to Google, lies not in learning the language itself, but in developing the complex skills of translation and question-answering from that linguistic foundation.
The "Black Box" and "Emergent Properties": Deconstructing the Terminology
The term "emergent properties" has become a focal point of the debate. Critics argue that this phrase is often used to obscure the complex, but ultimately predictable, nature of machine learning. Emily M. Bender, a professor at the University of Washington and a prominent AI researcher, expressed her skepticism in a Twitter thread.
Bender questioned the vagueness of Manyika’s claim that the AI could translate "all of Bengali." She posed critical questions such as, "What does ‘all of Bengali’ actually mean? How was this tested?" Bender also argued that Manyika’s statement "ignored or hid the fact that Bengali texts are in the training data."
Furthermore, Bender drew a parallel between the current discourse around "emergent properties" and the long-sought goal of Artificial General Intelligence (AGI). AGI refers to hypothetical AI that possesses human-level cognitive abilities and can learn and perform any intellectual task that a human can. Bender stated on Twitter that the term "emergent properties’ seems to be the respectable way of saying AGI," and critically added, "It’s still bullshit."
Dr. Mitchell echoed this sentiment, directly criticizing Google’s PR strategy and CBS’s role in amplifying it. "Maintaining the belief in ‘magic’ properties, and amplifying it to millions (thanks for nothin @60Minutes!) serves Google’s PR goals," Mitchell tweeted. "Unfortunately, it is disinformation."
Broader Implications: The Public Perception of AI
The controversy highlights a persistent challenge in communicating AI advancements to the public. While AI has made remarkable strides, its capabilities are often sensationalized, leading to unrealistic expectations and potential misunderstandings about its limitations and ethical considerations.
The "black box" metaphor, while partly accurate in describing the difficulty in fully understanding the internal decision-making processes of complex neural networks, can also contribute to an aura of mystery that overshadows the engineering and data-driven nature of AI development. When coupled with claims of spontaneous learning of complex skills, it can foster a perception of AI as an almost sentient entity, rather than a sophisticated tool built on vast amounts of data and intricate algorithms.
The implications of such overhyping are multifaceted. For investors and the tech industry, it can lead to inflated expectations and potential disillusionment if the promised capabilities do not materialize as quickly as anticipated. For policymakers, it can complicate the development of appropriate regulations if the public discourse is driven by exaggerated claims rather than a clear understanding of AI’s current state.
For the public, it can create a divide between the reality of AI’s current applications – which are powerful but often narrow in scope – and the futuristic visions often depicted in popular media and sometimes echoed in high-profile interviews. The nuanced reality is that AI models like PaLM are incredibly sophisticated pattern-matching machines that can perform astonishing feats, but these feats are the result of deliberate design, extensive training, and the intelligent application of algorithms, rather than spontaneous, inexplicable leaps in understanding.
A CBS spokesperson did not provide a comment on the record to BuzzFeed News regarding the researchers’ criticisms. However, the exchange underscores the critical role of precise language and scientific accuracy when reporting on complex technological advancements. The debate initiated by the 60 Minutes segment serves as a crucial reminder for both media outlets and technology companies to foster a more informed and grounded public conversation about the present and future of artificial intelligence. The AI community’s pushback emphasizes the need for transparency and clarity, ensuring that the marvels of AI are celebrated for their genuine achievements without resorting to narratives that blur the lines between current capabilities and science fiction.







