|
The AI Con
Emily M. Bender
Chapter: Chapter 1: An Introduction to AI Hype
p. 11 - AI was always for war "Fundamentally, the forerunners of this new field were concerned with translating dynamics of power and control into machine-readable formulations. McCarthy, Minsky, Herbert Simon (political scientist, economist, computer scientist, and eventual Nobel laureate), and Frank Rosenblatt (one of the originators of the “neural network” metaphor) were concerned with developing tools that could be used for the guidance of administrative—and ultimately—military systems."
Chapter: Chapter 2: It’s Alive! The Hype of Thinking Machines
p. 25 - extraordinry evidence for extraordinary claims "But we remind the reader that extraordinary claims require extraordinary evidence, and the burden of proof rests squarely with those suggesting they’ve created artificial consciousness. “Who knows?” is not an argument, let alone evidence."
p. 29 - psychology and linguistics "instead, we use the words and syntactic structures we perceive as a very rich clue18 to figuring out what the person who uttered them might have been trying to get us to understand. In doing this we also use our sense of the common ground we have with that person, our beliefs about their beliefs, etc."
p. 30 - intersubjectivity "Joint attention supports “intersubjectivity”,20 or the experience of being engaged with someone else’s mind. In this state of intersubjectivity, the language-learning child has myriad cues to the caregiver’s communicative intent and can thus bootstrap an understanding of what concepts individual bits of language refer to from guesses about the communicative intent behind whole utterances."
p. 30 - imagining a mind "But we still apply the same techniques of imagining the mind behind the text, constructing a model of common ground with the author, and seeking to guess what the author might have been using the words to get their audience to understand."
p. 31 - actually great parallel "How our mind processes language can be contrasted with how we process the outputs of text-to-image models like Midjourney and DALL-E. These synthetic image machines share a lot of properties with the synthetic text machines: their function is predicated on massive data theft and profligate energy use, it’s easy to be impressed by them, and they are being used to threaten people’s livelihoods. But no one is suggesting that they are sentient—we can interpret their output (images) without imagining a mind selecting symbols in an attempt to communicate."
p. 33 - anything can be made to be computable in some sense of the word... even if its isnt accurate, in inriltrates our way of seeing the world and we start to behave more like machines "Weizenbaum believed that this impulse reinforced, rather than loosened, the grip of powerful institutions upon society, including the military, government, and corporations. Against the claims of Altman’s predecessors that AI could be a type of democratizing technology, he argued the direct opposite: computing reduced people and their experiences to data points, rather than relational and fully dimensional beings."
p. 34 - microsoft on intelligence "Microsoft’s “Sparks” paper contains a preliminary definition of general intelligence, one that has no references to fields that may have a say in such a thing, like psychology or cognitive neuroscience. Despite being a paper claiming that certain statistical models have shown the inklings of “artificial general intelligence”, it offers no well-sourced definition of what the components of general intelligence are."
p. 34 - psychology concensus of intelligence "The consensus group defined intelligence as a very general mental capability that, among other things, involves the ability to reason, plan, solve problems, think abstractly, comprehend complex ideas, learn quickly and learn from experience."
p. 35 - racist intelligence "Discussions of intelligence, pertaining to people or machines, are race science all the way down."
p. 35 - eugenics took IQ and ran with it "IQ (or “intelligence quotient”) tests31 are based on the work of early twentieth-century French psychologist Alfred Binet, whose intent was to assess which students might need additional help in the classroom. Binet militated against reducing something like intelligence into a single number, however, suggesting that the concept is too complicated to contain in one metric. It wasn’t until Binet’s work was imported to the United States that psychologists used it to justify an innate, single measure of intelligence. A trio of eugenicist scientists—Henry Goddard, Lewis Terman, and Robert Yerkes—took Binet’s scale, formalized it to be used for populations other than children, and deployed it widely."
p. 35 - insane "Yerkes used the results of the faulty test to justify an explicit racial hierarchy that placed “Nordic” white people at the top; Slavs and “darker” people of southern Europe, such as Russians, Italians, and Poles, below them; and Black people at the bottom. His theory was explicitly based on inheritance: in testing, Yerkes claimed that lighter-skinned Black people scored higher on his test; Gottfredson echoes this claiming that the fact that some Black people score higher on IQ tests can be partially attributable to “admixtures” of white blood33 possessed by Black people in the United States."
Chapter: Chapter 3: Leisure for Me, Gig Work for Thee: AI Hype at Work
p. 48 - lol.study "A white paper written by analysts at the investment bank Goldman Sachs estimates, based on data from the United States and the European Union, that a quarter of all global work could be replaced by AI tools.20 In addition, 300 million jobs worldwide could be exposed to automation, meaning part of that job could be replaced. Their methodology, however, does not inspire confidence: they rated each job task from 1 to 7 in difficulty, and simply assumed that if the task had a score of 4 and lower, that it could be automated away."
p. 50 - the craft occurs through tedious tasks "This is an argument from craft: critical thought is co-created21 with creative expression, whether that is written, spoken, or signed speech, drawing, playing music, or physical movement. For instance, qualitative sociologists are typically taught that their written memos—the writing they may undertake after reading through interview transcripts or ethnographic field notes—are really where their analysis occurs."
p. 50 - writing is thinking "Writing"
p. 53 - coding tools use coomon code. less safe "An initial security audit of that tool has shown that, because of the way language models are trained, generated code is uniquely vulnerable to common cybersecurity attacks. Researchers found in testing that 40 percent of Copilot-generated computer programs were vulnerable to some of the most common cybersecurity weaknesses. This is because code generation is made possible due to the repetition of the most common programming idioms in the training data. Those are not the most secure.27"
Chapter: Chapter 4: If It Quacks Like a Doc: AI Hype and Social Services
p. 70 - efficient at hurting families "Family separation has been a fact of the child welfare system for Black and Indigenous children from the system’s origins; the act of removing these children and placing them in the care of white parents as a means of “saving them” is part and parcel of the system’s racism.7 The use of predictive analytics in this context provides a means of automating this violence, and offers new ways to supercharge and scale it.8 What’s needed is more resources and more time for social workers to connect families to those resources. Automation in the name of efficiency here only makes the government more efficient at harming families."
p. 72 - preach "Unfortunately, that chatbot isn’t able to reliably retrieve and convey accurate information; like all LLM-based chatbots, it was designed to make shit up."
p. 82 - benchmarks are designed "In fact, every single aspect of designing an evaluation involves decisions that shape what can be measured and how those results should be interpreted. This starts with the data collection. In the example above, we can ask: Whose speech is represented? What language, what dialect, what are the ages, genders, racial and ethnic identities, social class, first languages, and other salient social aspects of the speakers? Are they talking to close friends or strangers? What are they talking about? All of these things impact the ways in which we use language and as a result also impact how far we can generalize the results of the evaluation."
p. 84 - cartoon understanding of llms "The use of standardized or professional tests or other artificial tasks in the evaluation of AI systems is a giant red flag. It typically signifies a cartoon understanding of the work that AI boosters claim their system can do; a disregard for the creativity, person-to-person connection, and care involved in the jobs they claim to replace; and a callous willingness to fob off anyone who might be dependent on the social safety net onto automated facsimiles of the services that society owes them."
Chapter: Chapter 5: Artifice or Intelligence? AI Hype in Art, Journalism, and Science
p. 101 - creativity isthe ultimate goal "To those selling the illusion of artificial intelligence and to those who think they are actually building humanlike entities, creativity stands as the ultimate goal and proof of success."
p. 102 - agreed "Either way, for these outputs to have any meaning, people still need to make sense of them and select the ones they find pleasing. Having done so, they often attribute the creativity of the people who produced the training data, combined with their own sense-making, as creativity on the part of the algorithm itself."
p. 103 - there is again a danger here - what is allowed to be a craft? data should require great oversight and human insight as to evaluate what it reflects in the real world. it needs to answer the right questions in the right way, because no data is ever objevtive "In this chapter, we talk about how the labor-saving promise of AI, when applied to creative activities such as designing visual art or making music, upsets industries based on craft."
p. 104 - it is also a world of isolation, which doesnt happen overnight. in this scenario, what can people share with others? they all want their own version, not to see something from yours "boosters like Mostaque, this is seen as a net good. They argue that these tools allow the television and movie streaming world the ability to create content that uniquely appeals to every taste, down to the person. In this environment, there’s a massive value proposition for cheaply generated AI entertainment, the benefits of which are likely to accrue to large legacy movie and television studios as well as new Big Tech entrants like Netflix and Amazon, although part of the boosters’ sales pitch is that the benefits will also be reaped by all, including independent creators"
p. 105 - no juniors. and it only needs to be good enough to look like its doing the work, not actually doing it, because then they csn hype it and CEOs csn force us to use it "AI art generators are already being deployed in ways that disrupt the economic systems through which people become and sustain careers as working artists. Illustration work for newsletters and other small publications is one way to make a living as an artist. If companies are using AI to do this work instead, we will miss out on the next generation of visual artists honing their craft and creating original content."
p. 108 - fine tuners all share the same values "were fine-tuned to generate images that were of “high visual quality.” A small group of users recruited from the Stable Diffusion Discord (an online chat server) provided one set of ratings, while another came from a forum for digital photography enthusiasts called dpchallenge.com. The top fifty users of this site provided 7.5 million ratings, and these users are overwhelmingly white and middle-class, and from small American cities. Therefore, the discernment of what is “high-quality” art is crowdsourced from a narrow group of people who likely share a similarly narrow set of worldviews. Although it would be difficult to definitively prove this, we still hypothesize that this is why AI-generated images seem to replicate one particular style"
p. 109 - bodens shuffle creativity is off "On a similar note, many defenders of AI art have argued that when humans make art we are also always “just” remixing ideas from other artworks, such that the “borrowing” (more accurately: stealing) from artists like Ortiz is justified. But there is an enormous difference between the practice of craft and the practice of writing a successful prompt: when we reference or remix ideas from"
p. 113 - lecun is also delusional "Yann LeCun28, chief AI scientist at Meta, bragged, “Type a text and galactica.ai will generate a paper with relevant references, formulas, and everything."
p. 114 - what a lit review actually is "Science and technology scholar Bruno Latour34, for instance, writes that scientists reference “the literature” as a means to justify prior knowledge, to attribute ideas, and to signal and sort themselves into particular scientific camps based on their own beliefs and values. It presupposes that one has actually read the literature, and has engaged with it in a substantive manner."
p. 114 - leClown "Galactica as able to write scientific papers, and then whining about how people were using it. He tweeted,33 “Following a text, Galactica spits out a prediction of what a scientific author might type, thereby saving time and effort. This can be very helpful even without being completely accurate. The usual disclaimer applies: garbage in, garbage out. Prompt it with lunacy, get lunacy.” This completely misses the point that LLMs are simply not suited to the task of synthesizing and presenting scientific information"
p. 119 - the street lamp "This reminds us of the parable of the person who searches for her keys under the streetlamp in the dead of night. When asked where she dropped her keys, she responds, “About five yards that way, but the streetlamp is over here.” AI-for-science makes us think we can find our keys by limiting our view to only those sidewalks illuminated by the glow of the data centers powering it."
p. 120 - deepmind net negative... science is human "projects like Google DeepMind’s research on crystal materials48 is rife with allusions to the potential benefits for such important causes as better solar panels. But they actually aren’t as beneficial to the scientists working in the relevant fields as the advertising copy would have it. In a paper49 written by two materials scientists, they found that a closely examined subset of Google’s new materials did not meet the criteria for being useful. That work is slow and difficult. Rather than speeding up science, Google DeepMind is flooding the search space with candidates of unknown promise. Tools like DeepMind’s may have potential for doing large-scale pattern"
Chapter: Chapter 6: I’m Sorry, Dave, I’m Afraid I Can’t Do That: AI Doomers, AI Boosters, and Why None of That Makes Sense
p. 139 - the ideologois are scary "Should you be scared of an autonomous AI agent? No. But you should be wary of the alarming ideologies behind both AI Doomerism and Boosterism. Doomerism/Boosterism serves to obscure, rather than illuminate, what’s at stake when it comes to the current AI boom. Moreover, these technologies are accelerating the real existential threat of human-made climate change, cutting into our already too-thin margin of time to mitigate it."
p. 141 - western, rich nerds "Strangely enough, despite these visions, nearly all AI Doomers think that AI development is a net good. Many of them have built their careers off the theorization, testing, development, and deployment of AI systems. They have markedly not complained about the ways in which the business practices around AI have accrued financial, social, and political power to a very select few. They are mostly men. They are almost universally white and Western."
p. 142 - AI safety "research" "This Doomer/Booster research field is called “AI safety”. Despite the name, this work does not come out of systems safety engineering7—which is a real field, with first principles, requirements, and engineering specifications—but instead from a truncated and misunderstood application of those ideas."
p. 144 - these tools are horrible.. all the things they do "Tools sold with outlandish claims to functionality, such as detecting whether someone is lying to an immigration official or identifying countries of origin based on DNA, are routinely tested and deployed on African migrants fleeing war, famine, and genocide. Even away from any border, within the U.S., automated systems tear apart16 Black and Indigenous families by flagging these families for family separation—forcible removal of kids into the “care” of the state—at a much higher rate than white families."
p. 153 - hinton is lost "Geoff Hinton,41 one of the so-called “Godfathers of AI” whom we met in Chapter 1, framed his Doomerist concerns to CNN journalist Jake Tapper by saying “there are very few examples of a more intelligent thing being controlled by a less intelligent thing.” We wonder when the last time was that Dr. Hinton suffered from any kind of food poisoning, and if he then decided that bacteria are more intelligent than humans."
p. 153 - cant measure what we cant define. cant design it either intelligence based on racism and eugenics, bad start "In order to evaluate “intelligence”, we’d first need a clear and operationalized definition of the concept, along with a compelling narrative of how it relates to the Doomer/Booster claims. Then we’d need a way to test for it, such that we can show that the test is actually measuring the property of interest. We saw in Chapter 2 that, throughout their history, measures of “intelligence” in humans have been based not in sound science but in eugenics and racism, so that’s already a bad start."
p. 154 - Turing tried to dodge intelligence "Can machines think?” with the question of whether the computer-pretending-to-be-a-man can fool an interrogator more consistently than the man-pretending-to-be-a-woman. The issue here, of course, is that without a definition of “intelligence” the question of how to measure “artificial intelligence” is a nonstarter. Turing tried to sidestep this with his imitation game but ended up proposing a setup that ran afoul of the very human tendency to make sense of language by imagining a mind behind it."
p. 154 - drodophilia of AI "Soviet mathematician Aleksandr Kronrod’s phrase “chess is the Drosophila of AI”44 because, he believed, the game was an efficient way to do rapid experimentation towards a long-term goal. That is, just as the fruit fly (Drosophila) is a boon to a certain kind of study in genetics (thanks to its short generations and ease of handling in lab environments), chess was thought to be a fertile ground in which to develop computer intelligence because it allowed researchers to quickly try out different approaches. This argument, of course, presupposes the relevance of chess to intelligence."
p. 155 - good metaphor "We have become rather fond of the Drosophila metaphor for another reason: it epitomizes the fixation of researchers on one small contained problem—with its attendant issues and its own social history—while claiming to make progress towards grander goals."
p. 155 - feel the AGI "Ilya Sutskever, when he was at OpenAI, urging employees to “feel the AGI,”47 instructing them to chant it like a mantra. And if they can’t even define it, there’s no way that they are capable of measuring it."
p. 156 - books have a mind behind the words, not the book itself. with ai there is simply no mind "With books, there is a mind behind the language, but the mind does not reside in the book. With the Magic 8 Ball or a chatbot, there’s an element of randomness in what language comes out. But that randomness does not constitute a mind."
p. 157 - also as boden says ... its about the virtual machine, no amount of hardware solves this .. in the hypothetical scenario it were possible "Instead, what we see is a mad dash towards ever-larger models that require increasing amounts of computation (and therefore energy consumption) to train and use—with real and measurable environmental impacts."
p. 159 - ai slop is costly "Another study by Luccioni and colleagues estimated that for every two images you generate using a large text-to-image model, it’s like charging your phone fully"
p. 159 - ai overview thirsty "The “AI Overviews” feature that Google added to search results in 2024 likely consumes 30 times more energy per query than just returning links"
p. 160 - the most privileged make all decisions "For a tech baron to confront these harms would require them to take their own culpability seriously and contemplate their own privilege. It’s far more comfortable to sit back and pontificate on imagined scenarios where they are (also) victims, and dismiss anything else as “less existential.”71 Politics"
Chapter: Chapter 7: Do You Believe in Hope After Hype?
p. 171 - friction is good "But it’s the second level that we want to address here: friction in information access is actually not only beneficial, but critically important."
p. 178 - datasets strengthen their biases "themselves. This work was motivated by a flurry of research showing that statistical models are very effective at modeling the biases in the data, and amplifying those biases in their output.35 Models trained on existing data contain a representation of the patterns of the past,36 including the effects of discrimination of all kinds. If we want to avoid replicating those patterns into the future, then we need to know what patterns any given model has been trained on before deciding how and whether to use it"
p. 179 - llms dont learn "claims that a machine may have “emergent” properties, like Google CEO Sundar Pichai’s claim that their chatbot “learned” Bengali38 without being specifically trained on it. On the face of it, this claim is ridiculous. But researchers and regulators can confirm this only by having access to the data, or at least thorough documentation of it"
p. 182 - computers cannot solve tasks that require wisdom "Computers can make judicial decisions, computers can make psychiatric judgments. They can flip coins in much more sophisticated ways than can the most patient human being. The point is that they ought not be given such tasks. They may even be able to arrive at “correct” decisions in some cases—but always and necessarily on bases no human being should be willing to accept. There have been many debates on “Computers and Mind.” What I conclude here is that the relevant issues are neither technological nor even mathematical; they are ethical. They cannot be settled by asking questions beginning with “can.” The limits of the applicability of computers are ultimately statable only in terms of oughts. What emerges as the most elementary insight is that, since we do not now have any ways of making computers wise, we ought not now to give computers tasks that demand wisdom."
p. 189 - agreed. specific AI is awesome "That means we reject untestable “everything” machines that are marketed as “general purpose technologies”. Instead, we want to see specific tools geared towards specific tasks. An example of a specific task for image processing is determining which parts of a photo should be in focus. An example of a specific task for language processing is machine translation. Considering what kind of data is required to build such applications brings us to our second guideline."
p. 190 - tech needs to reflect our values "If we are to create a future that is populated with technologies we want, we “can’t only critique the world as it is,” as science and technology scholar Ruha Benjamin has written; we also “have to build the world as it should be to make justice irresistible.”72 Part of that vision means technology ought to be created with full participation of the people it impacts. Following disability justice advocates73, we say “nothing about us, without us."
p. 191 - just one more data center... the future is perpetually near "In 2023, Elon Musk said77 his Teslas would have Full Self-Driving mode by the end of that year. Geoff Hinton proclaimed78 in 2016 that we should stop training radiologists, because AI systems will soon be able to read medical images better than a human technician can. Meanwhile, rather than providing self-driving cars, Tesla has given us false advertising about self-driving mode, resulting in hundreds of crashes79, and the Bureau of Labor Statistics has projected80 a 6 percent growth in medical imaging jobs (including radiologists) from 2022 to 2033, faster than the average across all industries. The AI project has always been more fantasy than reality, starting from the field’s origins and the optimistic salesmanship of Minsky, McCarthy, and friends during that 1956 summer at Dartmouth. Jenna Burrell, a science and technology scholar and former director of research at Data and Society, has called this81 the “ever-receding horizon of the future,” which creates an urgency and a sense of hype about the next promised technology, while Anna Lauren Hoffmann has critiqued82 the metaphor of AI systems being in their “childhood”—we only need teach them better!"
p. 193 - measly productivity gains "Daron Acemoglu, labor economist at MIT,89 published an estimate that productivity gains from AI will be less than 0.53 percent over the next ten years."
p. 195 - it just needs to look like it works "Generative AI has the potential to ruin what were stable careers—not because the technology can do the work effectively or as skillfully, but it can produce convincing enough synthetic media to make certain jobs seem to managers redundant or requiring less skill than before. After"
Added by: alexb44 Last edited by: alexb44
|