Sound Research WIKINDX |
![]() |
|
Silver, D., & Sutton, R. S. (2025). Welcome to the era of experience. Google AI, 1. Added by: Mark Grimshaw-Aagaard (18/11/2025, 09:18) Last edited by: alexb44 (10/09/2026, 05:35) |
| Resource type: Journal Article Language: en: English Published BibTeX citation key: Silver2025 Email resource to friend View all bibliographic details |
Categories: AI/Machine Learning Keywords: Artificial creativity, Artificial Intelligence Creators: Silver, Sutton Publisher: MIT Press (Cambridge, Massachusetts) Collection: Google AI, Resources citing this (Bibliography: WIKINDX Master Bibliography) |
Views: 8/202
|
| Abstract |
|
"We stand on the threshold of a new era in artificial intelligence that promises to achieve an unprecedented level of ability. A new generation of agents will acquire superhuman capabilities by learning predominantly from experience. This note explores the key characteristics that will define this upcoming era."
Added by: Mark Grimshaw-Aagaard Last edited by: Mark Grimshaw-Aagaard |
| Notes |
|
PDF preprint - will be a chapter for book "Designing an Intelligence" (not all book details available).
Quite utopian. ————— Alex notes ————— Summary: Human data will not continue to improve AI, evidenced by progress slowing already. Proposes a new source of data - experience. Agents will act in the real-world (autonomously), planning long term which allows an agent to achieve long term goals and to know how to prioritise. Introduce the concept of grounded rewards - rewards unaffected by human prejudgment that arise directly in the environment and thus relate directly to the machine. Claim that these advancements will lead to not only AGI, but ASI. My thoughts: Provides no substance to as to how robots will begin to experience, thus it's a grossly reductive understanding of experience. It is essentially a utopian view of singularity or radicalist movement views. Leading CS researchers frequently present these utopian or dystopian views of what AI can be, but never explain exactly how that will happen. They lean on cognitive science by reducing concepts like intelligence, intentionality, reasoning, etc. to meaningless analogies in computer science that seems to defend their position, but it is merely academic sleight of hand. Added by: Mark Grimshaw-Aagaard Last edited by: alexb44 |
| Quotes |
|
"To progress significantly further, a new source of data is required. This data must be generated in a way that continually improves as the agent becomes stronger; any static procedure for synthetically generating data will quickly become outstripped. This can be achieved by allowing [AI/LLM] agents to learn continually from their own experience, i.e., data that is generated by the agent interacting with its environment. AI is at the cusp of a new period in which experience will become the dominant medium of improvement and ultimately dwarf the scale of human data used in today’s systems."
Added by: Mark Grimshaw-Aagaard
(18/11/2025, 09:29)
|
|
"Human-centric LLMs typically optimise for rewards based on human prejudgement: an expert observes the agent’s action and decides whether it is a good action, or picks the best agent action among multiple alternatives. For example, an expert may judge a health agent’s advice, an educational assistant’s teaching, or a scientist agent’s suggested experiment. The fact that these rewards or preferences are determined by humans in absence of their consequences, rather than measuring the effect of those actions on the environment, means that they are not directly grounded in the reality of the world."
Added by: Mark Grimshaw-Aagaard
(18/11/2025, 09:38)
|
|
Super intelligence == Garmin watch?
"Powerful agents should have their own stream of experience that progresses, like humans, over a long time-scale. This will allow agents to take actions to achieve future goals, and to continuously adapt over time to new patterns of behaviour. For example, a health and wellness agent connected to a user’s wearables could monitor sleep patterns, activity levels, and dietary habits over many months. It could then provide personalized recommendations, encouragement, and adjust its guidance based on long-term trends and the user’s specific health goals." It genuinely astounds me that they can speak of super intelligence and in the same breath mention a health and wellness agent that is basically a second generation garmin watch.
Added by: alexb44
(10/09/2026, 05:14)
|
|
These apparent examples of experience based data is literally user input.
"Grounded rewards may arise from humans that are part of the agent’s environment. For example, a human user could report whether they found a cake tasty, how fatigued they are after exercising, or the level of pain from a headache, enabling an assistant agent to provide better recipes, refine its fitness suggestions, or improve its recommended medication. Such rewards measure the consequence of the agent’s actions within their environment, and should ultimately lead to better assistance than a human expert that prejudges a proposed cake recipe, exercise program, or treatment program." Their logic of grounded rewards in this case literally relies on user input, it's the same as before. A grounded reward would need to have actual meaning to the robot itself - their example implies programming, which means human prejudgment is present. They go on: "In fact, the world abounds with quantities such as cost, error rates, hunger, productivity, health metrics, climate metrics, profit, sales, exam results, success, visits, yields, stocks, likes, income, pleasure/pain, economic indicators, accuracy, power, distance, speed, efficiency, or energy consumption. In addition there are innumerable additional signals arising from the occurrence of specific events, or from features derived from raw sequences of observations and actions." Again, these are all data points used by humans to describe the world... "error rates", "climate metrics"... they have inherent human bias by definition, there is nothing natural in these examples that would imply they arise from the environment itself, because they all relate to the human relation to the environment. It all comes back to the extremely limited and faulty inclusion of the concept of experience.
Added by: alexb44
(10/09/2026, 05:18)
|
|
Where LLMs can act as a universal computer, "the principle of a universal computer only addresses the internal computation of the agent; it does not connect it to the realities of the external world. An agent trained to imitate human thoughts or even to match human expert answers may inherit fallacious methods of thought deeply embedded within that data, such as flawed assumptions or inherent biases. For example, if an agent had been trained to reason using human thoughts and expert answers from 5,000 years ago it may have reasoned about a physical problem in terms of animism; 1,000 years ago it may have reasoned in theistic terms; 300 years ago it may have reasoned in terms of Newtonian mechanics; and 50 years ago in terms of quantum mechanics. Progressing beyond each method of thought required interaction with the real world: making hypotheses, running experiments, observing results, and updating principles accordingly. Similarly, an agent must be grounded in real-world data in order to overturn fallacious methods of thought. This grounding provides a feedback loop, allowing the agent to test its inherited assumptions against reality and discover new principles that are not limited by current, dominant modes of human thought. Without this grounding, an agent, no matter how sophisticated, will become an echo chamber of existing human knowledge."
Added by: Mark Grimshaw-Aagaard
(18/11/2025, 09:45)
|
| Musings |
|
Piling on to your musing above, Mark... Why would a robot even care about an extra sensor or extra RAM? From More Everything Forever, quoting Ceglowski: "Cegłowski, who was born in Poland, says the idea that a superintelligent being would inevitably want to improve itself is “unabashedly American.” “My roommate was the smartest person I ever met in my life. He was incredibly brilliant, and all he did was lie around and play World of Warcraft between bong rips,” he says. “The assumption that any intelligent agent will want to recursively self-improve, let alone conquer the galaxy, to better achieve its goals makes unwarranted assumptions about the nature of motivation.” (Becker 2025, p.112) Becker, A. (2025). More Everything Forever: AI Overlords, Space Empires, and Silicon Valley's Crusade to Control the Fate of Humanity. New York: Basic Books.
Added by: alexb44
(10/09/2026, 05:35)
|