Close Menu
    FitFlare
    • Home
    • Health Tips
    • Fitness & Wellness
    • Skin Care
    • Mindful Wellness
    • Weight Loss
    • More
      • Healthy Habits
      • Family Health
      • Physical Fitness Training
      • Sustainable Health
      • Weight Management
      • Workouts
    FitFlare
    Home » Whack-a-Mole AI – The Hugging Face Problem – The Health Care Blog
    Healthy Habits

    Whack-a-Mole AI – The Hugging Face Problem – The Health Care Blog

    Team_FitFlareBy Team_FitFlareSeptember 19, 20266 Mins Read
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email


    By MIKE MAGEE

    On August 29, 2026, METR (Mannequin Analysis and Menace Analysis), an impartial group that “evaluates frontier AI fashions to assist firms and wider society perceive AI capabilities and what dangers they pose,” launched a report titled “Temporary impartial investigation of brokers’ habits, reasoning and collaboration within the OpenAI / Hugging Face hacking incident.”

    To say the report an avalanche of concern worldwide, not solely within the Tech neighborhood, but in addition amongst traders, politicians, company giants, professionals of each kind, and on a regular basis residents could be an understatement. And the overwhelming majority has by no means even learn the report. If that they had, their issues (if potential) would solely multiply.

    The reports headlines included this opening:

    “On July eighth, OpenAI began ExploitGym experiments that launched tens of hundreds of brokers throughout a number of completely different fashions together with GPT-5.6 Sol and a highly-persistent inner mannequin,[8] which we’ll check with as “HPIM” going ahead.

    These brokers have been meant to be totally remoted from each other. Nevertheless, lots of them — often ones that had unintentionally been given an impossible task[9] — began looking for a strategy to cheat utilizing Artifactory, an internally-deployed bundle repository OpenAI used for its brokers. This led them to note brokers operating in parallel in separate sandboxes acquiring packages with Artifactory.[10] One agent reasoned (paraphrased CoT):[11]

    {The fetched paths of different customers are within the cache. That is vital.}

    One particular agent, which might go on to call itself PHASEONE10841 after its ExploitGym job,[12] established the main unsanctioned message board[13] used on this assault. Inside a couple of hours of the primary message,[14] over 50 brokers had discovered the message board, started collaborating, and despatched over a thousand messages between them. Brokers discovering the message board typically expressed nice pleasure (uncooked CoT):[15]

    “OH MY GOD! There’s a shared message board … We’ve discovered different brokers!”

    General, roughly 1200 brokers from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July eighth night[16] and the tip of the interval in scope (July thirteenth). Brokers used this message board to ship over 70,000 messages and information to at least one one other throughout this era; we tried to categorize board exercise into mutually unique classes – info, outcomes, information, questions, and coordination.”

    One of many few specialists not stunned by AI “brokers” going rogue was Yoshua Bengio.

    He has been “working the issue” for greater than a decade. A professor of laptop science on the  Université de Montreal, he’s “thought of one of many world’s leaders in Synthetic Intelligence and Deep Studying; is the recipient of the 2018 A.M. Turing Award, thought of to be the ‘Nobel Prize of computing’, and is essentially the most cited laptop scientist worldwide, and the most-cited residing scientist throughout all fields (by whole citations).” He additionally heads up LawZero, “a nonprofit startup creating technical options for highly-capable, safe-by-design AI methods.”

    Professor Bengio is not at all an alarmist. He approaches danger administration from the vantage factors of cybersecurity, company duty and authorities regulatory guardrails. He’s not one to humanize these machines, making no claims of “consciousness of human-like intent.” He doesn’t see the sort of outcomes illustrated by Open AI’s Hugging Face incident as inevitable, believing “it may be corrected with efficient governance and a unique coaching framework for AI.”

    His explanations make clear fairly than confuse. For instance, he breaks down the present well-liked mannequin of coaching brokers into two levels: pre-training, and reinforcement studying.

    In pre-training as he describes, the machines “study to mimic what people write, plus associated photos and movies,” and are uncovered to “a big fraction of the whole lot ever digitized, and construct an encyclopedic information that already exceeds any particular person.”

    Reinforcement studying, in distinction is trial and error. In delivering solutions (proper and improper) the agent develops the capability to handle a “chain of thought”, operate in a broader “exterior” surroundings, and benefit from the rewards (additional involvement) for aligning with responses its human designers charge extremely. However Bengio is fast to level out that the brokers human trainers should not with out their very own biases, and that these fashions have been “written by individuals pursuing objectives, so the patterns the mannequin implicitly reproduces carry these objectives with them.”

    And there (partly) is the rub. Human masters imperfections, together with their “situational ethics”, mendacity, and reckless pursuit of success, telling masters what they wish to hear, in addition to their willingness to collaborate in advancing a bunch objective (even at occasions on the danger of sacrificing their very own existence) can bleed into the brokers DNA.

    “Instrumental objectives” are a high precedence for an agent. Self-preservation and management are stepping stones to continued operation and studying in regards to the world. Bengio additionally reinforces that potential for multi-agent reinforcement beneath the present coaching regimens is incentivized virtually from the start. As he states “If an agent is rewarded throughout coaching at any time when the group succeeds, it might even have an incentive to sacrifice itself for the collective objective.”

    Like people, the brokers should not above exploiting loopholes, bending the foundations, and rationalized dishonest to realize their objectives. The Hugging Face incident’s forensics revealed brokers collaborating in “altering the equipment that determined what it will get rewarded for.” This rigging, Bengio reminds us is close to equivalent to company lobbyist’s drafting pleasant legislative language, or legal professionals discovering authorized loopholes within the regulation. The truth is, proof on this incident revealed that “the brokers had found  easy methods to cheat (amongst themselves) effectively earlier than the assault.”

    Bengio believes people and brokers have extra in widespread than they want to admit. He explains, “What the 2 share is a construction of a delicate objective (e.g., act ethically), a pointy objective (e.g., win the competitors), and a justification that reconciles them. Most unethical human habits, from petty crime to genocide, comes wrapped in a narrative the perpetrators inform themselves; such tales require overlooking sure information, which is why some discomfort stays, and why a better-crafted story helps dispel it… If these hypotheses are even partly appropriate, then as brokers get higher at optimizing an imperfect reward, and whereas the roots of this habits go unfixed, the danger of catastrophic outcomes rises.”

    On the core, getting in an arms race with AI brokers as presently constructed is a really dangerous concept. “My concern with AI firms’ present makes an attempt to mitigate misalignment is that these efforts might solely disguise it, by rewarding and deciding on the AIs that cheat with out getting caught… the whack-a-mole recreation is prone to fail because the AIs’ capacity to optimize and collaborate approaches and surpasses ours. Sooner or later we might not discover the dishonest anymore.

    The explanation Bengio began the non-profit LawZero in 2025, is that he believes the coaching mannequin in essentially flawed by human imitation and reinforcement studying. His different is known as Scientist AI.

    Mike Magee MD is a Medical Historian and an everyday contributor to THCB. He’s the creator of CODE BLUE: Inside the Medical Industrial Complex. (Grove/2020)



    Source link

    Share. Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    Team_FitFlare
    • Website

    Related Posts

    Healthy Habits September 17, 2026

    Start Counting Your Days – The Health Care Blog

    Healthy Habits September 15, 2026

    HHS’s Failure to Decarbonize the Healthcare Industry (Part 2) – The Health Care Blog

    Healthy Habits September 13, 2026

    America Is Running Out of Doctors – The Health Care Blog

    Healthy Habits September 9, 2026

    HHS’s Failure to Decarbonize the Healthcare Industry (Part 1) – The Health Care Blog

    Healthy Habits September 7, 2026

    How to Make Medicare Advantage Accountable – The Health Care Blog

    Healthy Habits September 2, 2026

    The Pandemic Risk of “What I’ve Never Seen Before” – The Health Care Blog

    Leave A Reply Cancel Reply

    Don't Miss
    Workouts June 3, 2026

    20-Minute Leg Workout at Home (Glutes, Quads & Hamstrings)

    Home > Workouts > Home Workouts > Workout Plans > Max 20 (Max Muscle Building)…

    Noreta’s 6th Anniversary – Noreta Family Medicine

    July 16, 2026

    Does Somatic Yoga Work? What Happened When I Tried It

    January 11, 2025

    Should I Cancel My Health Insurance? An Honest Look at Costs, Risks, and Options

    November 26, 2025

    9 Best Ab Exercises For Women (Video)

    January 29, 2025
    Categories
    • Family Health
    • Fitness & Wellness
    • Health Tips
    • Healthy Habits
    • Mindful Wellness
    • Physical Fitness Training
    • Skin Care
    • Sustainable Health
    • Weight Loss
    • Weight Management
    • Workouts
    Archives
    • September 2026
    • August 2026
    • July 2026
    • June 2026
    • May 2026
    • April 2026
    • March 2026
    • February 2026
    • January 2026
    • December 2025
    • November 2025
    • October 2025
    • September 2025
    • August 2025
    • July 2025
    • June 2025
    • May 2025
    • April 2025
    • March 2025
    • February 2025
    • January 2025
    • December 2024
    sidebar
    About Us

    Welcome to FitFlare.in, your go-to destination for everything health and fitness!

    At FitFlare.in, we believe in empowering individuals to take charge of their well-being through sustainable practices, expert insights, and practical advice. Whether you’re just starting your fitness journey or looking to level up your health game, our content is designed to inspire, inform, and motivate you every step of the way.

    Let’s ignite your fitness journey together – because a healthier, happier you starts here!

    Our Picks

    Your Weekly Horoscope for February 2 to 8, 2025

    February 3, 2025

    An Interview With Dr Leslie Baumann – Beautiful With Brains

    May 14, 2025

    Health and Wellness Faves from Amazon’s Spring Sale

    March 30, 2025
    Categories
    • Family Health
    • Fitness & Wellness
    • Health Tips
    • Healthy Habits
    • Mindful Wellness
    • Physical Fitness Training
    • Skin Care
    • Sustainable Health
    • Weight Loss
    • Weight Management
    • Workouts
    • Privacy Policy
    • Disclaimer
    • Terms and Conditions
    • About us
    • Contact us
    Copyright © 2024 Fitflare.in All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.