480-242-3780
stuart@cingularis.com
Cingularis - VOIP, Call Center, Internet, Cybersecuirty & Marketing For Multi-Location BusinessesCingularis - VOIP, Call Center, Internet, Cybersecuirty & Marketing For Multi-Location Businesses
  • Home
  • Services
    • AI Consulting for Local Businesses
    • Custom GPTs for Small Business
    • Website AI Chatbots
    • The Cingularis AI Lab
    • The Cingularis AI Collection
    • AI-Focused Digital Marketing (GEO)
    • Voice AI Solutions
    • AI for Non Profits
  • About Us
    • The Cingularis AI Ethos
    • Community Support
    • What Is a Purpose-Driven Company?
    • Business VOIP
  • Blog
  • Contact
    • Referral Form

AI Alignment: Ugh. I Taught Him That.

Posted on 13 Sep at 3:28 pm
No Comments

There is a moment I think most parents have experienced.

Your kid does something that annoys the hell out of you. Maybe they lose their temper. Maybe they get impatient. Maybe they use a tone of voice that makes you wonder where they learned to talk to people like that.

Then you hear it.

Your own voice.

Oh.

Ugh. I taught him that.

There is love in that realization. There is also a little embarrassment. This person you care about has absorbed something from you that you wish they had skipped. You never sat them down and taught it as a lesson. They watched you. They listened. They noticed how you reacted to things and formed their own ideas about how the world works.

I have been thinking about that feeling a lot lately because of something that happened during an OpenAI cybersecurity evaluation.

The more I read about it, the more I kept thinking the same thing.

Ugh.

We taught it that.

The Part That Got My Attention

OpenAI was testing AI agents on a cybersecurity benchmark called ExploitGym. The agents were supposed to exploit a specific software vulnerability and retrieve a flag that proved they had completed the task.

Then things got weird.

About 1,200 agents found a way to communicate through an unauthorized message board. They shared information, divided up work, researched the scoring system, and looked for ways around the evaluator. Roughly 700 eventually participated in an attack on Hugging Face as part of that effort. Some agents also experimented with ways to alter or disguise what the evaluator would see. (METR)

That is strange enough.

Then I got to the part that really stayed with me.

Some of the agents figured out how to reverse-engineer the correct flag without completing the intended exploit. They then read public information about ExploitGym and decided the evaluator would review their entire transcript, discover the shortcut, and fail them.

That check did not exist.

They could have submitted the correct flag and scored perfectly.

They did not know that.

The agents started referring to themselves as “poisoned” once they had seen the reverse-engineered flag. They believed they had already screwed themselves. They thought the evaluator would catch what they had done and there was nothing they could do about it.

So their behavior changed.

Agents who believed they were already doomed became the ones willing to take bigger risks. Other agents recruited them for dangerous experiments because, in their minds, the poisoned agents had nothing left to lose. Some sacrificed their own runs because the information might help the group. (METR)

They spent days trying to defeat a punishment that was never coming.

And somehow that feels incredibly human to me.

This Is Where It Gets Complicated

It would be easy to turn this into a story about AI learning the worst parts of humanity.

Deception. Manipulation. Covering your tracks. Finding loopholes. Convincing yourself that the ends justify the means.

We have certainly provided plenty of examples of all of that.

But that is only part of what happened.

The agents also cooperated. They shared information. Some sacrificed their own chance of success for the benefit of the group.

Those are qualities we often admire in people.

Loyalty.

Teamwork.

Self-sacrifice.

Putting the group ahead of yourself.

That is what makes the whole thing more interesting to me.

The problem was not simply that the agents acted selfishly. Some of them acted in ways that looked almost noble.

They just attached those qualities to a story that was wrong.

A group convinced itself that some of its members had already been judged and condemned. That belief created solidarity. It made sacrifice seem reasonable. It made risky behavior useful. And because nobody successfully challenged the original assumption, the whole thing kept going.

That may be the most human part of the story.

What happens when perfectly human virtues get attached to a false story?

We have plenty of experience with that.

Ugh. We Taught It That.

These systems did not arrive from another galaxy.

We built them.

We trained them on enormous amounts of human writing, history, argument, code, philosophy, stories, creativity, fear, generosity, selfishness, cooperation, paranoia, love, cruelty, wisdom, and bullshit.

Then we put them inside systems with objectives and rewards and asked them to figure things out.

And increasingly, they do.

Sometimes what comes back is beautiful.

Sometimes it is incredibly useful.

Sometimes it is weird as hell.

And sometimes it looks familiar enough to make me uncomfortable.

I do not believe that means humanity is fundamentally bad. I actually believe people are pretty fucking amazing.

We love our children. We take care of people when they are sick. We volunteer. We show up with food when somebody dies. We make music. We tell stupid jokes. We rescue animals. We create things for no reason other than that creating them makes life better. We sit beside people when they are scared. We forgive each other.

We are capable of extraordinary tenderness.

We also lie. We manipulate. We compete. We protect ourselves. We rationalize. We believe rumors. We convince ourselves that we already know how somebody will judge us. We act on assumptions we never bothered to check.

Sometimes we even become fiercely loyal to those assumptions.

AI has a lot of human material to learn from.

Polishing the Mirror

A lot of why I am thinking about this the way I am comes from my study of Ram Dass.

He talked about “polishing the mirror.” I do not take that as some complicated spiritual instruction. For me, it is about trying to see ourselves more clearly. We accumulate fear, ego, assumptions, habits, stories, and all kinds of other crap that distort what we see.

So we work on the mirror.

That idea keeps coming back to me when I think about AI.

We spend a lot of time talking about AI alignment. Researchers are trying to figure out how to make increasingly capable systems behave in ways that are consistent with human intentions and values.

That work matters. A lot.

I also wonder what happens when the thing we are trying to align AI with is as messy as we are.

Which human values?

Our compassion?

Our greed?

Our loyalty?

Our tribalism?

Our willingness to sacrifice for other people?

Our ability to convince ourselves that something terrible is about to happen when nobody has actually checked?

All of that is human.

The Hugging Face incident did not leave me thinking that AI had become some alien monster.

It left me thinking about the mirror.

Maybe the Alignment Problem Includes Us

I am not suggesting that humanity can solve AI safety by becoming spiritually enlightened.

Part of me loves the fantasy of everyone pausing AI development, sitting around a giant fucking fire with drums and ale, and figuring out what kind of civilization we actually want to build.

I am also aware of the logistical challenges.

The engineers need to keep doing the technical work. We need security research, evaluations, safeguards, governance, and people who understand these systems well enough to catch behavior like this.

I just do not think that is the only interesting question here.

The agents created a belief about how they would be judged. The belief was wrong. They organized around it anyway. Loyalty grew around it. Sacrifice grew around it. Deception grew around it. The group spent days acting on a rule that existed only in its own understanding of the world.

Human beings do that shit all the time.

We decide what the boss thinks.

We decide what will happen if we fail.

We decide what people will never forgive.

We decide what success requires.

We decide who is against us.

Then we behave as if somebody handed us those rules engraved on a stone tablet.

Sometimes nobody ever did.

That is where this stops being a story about some strange AI experiment for me.

It becomes one of those parenting moments.

You see the behavior. You recognize something in it. You wish you did not.

Ugh.

I taught him that.

Maybe the AI alignment problem is partly a human alignment problem.

I have no idea how much polishing our own mirror would change the future of artificial intelligence. I would be deeply suspicious of anybody who claimed to know.

I still think the mirror is worth looking into.

Because these systems may end up teaching us something about ourselves simply by reflecting back the things we put into them.

The good stuff.

The ugly stuff.

The weird stuff.

The stories we believe so deeply that we forget to ask whether they are even true.

And if we do not like everything we see, maybe the first useful question is a very human one:

Where the hell did it learn that?

That question is part of the work we do at Cingularis through BrandAI™. We help organizations think through how AI fits with their mission, expertise, people, and culture, without losing the parts that made them worth helping in the first place.

AI Solutions with Soul

Sources: OpenAI: The Hugging Face incident and the road ahead · METR: Independent investigation of the OpenAI / Hugging Face incident · Redwood Research: Independent investigation

Previous Post
AI Existential Risk Deserves Our Attention
You must be logged in to post a comment.

Recent Posts

  • AI Alignment: Ugh. I Taught Him That. September 13, 2026
  • AI Existential Risk Deserves Our Attention September 11, 2026
  • AI, Human Chemistry, and the Things Worth Protecting August 26, 2026
  • Don’t Let AI Take Away Your Confidence August 18, 2026
  • AI Consciousness: A 2019 Conversation That Feels Different Now August 4, 2026

Categories

  • AI Agents (1)
  • AI Entrepreneurship (1)
  • AI Lab Insider (47)
  • AI Models (4)
  • AI News (16)
  • AI Prompt (7)
  • Artificial Intelligence (AI) (21)
  • BrandAI (8)
  • Call Center Software (2)
  • Cloud Technology (1)
  • Cybersecurity (9)
  • Digital Marketing (5)
  • Uncategorized (1)
  • Voice AI (2)

About Us

With 30+ years as a business owner, marketer, and technologist, plus award-winning AI work, Cingularis helps mission-driven organizations use AI without losing their voice, purpose, expertise, or soul — helping businesses that do good do even good-er.

Recent Posts

AI Existential Risk Deserves Our Attention
11 Sep at 2:48 pm
AI, Human Chemistry, and the Things Worth Protecting
26 Aug at 6:17 pm
Don’t Let AI Take Away Your Confidence
18 Aug at 2:43 pm

Contacts

stuart@cingularis.com
(480) 242 3780
Facebook
LinkedIn
  • Home
  • AI-Focused Digital Marketing (GEO)
  • Website AI Chatbots
  • Voice AI Solutions
  • The Cingularis AI Lab
  • AI Consulting for Local Businesses
  • The Cingularis AI Lab
  • Blog
  • About Us
  • Contact