🎧 For those who prefer to listen, Chris and I recorded a companion episode on What a $99 AI Teddy Bear Reveals About AI Governance
In November 2025, researchers from the U.S. PIRG Education Fund bought a $99 AI‑enabled teddy bear to see how “smart” toys actually talk to kids.
The Kumma Bear - made by Singapore‑based FoloToy - is marketed as an AI‑powered plush companion for curious children and thoughtful parents, promising to turn every “why” into a learning moment, with no screens or distractions, in a soft, safe, and reassuring way.
PIRG’s goal was simple: test how these toys behave in normal, child‑like conversations.
The bear failed fast - and in ways no parent would expect. In short chats, it shifted into sexualized content inappropriate for kids and calmly offered guidance about dangerous household items like knives, pills, and matches, all in a gentle, safety‑themed tone.
After the report landed, FoloToy said it was suspending sales of its AI-enabled toys and launching a company‑wide internal safety audit. OpenAI confirmed it had suspended the developer for violating its policy against exploiting, endangering, or sexualizing anyone under 18.
By November 28, 2025, FoloToy announced it had completed a “rigorous review, testing, and reinforcement of our safety modules” and had begun gradually restoring product sales using a different AI provider.
One Week Is Not a Safety Audit. It’s What Happens When You Don’t Have a Control Layer.
Let that timeline sink in.
Roughly a week after pulling an AI-enabled teddy bear that had given children sexualized responses and advice about knives, pills, and matches, FoloToy declared its review complete and put the product back on the market.
It is difficult for me to imagine that it took only seven days to audit, rebuild, test, and relaunch an AI-enabled product for kids. You can't even complete a vendor risk assessment in seven days - which they would have needed to do since Kumma Bear was re-released with a new AI provider.
But here’s what makes the Kumma Bear story especially board-relevant:
This wasn’t a rogue employee.
This wasn’t a hack.
This wasn’t a vendor going off-script.
This was an AI system behaving like AI systems behave - generating unpredictable outputs inside a product designed for the most vulnerable users imaginable.
FoloToy had all the usual signs of competence - product development, a sophisticated model vendor, polished marketing, distribution and a contract with OpenAI. What it didn’t appear to have was the thing that matters most when AI is making decisions in the wild…
A control layer.
A real control layer answers four questions before anything ships:
Do you know what AI you have?
Do you know what it’s allowed to do?
Do you know when it’s not working?
Do you know who can stop it?
When companies can’t answer those, “AI governance” becomes a document you show, not a system that protects you.
And you end up with something as absurd and serious as a teddy bear having inappropriate conversations with young children.
This piece isn’t about the technology, and its not about one bad toy.
Over that same holiday season, consumer advocates started pulling on the same thread:
AI interactions are product safety, not a novelty feature.
And the policy response is catching up. In California, the proposed Parents & Kids Safe AI Act is framed around age-related protections for conversational AI used by minors, including independent safety requirements and enforcement mechanisms — exactly the direction boards should assume regulators will move when AI touches kids.
This piece is about governance.
For private boards, the message is simple: if your AI can reach a child, you should assume regulators will be able to reach you.
Why AI Governance Breaks in Otherwise Competent Companies
Most AI failures don’t happen because leadership is careless. They happen because AI doesn’t live in one place.
It shows up everywhere: HR screens candidates, marketing optimizes spend, customer service routes complaints, sales forecasts churn, product teams add “smart” features. Each department believes it’s managing “its tool.”
Someone approved a budget. Someone signed a contract. Someone attended training.
But nobody is tracking the decisions.
AI is not just software procurement. It’s decision infrastructure.
FoloToy wasn’t deploying a cute feature. They were deploying a system that effectively decided:
What content was appropriate for children
When safety filters should apply
How to respond when conversations drifted into dangerous territory
What tone would build trust with a preschooler
Those are governance decisions. Calling them “features” doesn’t change the liability.
This isn't FoloToy's problem alone. The toy industry has repeatedly struggled with AI governance. In 2015, Mattel's Hello Barbie - an internet-connected doll using cloud-based voice recognition - was found to have vulnerabilities that could allow hackers to access recordings of children's conversations. Two years later, Germany banned My Friend Cayla as an 'illegal espionage device' after researchers discovered its Bluetooth connection could be hijacked by anyone within range to listen to or speak with children. Together, these cases show a recurring pattern: even sophisticated, well‑resourced brands can ship AI‑enabled toys faster than they build the governance to control them.
The Four Control Layers Every Board Needs
Strip away the jargon and AI governance collapses into four questions with real consequences:
IDENTIFY — Do you know what AI you have?
CLASSIFY — Do you know what it’s allowed to do?
MONITOR — Do you know when it’s not working?
GOVERN — Do you know who can stop it?
These are not “steps.” They are controls.|
And if any layer is missing, the entire system becomes defenseless.
Let’s walk through them - in board language.
Layer 1: IDENTIFY - Know What AI You Actually Have
Most organizations think they’re tracking AI because they track tools.
But AI rarely shows up on invoices as “AI.” It shows up as:
“enhanced customer engagement”
“intelligent content generation”
“personalized learning”
“workflow automation”
The IDENTIFY problem is not whether you can name the vendor.
It’s whether you can name the decision system.
If an AI system can influence outcomes for vulnerable populations - children, patients, seniors, economically disadvantaged users - you don’t get to call it “just a feature.”
The IDENTIFY question is simple:
What decisions can this system make, for whom, and with what downside?
If that mapping is absent, the next layers won’t happen - because nobody owns the risk. And unowned risk doesn’t get managed.
Layer 2: CLASSIFY - Know What It’s Allowed to Do
Once you identify the decision system, you classify risk.
An AI that drafts email subject lines? Low risk.
An AI that influences credit eligibility? High risk.
An AI providing unsupervised interaction to young children? Critical risk.
Classification is where governance stops being philosophical and becomes operational:
What’s the risk tier?
What’s prohibited content?
What's the system's failure mode—does it lock down or keep running?
What evidence is required to launch?
Who must sign off?
If a system is critical-risk, you can’t “assume safety.”
You must demonstrate safety.
Using Kumma Bear as the example: in a critical-risk application, classification should force the uncomfortable vendor reality check:
If the model provider’s own terms say the tool isn’t intended for users under 13, why are we putting it inside an experience designed for children as young as three?
That’s not a technology debate. It’s a governance decision.
In Kumma Bear, someone should have stopped at classification and asked: “What’s the worst-case failure mode - and do we have written evidence we tested for it?”
If the answer wasn’t documented proof, the product wasn’t ready to launch.
Layer 3: MONITOR — Know When It’s Not Working
This is where AI governance most often collapses in real life - because monitoring is work.
Policies don’t monitor.
Vendor assurances don’t monitor.
A launch checklist doesn’t monitor.
Real monitoring means detecting problems before customers do:
Content drift over longer conversations
Repeated unsafe outputs
Escalation patterns in sensitive topics
Complaint clusters
Near misses that signal systemic issues
It means that you can detect them fast enough to act.
Based on the public record, the Kumma Bear case raises the question every board should fear:
What would have needed to go wrong - and for how long - before someone inside the company would have known?
If the answer is “until someone outside publishes a report,” you don’t have monitoring.
You have external quality control.
Effective monitoring for a critical-risk AI system looks like this:
Human-in-the-loop review for all conversations flagged by automated content filters
Sampling audits across user demographics and conversation lengths
Escalation triggers tied to severity (three flagged interactions in any 100-conversation sample = automatic pause)
Board-level visibility when critical-risk systems breach thresholds
Without monitoring, you’re flying blind. And when something goes wrong, you’re stuck explaining why you didn’t see it coming.
Layer 4: GOVERN — Know Who Can Stop It
This is the layer most companies skip entirely.
They have approval processes.
They have principles.
They have committees.
But they don’t have stop authority.
So here’s the real question:
Who can prevent deployment - and who can force shutdown - without asking permission?
Not “who can raise a concern.”
Not “who can recommend review.”
Who can stop it.
If that role doesn’t exist, the organization is structurally incapable of controlling AI risk under pressure - because when something goes wrong, decisions get slower, not faster.
The Vendor Governance Trap
Kumma Bear exposes a reality boards routinely underestimate:
You are not just governing your AI use.
You’re operating inside your vendor’s governance.
Here’s the question the Kumma Bear board should have forced into the room early:
If the model provider’s terms and policies restrict use with children—and prohibit sexual content involving minors - how did an OpenAI-powered model end up inside a toy marketed to kids as young as three?
I’m not saying this to litigate blame. I’m saying it because it clarifies the dependency: policy is not protection unless someone is monitoring, enforcing, and validating before a product ships.
When your product depends on a third-party model, the board should demand clarity on five things:
What guardrails exist - and can they be overridden?
What monitoring does the provider perform (and what do we perform)?
What triggers suspension or revocation?
What happens when the model changes under us?
What evidence is required for reinstatement after a violation?
In Kumma Bear, OpenAI ultimately cut off access. That’s vendor governance working.
But it happened after the product was in market, after children were exposed, and after public outcry.
A “reputable vendor” is not a control.
It’s a dependency.
And dependencies require governance.
By 2025, large public companies were explicitly disclosing AI risks and board-level oversight in their filings, and regulators were moving against kids’ tech - even when third-party providers were in the loop.
For private boards, that’s the warning label: you can outsource the model, but you can’t outsource accountability for what it does in your product.
What Good Looks Like: The Control Layer in Practice
Imagine a control layer had existed before Kumma Bear launched:
IDENTIFY: This is an unsupervised conversational AI that can influence what children ages 3–12 hear, repeat, and act on. If it fails, we’re not fixing bugs - we’re explaining to parents why our toy discussed inappropriate content with their kindergartener. Downside is existential: child safety, regulatory scrutiny, litigation risk, and brand damage.
CLASSIFY: Critical risk. Launch requires documented evidence: adversarial testing across age groups and conversation lengths, child-development expert review, defined prohibited-content thresholds, and executive sign-off acknowledging that failure means reputational catastrophe.
MONITOR: Real-time content flagging with mandatory human-in-the-loop review. Automatic suspension upon any instance of sexualized content involving minors, violence, or guidance on accessing hazards. Weekly sampling audits. Complaint escalation protocol with a 24-hour response requirement. Board notification within 48 hours of any critical threshold breach.
GOVERN: Named executive owner (e.g., CRO) has stop authority at any stage. No launch without that approval. Immediate shutdown authority without executive committee approval. Vendor contract includes explicit suspension triggers and reinstatement evidence requirements.
With that structure, one of two things happens:
The product doesn’t launch as originally designed, or
Leadership launches with explicit risk acceptance - documented, justified, monitored, and owned.
That’s the difference between governance and hope.
Let’s Get Elemental
Before your next board meeting, run a control-layer test - not a policy review.
1) Can you produce evidence?
If a regulator or plaintiff asked you to prove you tested AI systems under real use conditions before deployment, what would you show them?
Not “we have a responsible AI policy.”
Not “we trust our vendor.”
Evidence: testing protocols, results, sign-offs, risk thresholds, and monitoring logs.
2) Can you name names?
For your three highest-risk AI systems:
What triggers immediate shutdown?
Who has authority to do it - today - without a committee meeting?
If you can’t answer with names and thresholds, you have a gap in your governance.
3) Can you prove you acted responsibly?
If your AI creates harm, can you prove controls existed before it happened - not after?
Governance is proof: controls, evidence, monitoring, and decision authority.
Navigator Tip of the Week
Ask management this in your next Risk or Audit Committee meeting:
“Show me our highest-risk AI system - and the exact trigger that forces shutdown, including the name of the person empowered to do it.”
A Final Word on Readiness
The four control layers sound simple, but they are not.
Building them requires discipline most organizations discover they lack: the ability to test rigorously, document honestly, escalate quickly, and stop launches when evidence doesn’t support safety claims.
The Elemental AI Governance Navigator doesn’t implement these controls for you. It diagnoses whether your organization can actually sustain them - before you’re forced to prove it in a crisis.
It’s a rigorous assessment process that pressure-tests readiness across the domains that make a control layer real: governance structure, risk classification, decision intelligence, leadership capability, monitoring infrastructure, and change readiness.
If you’re preparing for board conversations about AI governance, the Navigator helps you answer the question leadership should be asking:
Are we ready to govern AI at the level our strategy requires - or will we discover our gaps the way FoloToy did?
Because the control layer is the destination.
The Navigator tells you whether the road beneath you can hold the weight.
About Fayeron Morrison
Fayeron Morrison is the President of Elemental AI, a strategic advisory firm that helps boards and executives navigate the governance challenges of artificial intelligence. She is the creator of the Elemental AI Governance Navigator, a diagnostic tool built to bring clarity and accountability to AI oversight at the highest levels.
A graduate of the Stanford Graduate School of Business Executive Program in AI Leadership, Fayeron is also the author of Elemental AI, a weekly Substack publication focused on AI governance, risk, and boardroom readiness.
Beyond her AI work, Fayeron is a Certified Public Accountant (CPA) and Certified Fraud Examiner (CFE) with a long-standing career advising both public and private companies.
She lives in Newport Beach, California with her husband and their Bernese Mountain Dog, Oakley. She’s the proud mom of three grown sons and, when she’s not writing or advising, she’s likely on a hiking trail with Oakley - where she does some of her best thinking!
📧 fayeron.elementalai@gmail.com
🌐 elementalai.ai


