Compliance

Shadow AI Apps: How to Vet Employee-Built AI Tools

Shadow AI Apps: How to Vet Employee-Built AI Tools

Know what attackers see before they do. See a live Radar report →

The Request You Didn't Train For

An employee or client pings you: "Hey, I built a little app with Claude to help with [scheduling / invoicing / customer lookups]. Can you install it on my work laptop?" There's no vendor page, no changelog, no security whitepaper. Just a zip file, a GitHub link, or a single executable someone generated through a chat prompt. This exact scenario showed up on Reddit when an IT admin got a request to install a client's homemade Claude app and had no idea where to start evaluating it.

This is shadow AI in its most literal form: not an employee pasting data into ChatGPT, but an actual piece of software — built by someone with no development background, using an AI model as the compiler — now asking for a place on your managed fleet. It looks harmless. It usually isn't vetted at all. And increasingly, the industry data says you should treat it as hostile until proven otherwise.

Why This Is a Bigger Deal Than It Looks

AI-assisted coding tools have made it trivially easy for non-developers to produce working software, but "working" and "safe" are not the same thing. A few recent developments make the stakes concrete:

According to The Hacker News, Apple is tightening macOS Full Disk Access controls specifically because AI agents and apps have been quietly reading files, mail, messages, and browser history without users fully realizing it. If Apple — with its tightly controlled app ecosystem — is responding to this problem, it should tell you everything about the risk level of an unreviewed, homegrown Claude app that has no sandboxing at all.

Integration layers compound the risk. According to The Hacker News, GitLab recently had to patch a critical 9.9-severity flaw in its AI Gateway that allowed command execution on self-hosted servers. That's a vulnerability from a mature, well-resourced vendor. A client-built Claude app connecting to an AI backend carries the same class of exposure, just with none of the patching discipline or security review GitLab has in place.

The people using these tools are also targets. According to The Hacker News, the China-aligned group TA419 has been running adversary-in-the-middle phishing campaigns specifically against people working in AI policy and tooling, harvesting credentials through Microsoft authentication flows. An unvetted AI app that stores API keys or session tokens locally hands an attacker exactly the kind of foothold these campaigns are designed to exploit.

And finally, there's a mindset problem worth naming. A recent piece titled "Is It Fair to Blame 'Rogue' AI for Security Failures?" makes the point that AI agents should be evaluated as untrusted, nondeterministic software — not as productivity shortcuts that get a pass because they're convenient. That's the frame every sysadmin needs before saying yes to a homegrown AI tool.

The Vetting Checklist

Before any employee- or client-built AI app touches a managed device, run it through the same process you'd apply to any unknown third-party binary.

  1. Ask what it actually touches. Get the builder to list every file, folder, API, and credential the app reads or writes. If they can't answer precisely, that's your first red flag.
  2. Check for Full Disk Access or elevated permission requests. Given Apple's own response to AI agents overreaching on file access, treat any permission prompt beyond the bare minimum as suspicious.
  3. Inspect network calls. Does the app phone home to a third-party AI API? Is that traffic encrypted? Is there a hardcoded API key sitting in plain text in the source?
  4. Test it in isolation first. Run the app in a sandboxed VM or an isolated test account, never on a production device with live customer data.
  5. Scan the binary or codebase. If source is available, review it or run it through static analysis. If it's a compiled executable with no source, that alone should disqualify it from production use.
  6. Apply least privilege. If it must run, give it its own service account with the narrowest possible access, not the employee's or client's existing credentials.
  7. Log and monitor. Treat the app like any new endpoint software — log its network activity and file access for the first several weeks.

This is the same discipline we outlined in our deeper breakdown of client-built AI apps and the security risks before you say yes, and it pairs well with the broader governance approach covered in our guide to AI governance and shadow AI data leakage.

Where This Fits in Your Broader Endpoint Strategy

Shadow AI apps aren't a one-off problem — they're a symptom of a larger shadow IT trend where employees and clients bypass security review because the official process feels slower than just building the thing themselves. If you haven't already, pair this checklist with a documented policy on application whitelisting versus endpoint privilege management so unapproved software, AI-built or otherwise, can't silently install itself in the first place. It's also worth revisiting your detection tooling against the pattern described in our post on detecting unauthorized software and rogue agents on your network — because the next homegrown app might not come with a polite request first.

Take Action

Shadow AI apps are easy to miss until one of them leaks customer data or opens a path for a credential-harvesting attack. Proactive scanning catches these gaps before attackers — or a well-meaning client — do it for you. Oscar Six Security's Radar gives you an affordable way to check your environment for exactly these kinds of exposures, starting at $99 a scan. Run a scan today and see what's quietly sitting on your managed devices.

Focus Forward. We've Got Your Six.

Frequently Asked Questions

What is shadow AI and why is it a security risk?

Shadow AI refers to AI tools or apps that employees or clients build and use without IT review or approval, often created through chat-based coding assistants like Claude. It's risky because these apps can access files, credentials, and networks with no security testing, sandboxing, or vendor accountability behind them.

Should I let employees install AI apps they built themselves?

Not without vetting. Treat any employee- or client-built AI app as untrusted third-party software: review its permissions, network calls, and data access before it touches any managed device, and run it in an isolated test environment first.

How do I check what data an AI app can access on a device?

On macOS, review Full Disk Access and permission prompts, which Apple is now tightening specifically because AI agents have been quietly reading files, mail, and browser history. On Windows, monitor network calls and file system activity through your EDR tooling before approving broader rollout.

What tool should I use to find unauthorized AI apps on my network?

A combination of endpoint detection tooling and periodic vulnerability scanning works best. Oscar Six Security's Radar scan, priced at $99, helps small businesses and MSPs identify unauthorized software and exposed attack surface before it becomes a breach.

Can a homegrown AI app lead to a data breach?

Yes. Unvetted AI apps can store API keys in plain text, request excessive file access, or connect to AI backends with unpatched vulnerabilities, any of which can lead to data leakage or remote code execution. Recent incidents like GitLab's critical AI Gateway flaw show this risk applies even to professionally built AI integrations, let alone homemade ones.

Step-by-Step Guide

  1. Ask for a full data access list

    Have the app's builder document every file, API, and credential the app reads or writes before any installation discussion continues.

  2. Check permission requests

    Review whether the app requests Full Disk Access or other elevated permissions, and treat anything beyond the bare minimum as a red flag.

  3. Inspect network traffic

    Identify what external services the app calls, confirm the traffic is encrypted, and check for hardcoded API keys in the source or binary.

  4. Test in an isolated environment

    Run the app in a sandboxed VM or test account first, never on a production device with live business or customer data.

  5. Scan the code or binary

    Run available source through static analysis, or disqualify compiled executables with no source from production use.

  6. Apply least privilege

    If approved, run the app under a dedicated service account with the narrowest possible permissions rather than existing employee credentials.

  7. Monitor after deployment

    Log the app's network and file activity for several weeks post-install to catch unexpected behavior early.

Find out what's exposed. Radar scans your external attack surface and shows you exactly what needs fixing. See a sample report →