Antivirus, for AI agents. Except a hijacked agent looks exactly like a healthy one — same app, same permissions. Only its behaviour changes. So AariaSec learns what each of your agents normally does, flags the one that stops behaving like itself, and leaves you something an auditor will accept: proof of what your agents did, without keeping what they said.
Not a platform diagram. These are the product's own navigation groups, so what we call it here is what it is called once you install it.
Find the AI agents already running — on this machine, or across a network from VPC flow logs with nothing installed on the endpoint — then learn what each one normally does. A hijacked agent is the same app with the same permissions; only its behaviour changes, so behaviour is what gets measured. Discovery, fleet, cost, alerts, threat intel.
Follow a single agent through what it actually did — session by session, tool call by tool call, with the risk score that flagged it and the tool-chain graph it walked. Every critical alert is argued by three isolated models before it reaches you, and you can read the transcript.
Kernel-enforced scope on macOS and Linux — Seatbelt, and Landlock with a seccomp filter — so an out-of-scope connection is refused by the kernel rather than declined by a model. Windows has no kernel tier yet; there, egress is held at the proxy. Two conditions that matter for an existing fleet: kernel scope is Enterprise-only, and it requires the agent to be launched under AariaSec's scope lock — it doesn't retroactively cover an agent already running before that. Plus network policy and a kill switch. Step through the real probes →
A hash-chained record of what every agent did, and the report your auditor asks for — SOC 2, ISO 27001, HIPAA, DORA — generated from it. We do not hold those certificates, and say so plainly below. We produce the evidence they ask for.
A real slice of the AariaSec console, running on sample data right here in your browser. Run a session, and see the content firewall wave it through while behavior gives it away.
Click through discovery, the live catch, session replay, and enterprise onboarding — in about two minutes.
Zero-instrumentation monitoring for your AI agent. No agents to rewrite, no cloud to configure. Point your AI traffic through AariaSec and it does the rest.
Install the app and route an agent with one line — HTTPS_PROXY. Discovery finds the rest of your AI apps automatically.
AariaSec learns each agent's normal — tools, egress, timing, token rhythm — recognized from day one, fully baselined in about three days.
When an agent drifts or is hijacked, a Red/Blue/White panel returns a clear verdict and risk score (0–100) — full debate transcript included as evidence you can replay — and can contain it, automatically.
Most tools scan the words an agent sends. AariaSec watches how it behaves — so an agent that's been tricked, and still has all your access, gets caught. Signatures miss that. Behavior doesn't.
Tool cadence, egress targets, token rhythm — measured continuously, scored against its own normal, not a generic rule set.
An independent AI panel argues both sides of every alert and returns a clear verdict with a risk score — not a black-box flag you're asked to trust.
Agents don't only get hacked — they wander. AariaSec watches for the slow drift away from normal, not just one-off injection attempts.
When one deployment flags a bad agent, every other recognizes it on day one — collective defense, while only anonymous hashes ever leave a device.
An estimated per-agent token spend — including browser and web-app AI that never shows up on any bill. A sudden cost spike is often the first sign an agent has gone off-script.
Nothing leaves your device but an anonymous hash — backed by a tamper-evident audit trail to prove it.
A fixed limit is a sphere — one radius, every direction, blind to orientation. A learned baseline is an ellipsoid, shaped by what this agent actually does. An agent can drift well inside the sphere while walking far outside the ellipsoid: every permission intact, nothing over any limit, and still not behaving like itself.
Every regulation heading your way asks the same question: show us what the system did. The obvious way to answer it is to log the prompts — and then your evidence file is the worst thing you own. We never had that option, because we never store prompt or response text at all. Which turns out to be the only comfortable place to be standing when someone asks for records.
Prompt and response text is reduced to a SHA-256 fingerprint at the moment of capture. What remains is behaviour: which tools ran, in what order, how much moved, where it went. The log is hash-chained — each entry carries the hash of the one before it — so a missing or altered entry is arithmetic, not opinion.
There is no content in it to leak, subpoena, or breach. That is why we can hand you the whole trail and why our benchmark corpus is public when a content-reading vendor's never could be.
The honest tradeoff: you get behaviour, not the literal text — which tool ran, in what order, how much moved, where. If your incident process needs the actual prompt content for forensic replay, or a regulated program requires content logging, this architecture will not give you that, on purpose. Some organizations will decide that trade is wrong for them, and that is a legitimate call to make before you deploy, not after a breach.
An AI agent that has been tricked is still using access you deliberately granted it. Every control below is doing its job correctly and still cannot tell you that happened.
| What you have | What it decides | Why the agent slips past |
|---|---|---|
| Network firewall / SSE | Which destinations are reachable | The agent calls an approved LLM endpoint. Allowed, and correctly so — the destination was never the problem. |
| EDR / XDR | Whether a process behaves like malware | Same signed binary, same permissions, no exploit. An agent reading a file it has rights to is indistinguishable from normal work. |
| DLP / CASB | Whether content leaving matches a pattern | Requires reading the content. The agent's risk is a sequence of individually-authorised actions, not one document. |
| LLM gateway / router | Which model, which key, what it cost | Sees the traffic that was routed through it. Says nothing about whether the agent's behaviour changed. |
| Prompt firewall / guardrails | Whether a message looks malicious | Scans words. A hijacked agent's next request is ordinary text — the instruction arrived hours ago, inside a document. |
| AariaSec | Whether this agent is behaving like itself | Learns each agent's own baseline — tools, egress, timing, token rhythm — so an authorised agent doing something out of character is visible without reading a single prompt. |
We are not asking you to replace any of these. AariaSec answers a question none of them are built to answer, and exports to the stack you already run — Splunk, Sentinel, Elastic, OpenTelemetry.
The reason AariaSec works in a regulated environment is architectural, not a setting: detection runs on your own machine, and prompt and response text is reduced to a SHA-256 fingerprint at the moment of capture. There is no content to export, subpoena, or breach.
KYC and AML screening, fraud and claims triage, trade reconciliation — agents handling customer money and customer data, watched without their prompts ever being written down. DORA wants anomalous activity detected promptly and an incident record to hand over: behavioural detection produces the first, the tamper-evident hash-chained trail is the second, and the DORA report assembles both against Articles 10, 17 and 19.
Patient intake, clinical documentation, prior authorisation, coding and billing — agents working on records you are accountable for. HIPAA's audit-controls requirement wants that activity recorded and reviewable; ours is hash-chained and tamper-evident, and because no prompt or response content is stored anywhere, ePHI never enters the monitoring system in the first place. Air-gap mode switches off every outbound feature the product has.
Coding and code-review agents, CI and deployment automation, internal tooling — the agents with the most access and the least supervision. Article 12 wants logging and traceability for high-risk AI systems: the evidence pack is generated from what actually ran, alongside ISO 27001 and SOC 2 collectors that assemble control evidence rather than a questionnaire.
Support and SDR agents — drafting and sending responses, scoring leads, updating records, sometimes deciding what a customer sees next. GDPR Article 22 gives people the right not to be subject to a decision based solely on automated processing without meaningful human oversight. The GDPR report reconstructs, for any agent and window, how many decisions were automated, how many a human actually reviewed, and where customer PII entered the loop — generated from what actually ran, not a questionnaire.
Most of a security review is finding out what a vendor will not put in writing. Here is our position, including the parts that are not finished.
| DPA, EULA & NDA | Shipped. All three are presented at activation and must be accepted before the product runs. Acceptance is recorded as a SHA-256 of the exact text shown, and a document whose hash does not match is rejected — so what you agreed to is provable later. |
| SOC 2 | Readiness assessed, audit not started — this is our own control-by-control review, not a third-party opinion. As of 2026-09-07: 0 outright gaps, 12 of 15 tracked criteria implemented, 3 partial. None of the three close with more writing: change authorization (a two-person team cannot show segregation of duties — the gate produces a signed receipt of who ran it, not an independent approver), vendor-review cadence (the program and calendar are real, but the next checkpoint, 2026-11-01, has to actually happen before we can call the cadence demonstrated), and capacity/scaling (the storm-breaker concurrency bug is fixed and load-tested for real, but Helm autoscaling still can't be verified without a live Kubernetes cluster). We would rather show you that table than imply a certificate we do not hold. |
| Evidence generation | SOC 2 and ISO 27001 collectors ship in the product and assemble control evidence from what actually ran on your fleet. |
| NIST AI RMF | Mapped, not certified — there is no certifying body for a voluntary framework. Detection that runs on your machine and never stores raw content is Govern by architecture, not policy; the behavioral fingerprint and discovery scan are Map; anomaly scoring and the CVR score are Measure; the debate panel, containment, and hash-chained audit trail are Manage. |
| Data residency | Your machines only. No account, no cloud plane, no telemetry of content. One channel is on by default: anonymous hash-only detection signatures. One environment variable turns it off, and the app reports its own egress posture on the dashboard so you don't have to take our word for it. |
| Data retention | 30 days by default, and you control it. Raw event telemetry is pruned
automatically on a configurable window (AARIASEC_EVENTS_RETENTION_DAYS); alerts,
verdicts, fingerprints, and the hash-chained audit log are never touched by that job — the record
that proves what happened outlives the raw data that produced it. |
| Subprocessors | None. There is no cloud plane for a subprocessor to sit behind — detection runs on your own machine, so there is nothing to list rather than a list we ask you to trust. |
| Air-gapped deployment | Supported. Air-gap mode switches off every outbound feature the product has — external threat-intel polling, TAXII export, licence-revocation checks. Nothing in AariaSec calls out. |
| Identity | SAML / OIDC with MFA, SCIM provisioning, custom roles, API keys. |
| Deletion & exit | Uninstall reverses the proxy setting and removes the certificate. A purge deletes detection data and issues a deletion receipt. The hash-chained audit log deliberately survives — an audit trail you can quietly erase is not an audit trail. |
| Support SLA | None defined. An uptime commitment describes a service we host; this runs on your own infrastructure, so there is no uptime for us to promise. We would rather say that plainly than publish a number that doesn't mean what it looks like it means. |
| Known limits | Published, not buried. Our detection benchmark shows the attack class we handle worst, and our platform notes list the bypasses a host-emplaced sensor cannot prevent. See the research → |
Trust Services Criteria, as of 2026-09-07. This is the same table our own readiness tracking uses — nothing summarized away, nothing renamed to look better. ✓ implemented, ◖ partial (with the specific remaining gap), no row is an outright ✗.
We are not third-party audited. This is our own control-by-control self-assessment, disclosed rather than asserted as a summary number. Ask for anything you want to independently verify — we'll tell you exactly what evidence exists for it.
Reviewing us and need something not listed? Ask directly — a real answer, including “we do not have that yet”.
Currently raising to expand the team and fund an independent SOC 2 audit — capital and headcount are what actually move the remaining partial criteria, not more documentation.
The free app secures the machine it runs on. Enterprise runs it across your whole fleet from one console — a single pane of glass that never takes your data to anyone's cloud. Prompts and behavioral data stay on each endpoint; the console only ever sees health, versions, and the anonymous hashes you consent to share. Straight about where we are: every capability below ships in the product today, but we're pre-revenue on Enterprise and taking on our first design partners now — direct access to the founders, and real influence over what ships next, not a queue behind existing accounts. What is signed today is three small-business pilots — agentic-orchestrator and real-estate verticals — running with ongoing check-ins, not one-time installs, with more in the pipeline. No enterprise contract is signed yet, and we would rather say that than blur a pilot into a logo.
See every enrolled install's health, version, and drift in one place. Push policy and staged version rollouts to a group, promote canary → broad, and roll back centrally — all without touching an endpoint.
Spot AI agents on machines you haven't put an endpoint on yet — from VPC flow logs or a gateway proxy at the network edge — then one-click deep-inspect the ones that matter. An agent that never identifies itself is still named: we ask the OS which process owns the connection, so policy binds to it without its cooperation.
SAML / OIDC single sign-on (MFA enforced by your IdP), SCIM user & group provisioning that maps directory groups to roles, and org-defined custom roles with least-privilege permission sets.
Stream detections to Splunk, Sentinel or Elastic, export fleet metrics and a trace span per alerted session to any OpenTelemetry collector, open tickets in Jira / ServiceNow, and bulk-export events & alerts to your own warehouse. Metadata only — never prompt or response content.
Each install's management identity is kept cryptographically separate from the anonymous threat-intelligence it contributes. Enterprise-grade fleet manageability, with a privacy guarantee most EDR platforms can't make.
Five clips, under four minutes total — see it work, then why it matters.
AARIASEC_CONTRIBUTE_SIGNATURES=0) turns it off, and the app reports its own egress posture on the dashboard so you never have to take our word for it. Everything else — SIEM export, TAXII, webhooks, compliance bundles — stays off until you enable it.delete_*, drop_*, wipe_*, and 17 more) gets a 403 before the agent ever receives it. At the kernel (Enterprise): an agent launched under AariaSec's scope lock runs inside Landlock+seccomp (Linux) or Seatbelt (macOS) — stepping outside the sanctioned scope is impossible at the syscall, and fail-closed if the sandbox can't be established. Between watching and stopping: containment used to be all-or-nothing. You can now pin one agent's egress to an authorized exit, cap its bandwidth, or cut just the channel two coordinating agents share — both keep doing their real work while the shared channel closes. Without the certificate: a quarantined agent is refused at the connection itself, so containment works even in discover-only mode with nothing decrypted and no certificate installed. Four honest limits: the agent must be launched through AariaSec for scope lock, scope lock is Enterprise-only, proxy-level network containment fails open if its policy file is missing or corrupt, and bandwidth caps are applied by the proxy — there is no kernel-side shaping, so an agent that leaves the proxy is unshaped.agents:write permission), any reason, whether that's a planned model upgrade or a baseline you no longer trust. The prior baseline is kept for comparison, not discarded, and forbidden-action alerts stay on through the whole re-learning window regardless..sha256 sidecar you can check before you run it.Most tools assume they cover everything. AariaSec measures it — and shows you the gap. New AI gateways and routers (OpenRouter, LiteLLM, Azure OpenAI, Bedrock) quietly move agent traffic off the paths other tools watch. We detect them, flag any app pointed at an unmonitored one, and let you bring it under watch in one click.
api.openrouter.aiCustom LLM proxyTwo public artifacts, free to anyone. One documents how AI agents actually break in the wild. The other measures whether a monitoring system would catch it — ours included, and open for any other vendor's to be scored right next to it.
A curated library of real-world incidents — prompt-injection campaigns, agent breaches, espionage patterns — each mapped to the behavioural rules that catch it and the controls we recommend. 25 incidents tracked and growing. No login, free on every tier.
Every other agent benchmark scores the agent — whether it can be jailbroken. PRAMANA scores the detector. It exists because our architecture never stores prompt content, so the corpus is publishable where a content-reading vendor's never could be.
We do not top our own leaderboard. On season v3 our weakest attack class scores 0.71 recall (roughly 3 in 10 missed on that one class) against a 4.6% false-positive rate (about 46 false alerts per 1,000 benign traces). Real gaps, published deliberately — a benchmark that only embarrasses other people is marketing. And to be precise about what these numbers are: benchmark-corpus figures, not a demonstrated production rate — real-world numbers on your own fleet may differ in either direction, and we don't yet have data claiming they won't.
The live season's labels, the labeled dev split, real attack-class names, and the trace generators are withheld on purpose, not to obscure the method — publishing them would let anyone mint unlimited training data and destroy the holdout. The scoring code and full methodology are open; what's withheld is only the answer key.
Honestly, right now every row is ours — our detector plus
five baselines we built ourselves, because no other vendor has submitted yet.
That's a real, current gap, not a hidden one: if you build agent-monitoring
tooling, a helper script formats the submission for you — you run your
detector on your own hardware, we never execute submitted code, and losing a
column is worth more to this benchmark than another vendor claim would be.
python3 submit_helper.py --detector you:detect --name yourco
Two people build AariaSec. One spent seventeen years shipping security into the enterprise infrastructure you already run. The other measures behavioral change for a living — clinically, from observable behavior alone, without ever reading a mind. That second discipline is why this product works the way it does.
Seventeen years at a Fortune 100 networking company, most recently technical lead on the next-generation architecture for a $10B+ enterprise switching portfolio — the class of equipment that is probably carrying this page to you.
Product security across that portfolio: hardware root of trust, FIPS certification alongside government certification teams, secure-boot enablement, and anti-counterfeit cryptographic validation. That work runs on one assumption — a device will not tell you it has been tampered with, so you verify from the outside. AariaSec applies the same assumption to AI agents.
Named inventor on five US patent filings in networking and systems, before the 223 claim families behind this product. MS Electrical & Computer Engineering, University of Wisconsin–Madison. Earlier: patient-monitoring systems at GE Healthcare, DVD+RW silicon at Philips Semiconductors, embedded ARM systems at Nalanda Telematics.
Clinical supervisor in applied behavior analysis, working with children on the autism spectrum — behavior intervention plans and the daily measurement they depend on. She holds a graduate certificate in Applied Behavior Analysis from UMass Global and is a Registered Behavior Technician certified by the Behavior Analyst Certification Board; the coursework and supervised-fieldwork requirements for Board Certified Behavior Analyst (BCBA) certification are complete, with the board exam next.
ABA is a field built on a constraint that should sound familiar: you never get access to internal state. You cannot ask what something was thinking. You infer function from what is observable and measurable, and the discipline has spent a long time making that rigorous.
That is the same constraint a monitor operates under when it is forbidden from reading prompts. She mapped the clinical toolkit — Functional Behavior Assessment, behavioral momentum, single-subject design, inter-rater reliability — onto agent telemetry, and it produced the claim family the detection engine is built on.
It is not a metaphor in a slide deck. It is in the schema: every alert carries an FBA function classification (escape · access · attention · sensory), a 0–100 behavioral momentum score, and an inter-rater reliability figure between the Red and Blue debate agents. Those are column names in the product, not talking points.
She is also founder and CEO of Aaria's Blue Elephant, a 501(c)(3) running inclusive programs for neurodivergent children — the applied end of the same discipline.
Open now: marketing and sales — the first commercial hires, working directly with the founders. A security or infrastructure background helps. Caring that every claim on this page can be checked matters more.
.sha256 sidecar — run shasum -a 256 -c <file>.sha256 (macOS/Linux) or Get-FileHash <file> (Windows).