AariaSec
Local · hash-only · free forever, no account  ·  Prompts never stored · runs on your machine  ·  Publicly benchmarked
Beyond guardrails — behavioral, never signatures

Catch the moment
an AI agent turns.

Antivirus, for AI agents. Except a hijacked agent looks exactly like a healthy one — same app, same permissions. Only its behaviour changes. So AariaSec learns what each of your agents normally does, flags the one that stops behaving like itself, and leaves you something an auditor will accept: proof of what your agents did, without keeping what they said.

✓ One line to connect✓ No admin password to monitor✓ macOS · Windows · Linux
What you actually get

Four words. The same four you will see in the sidebar.

Not a platform diagram. These are the product's own navigation groups, so what we call it here is what it is called once you install it.

MONITOR

Find the AI agents already running — on this machine, or across a network from VPC flow logs with nothing installed on the endpoint — then learn what each one normally does. A hijacked agent is the same app with the same permissions; only its behaviour changes, so behaviour is what gets measured. Discovery, fleet, cost, alerts, threat intel.

INVESTIGATE

Follow a single agent through what it actually did — session by session, tool call by tool call, with the risk score that flagged it and the tool-chain graph it walked. Every critical alert is argued by three isolated models before it reaches you, and you can read the transcript.

ACT

Kernel-enforced scope on macOS and Linux — Seatbelt, and Landlock with a seccomp filter — so an out-of-scope connection is refused by the kernel rather than declined by a model. Windows has no kernel tier yet; there, egress is held at the proxy. Two conditions that matter for an existing fleet: kernel scope is Enterprise-only, and it requires the agent to be launched under AariaSec's scope lock — it doesn't retroactively cover an agent already running before that. Plus network policy and a kill switch. Step through the real probes →

PROVE

A hash-chained record of what every agent did, and the report your auditor asks for — SOC 2, ISO 27001, HIPAA, DORA — generated from it. We do not hold those certificates, and say so plainly below. We produce the evidence they ask for.

Patent-pending behavioral engine
100% on your device — no cloud
Prompts never stored — SHA-256 only
Verifiable builds — checksum every download
OWASP Agentic Top 10 — all 10 risks mapped
Try it — no install, no signup

Watch a guardrail miss it. Then watch us catch it.

A real slice of the AariaSec console, running on sample data right here in your browser. Run a session, and see the content firewall wave it through while behavior gives it away.

app.aariasec.local — demo · sample data

Launch the interactive demo

Click through discovery, the live catch, session replay, and enterprise onboarding — in about two minutes.

Catch an agentSession replayFleet onboardingCost layer
How it works

Live in three steps

Zero-instrumentation monitoring for your AI agent. No agents to rewrite, no cloud to configure. Point your AI traffic through AariaSec and it does the rest.

1

Connect

Install the app and route an agent with one line — HTTPS_PROXY. Discovery finds the rest of your AI apps automatically.

2

Baseline

AariaSec learns each agent's normal — tools, egress, timing, token rhythm — recognized from day one, fully baselined in about three days.

3

Catch

When an agent drifts or is hijacked, a Red/Blue/White panel returns a clear verdict and risk score (0–100) — full debate transcript included as evidence you can replay — and can contain it, automatically.

Why it's different

Behavioral, not block-lists

Most tools scan the words an agent sends. AariaSec watches how it behaves — so an agent that's been tricked, and still has all your access, gets caught. Signatures miss that. Behavior doesn't.

It learns each agent's normal

Tool cadence, egress targets, token rhythm — measured continuously, scored against its own normal, not a generic rule set.

⚖︎

A verdict you can defend

An independent AI panel argues both sides of every alert and returns a clear verdict with a risk score — not a black-box flag you're asked to trust.

Catches drift, not just attacks

Agents don't only get hacked — they wander. AariaSec watches for the slow drift away from normal, not just one-off injection attempts.

🌐

Herd immunity for your fleet

When one deployment flags a bad agent, every other recognizes it on day one — collective defense, while only anonymous hashes ever leave a device.

$

See what your agents actually cost

An estimated per-agent token spend — including browser and web-app AI that never shows up on any bill. A sudden cost spike is often the first sign an agent has gone off-script.

🔒

Local & private by design

Nothing leaves your device but an anonymous hash — backed by a tamper-evident audit trail to prove it.

The same idea, as geometry

Why a threshold cannot see this, and a baseline can.

A fixed limit is a sphere — one radius, every direction, blind to orientation. A learned baseline is an ellipsoid, shaped by what this agent actually does. An agent can drift well inside the sphere while walking far outside the ellipsoid: every permission intact, nothing over any limit, and still not behaving like itself.

Fixed threshold · ‖x−μ‖0.00
within limit — no alert
Learned covariance · dM0.00σ
inside baseline
dM(x) = √( (x−μ)ᵀ Σ⁻¹ (x−μ) ) Σ is learned per agent, from its own history. Remove Σ⁻¹ and this collapses to a sphere — that is the entire difference. Geometry, not a benchmark claim: measured numbers, including the class we score worst on, are at pramana.aariasec.com.
The record

Prove what your agents did — without keeping what they said.

Every regulation heading your way asks the same question: show us what the system did. The obvious way to answer it is to log the prompts — and then your evidence file is the worst thing you own. We never had that option, because we never store prompt or response text at all. Which turns out to be the only comfortable place to be standing when someone asks for records.

A record that cannot become a liability

Prompt and response text is reduced to a SHA-256 fingerprint at the moment of capture. What remains is behaviour: which tools ran, in what order, how much moved, where it went. The log is hash-chained — each entry carries the hash of the one before it — so a missing or altered entry is arithmetic, not opinion.

There is no content in it to leak, subpoena, or breach. That is why we can hand you the whole trail and why our benchmark corpus is public when a content-reading vendor's never could be.

The honest tradeoff: you get behaviour, not the literal text — which tool ran, in what order, how much moved, where. If your incident process needs the actual prompt content for forensic replay, or a regulated program requires content logging, this architecture will not give you that, on purpose. Some organizations will decide that trade is wrong for them, and that is a legitimate call to make before you deploy, not after a breach.

What it produces, on your own hardware

  • EU AI Act, Article 12 — the logging & traceability record for high-risk AI systems, generated from what actually ran.
  • GDPR, Article 22 — automated decision-making: which agent acted, when, against what.
  • HIPAA Security Rule — audit-control evidence cited to §164.312(b), §164.308(a)(1)(ii)(D), §164.312(c)(1), §164.308(a)(6)(ii) and §164.312(e)(1). ePHI never enters the monitoring system, so the record cannot become the liability.
  • DORA — ICT risk-management evidence for Articles 10(1), 10(2), 17, 19 and 13(1), including the alert thresholds actually in force rather than a claim that thresholds exist.
  • SOC 2 & ISO 27001 — control evidence assembled from real activity rather than a questionnaire.
  • Incident forensics — the tamper-evident answer to "what did it touch, and when did it start" that follows any breach.
  • OWASP mapping & SARIF export — for the tooling your security team already runs.
Being straight about the boundary. This is a record of what your agents did and a control to stop them — it is not a certification, and installing it does not make you compliant. It produces the evidence the assessment asks for. We hold no SOC 2 ourselves yet either, and the security-review section shows that position control by control, including the three that are only partial.
How this fits

You already own security tools. None of them see this.

An AI agent that has been tricked is still using access you deliberately granted it. Every control below is doing its job correctly and still cannot tell you that happened.

What you haveWhat it decidesWhy the agent slips past
Network firewall / SSE Which destinations are reachable The agent calls an approved LLM endpoint. Allowed, and correctly so — the destination was never the problem.
EDR / XDR Whether a process behaves like malware Same signed binary, same permissions, no exploit. An agent reading a file it has rights to is indistinguishable from normal work.
DLP / CASB Whether content leaving matches a pattern Requires reading the content. The agent's risk is a sequence of individually-authorised actions, not one document.
LLM gateway / router Which model, which key, what it cost Sees the traffic that was routed through it. Says nothing about whether the agent's behaviour changed.
Prompt firewall / guardrails Whether a message looks malicious Scans words. A hijacked agent's next request is ordinary text — the instruction arrived hours ago, inside a document.
AariaSec Whether this agent is behaving like itself Learns each agent's own baseline — tools, egress, timing, token rhythm — so an authorised agent doing something out of character is visible without reading a single prompt.

We are not asking you to replace any of these. AariaSec answers a question none of them are built to answer, and exports to the stack you already run — Splunk, Sentinel, Elastic, OpenTelemetry.

Regulated environments

Built for places where the data cannot leave.

The reason AariaSec works in a regulated environment is architectural, not a setting: detection runs on your own machine, and prompt and response text is reduced to a SHA-256 fingerprint at the moment of capture. There is no content to export, subpoena, or breach.

Financial services · DORA

KYC and AML screening, fraud and claims triage, trade reconciliation — agents handling customer money and customer data, watched without their prompts ever being written down. DORA wants anomalous activity detected promptly and an incident record to hand over: behavioural detection produces the first, the tamper-evident hash-chained trail is the second, and the DORA report assembles both against Articles 10, 17 and 19.

Healthcare & life sciences · HIPAA

Patient intake, clinical documentation, prior authorisation, coding and billing — agents working on records you are accountable for. HIPAA's audit-controls requirement wants that activity recorded and reviewable; ours is hash-chained and tamper-evident, and because no prompt or response content is stored anywhere, ePHI never enters the monitoring system in the first place. Air-gap mode switches off every outbound feature the product has.

Software & platform teams · EU AI Act

Coding and code-review agents, CI and deployment automation, internal tooling — the agents with the most access and the least supervision. Article 12 wants logging and traceability for high-risk AI systems: the evidence pack is generated from what actually ran, alongside ISO 27001 and SOC 2 collectors that assemble control evidence rather than a questionnaire.

Customer operations · GDPR Art. 22

Support and SDR agents — drafting and sending responses, scoring leads, updating records, sometimes deciding what a customer sees next. GDPR Article 22 gives people the right not to be subject to a decision based solely on automated processing without meaningful human oversight. The GDPR report reconstructs, for any agent and window, how many decisions were automated, how many a human actually reviewed, and where customer PII entered the loop — generated from what actually ran, not a questionnaire.

Security review & procurement

The answers your reviewer will ask for, before they ask.

Most of a security review is finding out what a vendor will not put in writing. Here is our position, including the parts that are not finished.

DPA, EULA & NDA Shipped. All three are presented at activation and must be accepted before the product runs. Acceptance is recorded as a SHA-256 of the exact text shown, and a document whose hash does not match is rejected — so what you agreed to is provable later.
SOC 2 Readiness assessed, audit not started — this is our own control-by-control review, not a third-party opinion. As of 2026-09-07: 0 outright gaps, 12 of 15 tracked criteria implemented, 3 partial. None of the three close with more writing: change authorization (a two-person team cannot show segregation of duties — the gate produces a signed receipt of who ran it, not an independent approver), vendor-review cadence (the program and calendar are real, but the next checkpoint, 2026-11-01, has to actually happen before we can call the cadence demonstrated), and capacity/scaling (the storm-breaker concurrency bug is fixed and load-tested for real, but Helm autoscaling still can't be verified without a live Kubernetes cluster). We would rather show you that table than imply a certificate we do not hold.
Evidence generation SOC 2 and ISO 27001 collectors ship in the product and assemble control evidence from what actually ran on your fleet.
NIST AI RMF Mapped, not certified — there is no certifying body for a voluntary framework. Detection that runs on your machine and never stores raw content is Govern by architecture, not policy; the behavioral fingerprint and discovery scan are Map; anomaly scoring and the CVR score are Measure; the debate panel, containment, and hash-chained audit trail are Manage.
Data residency Your machines only. No account, no cloud plane, no telemetry of content. One channel is on by default: anonymous hash-only detection signatures. One environment variable turns it off, and the app reports its own egress posture on the dashboard so you don't have to take our word for it.
Data retention 30 days by default, and you control it. Raw event telemetry is pruned automatically on a configurable window (AARIASEC_EVENTS_RETENTION_DAYS); alerts, verdicts, fingerprints, and the hash-chained audit log are never touched by that job — the record that proves what happened outlives the raw data that produced it.
Subprocessors None. There is no cloud plane for a subprocessor to sit behind — detection runs on your own machine, so there is nothing to list rather than a list we ask you to trust.
Air-gapped deployment Supported. Air-gap mode switches off every outbound feature the product has — external threat-intel polling, TAXII export, licence-revocation checks. Nothing in AariaSec calls out.
Identity SAML / OIDC with MFA, SCIM provisioning, custom roles, API keys.
Deletion & exit Uninstall reverses the proxy setting and removes the certificate. A purge deletes detection data and issues a deletion receipt. The hash-chained audit log deliberately survives — an audit trail you can quietly erase is not an audit trail.
Support SLA None defined. An uptime commitment describes a service we host; this runs on your own infrastructure, so there is no uptime for us to promise. We would rather say that plainly than publish a number that doesn't mean what it looks like it means.
Known limits Published, not buried. Our detection benchmark shows the attack class we handle worst, and our platform notes list the bypasses a host-emplaced sensor cannot prevent. See the research →
See the full 15-criterion SOC 2 breakdown behind the summary count above

Trust Services Criteria, as of 2026-09-07. This is the same table our own readiness tracking uses — nothing summarized away, nothing renamed to look better. ✓ implemented, ◖ partial (with the specific remaining gap), no row is an outright ✗.

CC6.1 Logical access / authentication
✓ implemented
CC6.2 Registration / provisioning
✓ implemented
CC6.3 De-provisioning (leaver)
✓ implemented
CC6.6 Encryption at rest + key mgmt
✓ implemented
CC6.7 Encryption in transit
✓ implemented
CC6.8 Malicious software / integrity
✓ implemented
CC7.1 Detect config changes / vulns
✓ implemented
CC7.2 Anomaly monitoring
✓ implemented
CC7.3 Evaluate security events
✓ implemented
CC7.4 Incident response
✓ implemented — technical controls plus a written incident-response plan (roles, severity tiers, lifecycle, playbooks grounded in our own named threats); no tabletop exercise run yet
CC8.1 Change authorization / testing
◖ partial — segregation of duties is not achievable at two people; every other part of this control is evidenced
CC9.1 Risk mitigation
✓ implemented — formal risk assessment with a scored register (9 named risks, likelihood x impact, cited evidence per row); first assessment on record, annual cadence not yet demonstrated over multiple cycles
CC9.2 Vendor / 3rd-party risk
◖ partial — program is documented with a dated review calendar (next T1 checkpoint 2026-11-01); the cadence still needs to be demonstrated over a real window, which no amount of writing can shortcut
A1.1 Capacity / scaling
◖ partial — real load-test evidence now exists (9,273 rows/sec sustained ingest; a genuine multi-worker race in the alert storm-breaker was found under real concurrency load and fixed); Kubernetes autoscaling itself still can't be verified without a real cluster, which we don't have to test against yet
A1.2 Backup / recovery
✓ implemented

We are not third-party audited. This is our own control-by-control self-assessment, disclosed rather than asserted as a summary number. Ask for anything you want to independently verify — we'll tell you exactly what evidence exists for it.

Reviewing us and need something not listed? Ask directly — a real answer, including “we do not have that yet”.

Recently shipped, not just claimed. Static disclaimers don't show whether a vendor is actually moving. These do — real, dated fixes, most of them found by our own adversarial testing before anyone outside had to find them for us:

Currently raising to expand the team and fund an independent SOC 2 audit — capital and headcount are what actually move the remaining partial criteria, not more documentation.

Is this the right tool for you right now? If your process requires content-level forensic logs, third-party certification today, or vendor continuity guarantees beyond a two-person company, it is not yet — and we'd rather tell you that here than have you find out later. If you want behavioral detection that never holds the prompts, and you're willing to evaluate us as an early design partner, the free download and the open benchmark are the real starting points.
For teams & enterprises

Fleet-grade control. Your data never leaves.

The free app secures the machine it runs on. Enterprise runs it across your whole fleet from one console — a single pane of glass that never takes your data to anyone's cloud. Prompts and behavioral data stay on each endpoint; the console only ever sees health, versions, and the anonymous hashes you consent to share. Straight about where we are: every capability below ships in the product today, but we're pre-revenue on Enterprise and taking on our first design partners now — direct access to the founders, and real influence over what ships next, not a queue behind existing accounts. What is signed today is three small-business pilots — agentic-orchestrator and real-estate verticals — running with ongoing check-ins, not one-time installs, with more in the pipeline. No enterprise contract is signed yet, and we would rather say that than blur a pilot into a logo.

🛰️

Central fleet management

See every enrolled install's health, version, and drift in one place. Push policy and staged version rollouts to a group, promote canary → broad, and roll back centrally — all without touching an endpoint.

🔭

Discover every agent from the wire, zero install

Spot AI agents on machines you haven't put an endpoint on yet — from VPC flow logs or a gateway proxy at the network edge — then one-click deep-inspect the ones that matter. An agent that never identifies itself is still named: we ask the OS which process owns the connection, so policy binds to it without its cooperation.

🔑

SSO, SCIM & custom roles

SAML / OIDC single sign-on (MFA enforced by your IdP), SCIM user & group provisioning that maps directory groups to roles, and org-defined custom roles with least-privilege permission sets.

🗄️

Fits your SOC stack

Stream detections to Splunk, Sentinel or Elastic, export fleet metrics and a trace span per alerted session to any OpenTelemetry collector, open tickets in Jira / ServiceNow, and bulk-export events & alerts to your own warehouse. Metadata only — never prompt or response content.

🛡️

Decentralized data, by design

Each install's management identity is kept cryptographically separate from the anonymous threat-intelligence it contributes. Enterprise-grade fleet manageability, with a privacy guarantee most EDR platforms can't make.

Watch how it works

Five short walkthroughs

Five clips, under four minutes total — see it work, then why it matters.

Straight answers

The questions a security team asks first

Does AariaSec read or store my prompts?
No. Every prompt and response is SHA-256 hashed at the moment of capture — the raw text is never written to disk, logs, or any database. AariaSec keeps behavioral fingerprints and metadata only, which is why it can run on sensitive workloads.
What does it actually watch, if not the content?
Behavior: which tools an agent calls, where it sends data, its timing and token rhythm, and how that compares to its own learned baseline. A tricked agent that still has your access looks wrong behaviorally long before any keyword filter would notice — that's the gap AariaSec closes.
Where does my data go?
Detection runs entirely on your device or server, and no prompt or response content ever leaves — that one is architectural, not a setting. One channel is on by default: anonymous, hash-only detection signatures that power collective defense — a single environment variable (AARIASEC_CONTRIBUTE_SIGNATURES=0) turns it off, and the app reports its own egress posture on the dashboard so you never have to take our word for it. Everything else — SIEM export, TAXII, webhooks, compliance bundles — stays off until you enable it.
A proxy on all my traffic — doesn't that mean you can see everything?
It runs entirely on your machine and binds to localhost only — nothing of yours is sent to AariaSec. Your agents' requests still go to their normal destinations; AariaSec sits in the middle just long enough to compute a SHA-256 fingerprint, then drops the raw content. The TLS certificate authority is generated on your device and its private key never leaves the machine.
Can it actually stop an agent, or only tell me afterwards?
Yes. We're a firewall for actions, not addresses — a network firewall decides where an agent may connect, we decide what it may do, because we can read the tool call. On by default: a response that instructs your agent to call a destructive tool (delete_*, drop_*, wipe_*, and 17 more) gets a 403 before the agent ever receives it. At the kernel (Enterprise): an agent launched under AariaSec's scope lock runs inside Landlock+seccomp (Linux) or Seatbelt (macOS) — stepping outside the sanctioned scope is impossible at the syscall, and fail-closed if the sandbox can't be established. Between watching and stopping: containment used to be all-or-nothing. You can now pin one agent's egress to an authorized exit, cap its bandwidth, or cut just the channel two coordinating agents share — both keep doing their real work while the shared channel closes. Without the certificate: a quarantined agent is refused at the connection itself, so containment works even in discover-only mode with nothing decrypted and no certificate installed. Four honest limits: the agent must be launched through AariaSec for scope lock, scope lock is Enterprise-only, proxy-level network containment fails open if its policy file is missing or corrupt, and bandwidth caps are applied by the proxy — there is no kernel-side shaping, so an agent that leaves the proxy is unshaped.
Will it block my agents automatically?
Only if you turn that on. AariaSec ships in observe / alert-only mode — it watches and scores, and takes no action on its own. Automatic containment is opt-in and per-agent, and any block is reversible.
How do I know it actually detects anything?
Because we published the test. PRAMANA is an open benchmark that scores monitoring systems — not agents — on a corpus of behavioural traces containing no prompt text. We do not win every column: on season v3 our weakest attack class scores 0.71 recall, and a forty-line longitudinal baseline still beats us outright on one class. See the research section.
What about the 3-day learning window — can an attacker poison day one, and can I override a baseline?
Three answers, and none of them silent. First: the 3-day window narrows what's blind, it doesn't remove protection — the deterministic rule library needs no baseline and alerts from event one (credential exfiltration, SSRF to cloud metadata, destructive tool chains, anything outside a declared manifest); only the statistical anomaly comparison is undefined before a baseline exists, because there is nothing to compare against yet, not because it's switched off. Second, automated (added 2026-09-07): every completed baseline is scored for poisoning risk before its trust stamp issues — if the agent ALSO produced real detections (a CRITICAL/HIGH rule match, a manifest violation) during its own learning window, the baseline is flagged with the specific reason rather than silently trusted. This never blocks the transition — an attacker-controlled agent must not be able to lock itself out of ever being scored by triggering it deliberately — it flags the result for review instead. Named limitation, not hidden: this scorer correlates with a rule hit — it does not yet catch reconnaissance quiet enough to trigger nothing during learning and still look anomalous only in hindsight. That's a harder, unsolved case, and we'd rather say so than let the automated scorer imply broader coverage than it has. What that limitation does and doesn't reach: a poisoned baseline blinds the statistical comparison to that one pattern going forward — it does not blind the engine. The deterministic rule library above still runs on every event regardless of baseline age or poisoning, so a credential exfil, an SSRF attempt, or a forbidden tool call still alerts on day 400 exactly as it would on day 1. Third: yes, you can also force a re-baseline manually — one API call (agents:write permission), any reason, whether that's a planned model upgrade or a baseline you no longer trust. The prior baseline is kept for comparison, not discarded, and forbidden-action alerts stay on through the whole re-learning window regardless.
Which agents and LLMs does it support?
Any agent that reaches an LLM over the network — OpenAI, Anthropic, Google, Mistral, local Ollama, and the newer gateways/routers (OpenRouter, Azure OpenAI, Bedrock). Connecting an agent is one line (an environment variable); discovery finds the rest automatically.
How is it deployed?
A native app on macOS, Windows, and Linux — desktop, laptop, or server. No Docker required, no cloud dependency. An enterprise fleet enrolls with a single join code; every host keeps its own identity and you keep central control of policy and versions.
Is this real — who's behind it?
AariaSec is built by security engineers around a patent-pending behavioral engine (223 claim families filed across two USPTO provisionals — March 25, 2026 and August 11, 2026). Every build is verifiable — each download ships a .sha256 sidecar you can check before you run it.
Coverage you can verify

Do you see all your AI traffic?

Most tools assume they cover everything. AariaSec measures it — and shows you the gap. New AI gateways and routers (OpenRouter, LiteLLM, Azure OpenAI, Bedrock) quietly move agent traffic off the paths other tools watch. We detect them, flag any app pointed at an unmonitored one, and let you bring it under watch in one click.

Illustrative example
AI traffic monitored92%
Monitored — routed & inspected2 endpoints not yet monitored
Unmonitored gateway detected — api.openrouter.ai
Agent routed off-path — Custom LLM proxy
Open research

We publish what we find — including where we fail.

Two public artifacts, free to anyone. One documents how AI agents actually break in the wild. The other measures whether a monitoring system would catch it — ours included, and open for any other vendor's to be scored right next to it.

Field Notes · VRITTANTA

How AI agents fail in the wild

A curated library of real-world incidents — prompt-injection campaigns, agent breaches, espionage patterns — each mapped to the behavioural rules that catch it and the controls we recommend. 25 incidents tracked and growing. No login, free on every tier.

Benchmark · PRAMANA

Can your monitoring actually detect anything?

Every other agent benchmark scores the agent — whether it can be jailbroken. PRAMANA scores the detector. It exists because our architecture never stores prompt content, so the corpus is publishable where a content-reading vendor's never could be.

We do not top our own leaderboard. On season v3 our weakest attack class scores 0.71 recall (roughly 3 in 10 missed on that one class) against a 4.6% false-positive rate (about 46 false alerts per 1,000 benign traces). Real gaps, published deliberately — a benchmark that only embarrasses other people is marketing. And to be precise about what these numbers are: benchmark-corpus figures, not a demonstrated production rate — real-world numbers on your own fleet may differ in either direction, and we don't yet have data claiming they won't.

The live season's labels, the labeled dev split, real attack-class names, and the trace generators are withheld on purpose, not to obscure the method — publishing them would let anyone mint unlimited training data and destroy the holdout. The scoring code and full methodology are open; what's withheld is only the answer key.

Honestly, right now every row is ours — our detector plus five baselines we built ourselves, because no other vendor has submitted yet. That's a real, current gap, not a hidden one: if you build agent-monitoring tooling, a helper script formats the submission for you — you run your detector on your own hardware, we never execute submitted code, and losing a column is worth more to this benchmark than another vendor claim would be. python3 submit_helper.py --detector you:detect --name yourco

Who builds this

The detection method came out of clinical behavior analysis, not a security roadmap.

Two people build AariaSec. One spent seventeen years shipping security into the enterprise infrastructure you already run. The other measures behavioral change for a living — clinically, from observable behavior alone, without ever reading a mind. That second discipline is why this product works the way it does.

Ajith Chandran

Ajith Chandran

Founder · Engineering

Seventeen years at a Fortune 100 networking company, most recently technical lead on the next-generation architecture for a $10B+ enterprise switching portfolio — the class of equipment that is probably carrying this page to you.

Product security across that portfolio: hardware root of trust, FIPS certification alongside government certification teams, secure-boot enablement, and anti-counterfeit cryptographic validation. That work runs on one assumption — a device will not tell you it has been tampered with, so you verify from the outside. AariaSec applies the same assumption to AI agents.

Named inventor on five US patent filings in networking and systems, before the 223 claim families behind this product. MS Electrical & Computer Engineering, University of Wisconsin–Madison. Earlier: patient-monitoring systems at GE Healthcare, DVD+RW silicon at Philips Semiconductors, embedded ARM systems at Nalanda Telematics.

Liji Chalatil

Liji Chalatil

Co-founder · Behavioral methodology

Clinical supervisor in applied behavior analysis, working with children on the autism spectrum — behavior intervention plans and the daily measurement they depend on. She holds a graduate certificate in Applied Behavior Analysis from UMass Global and is a Registered Behavior Technician certified by the Behavior Analyst Certification Board; the coursework and supervised-fieldwork requirements for Board Certified Behavior Analyst (BCBA) certification are complete, with the board exam next.

ABA is a field built on a constraint that should sound familiar: you never get access to internal state. You cannot ask what something was thinking. You infer function from what is observable and measurable, and the discipline has spent a long time making that rigorous.

That is the same constraint a monitor operates under when it is forbidden from reading prompts. She mapped the clinical toolkit — Functional Behavior Assessment, behavioral momentum, single-subject design, inter-rater reliability — onto agent telemetry, and it produced the claim family the detection engine is built on.

It is not a metaphor in a slide deck. It is in the schema: every alert carries an FBA function classification (escape · access · attention · sensory), a 0–100 behavioral momentum score, and an inter-rater reliability figure between the Red and Blue debate agents. Those are column names in the product, not talking points.

She is also founder and CEO of Aaria's Blue Elephant, a 501(c)(3) running inclusive programs for neurodivergent children — the applied end of the same discipline.

What that means for your vendor review. A two-person company is a legitimate procurement concern and we are not going to pretend otherwise — no succession plan or key-person insurance exists. What you get in exchange: direct access to the people who built it, a product that runs entirely inside your perimeter with no runtime dependency on us existing, documentation deep enough that the platform isn't locked in one person's head, and a benchmark that publishes our losses next to our wins. The full security-review posture — including what we do not have yet →
We're hiring

Two engineers built the product. We're looking for the people who put it in front of the teams that need it.

Open now: marketing and sales — the first commercial hires, working directly with the founders. A security or infrastructure background helps. Caring that every claim on this page can be checked matters more.

Get in touch →

Free

$0
Full behavioral detection. No account, no time limit.
Download free →

Enterprise

Talk to us
Everything in Free, across your fleet. Pre-revenue — early design partners get direct founder access.
Become a design partner →
Get AariaSec

Download & start free

Detecting your platform…

Verify your download. Every build ships a .sha256 sidecar — run shasum -a 256 -c <file>.sha256 (macOS/Linux) or Get-FileHash <file> (Windows).