
AI agent monitoring and insider threat detection share the same architecture: profile a baseline, flag deviations, encode assumptions about trust. We make this concrete across thirteen experiments. Using a Unified Behavioural Feature Schema (UBFS) that maps both employee activity logs and agent execution traces into a shared representation, we apply three anomaly detection models-Isolation Forest, LSTM Autoencoder, and Deep Clustering-across five domains. Cross-domain transfer works: an Isolation Forest trained on 329,000 insider threat user-days retains 97% of detection power on agent traces, and transfer to MCP tool-calling benchmarks exceeds within-domain performance (104.8% retention). But the blind spots transfer too. Model distillation attacks span a detection spectrum from trivially caught (FOCUSED extraction: 0.99 AUC-ROC) to nearchance (HYDRA distributed: 0.54), and task decompositionsplitting a malicious objective into innocent subtasks-costs 5-6% detection power for IF and DC (p=0.031, Wilcoxon signed-rank), while LSTM Autoencoder shows near-invariance (-0.7%, p=0.094). Synthetic OWASP profiling identifies Tool Misuse (ASI02) as a blind spot (∼0.52 AUC-ROC), but real-data validation on 500 ATBench trajectories reveals that this blind spot is an artifact of circular synthetic methodology: real ASI02 achieves 0.81-0.94 AUC-ROC across all models and feature sets. Augmenting UBFS with eight semantic features (UBFS-28) improves detection by up to 13% on real data. Distillation sensitivity analysis shows detection degrades monotonically with attack intensity for FOCUSED/BROAD/COT profiles, while HYDRA remains near-chance regardless of scale. Adversarial evasion testing reveals that mimicry attacks degrade detection by up to 25%, with Memory Poisoning (ASI05) most vulnerable. Temporal window ablation finds that 10 spans are optimal for Isolation Forest on both TRAIL and ATBench. Mapping five MITRE ATLAS techniques [1] into the UBFS confirms that the framework generalises beyond OWASP: all five techniques are strongly detectable (>0.93 AUC-ROC for IF/LSTM), introducing no new blind spots. The detection models port across domains. So do their biases, and so do their blind spots.
OWASP, insider threat detection, cross-domain transfer, MITRE ATLAS, UBFS, AI agent monitoring, anomaly detection, AI governance
OWASP, insider threat detection, cross-domain transfer, MITRE ATLAS, UBFS, AI agent monitoring, anomaly detection, AI governance
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
