Security (Sekyuritii セキュリティ) is the fifth of 10 Principles per Japan AI Guidelines for Business + addresses cybersecurity throughout AI lifecycle including adversarial attacks specific to ML + traditional cyber threats to AI infrastructure. (1) Security Principle Definition: (a) cybersecurity throughout AI lifecycle; (b) protection against adversarial + traditional attacks; (c) resilience + recovery capability; (d) supply chain security; (e) coordination with national cybersecurity framework. (2) AI-Specific Attack Surface: (a) Training Data Poisoning - adversarial manipulation of training data; (b) Backdoor Attacks - hidden triggers in trained models; (c) Adversarial Examples - imperceptible perturbations causing misclassification; (d) Model Extraction - reverse engineering proprietary models; (e) Model Inversion - reconstruction of training data; (f) Membership Inference - determining whether record in training set; (g) Prompt Injection (direct + indirect) - LLM-specific manipulation; (h) Jailbreak - safety guardrail bypass; (i) Data Exfiltration via LLM outputs; (j) Supply Chain - compromised foundation model + library + dataset. (3) MLSecOps + AI Security Pipeline: (a) Threat modelling per AI system (STRIDE + MITRE ATLAS + LINDDUN); (b) Secure SDLC for AI - design + development + deployment + operation + decommission; (c) AI-aware code review + SAST + DAST; (d) Model dependency scanning - foundation model + library + dataset; (e) Container + Kubernetes security; (f) Cloud AI service security; (g) Edge AI device security; (h) Federated Learning security. (4) Adversarial Robustness Testing: (a) Adversarial Example Generation - FGSM + PGD + AutoAttack + APGD; (b) Defensive Techniques - adversarial training + input preprocessing + detection; (c) Certifiable Robustness - randomized smoothing + interval bound propagation; (d) Distribution Shift Testing - covariate + concept shift; (e) Common Corruptions - ImageNet-C style; (f) Real-World Adversarial - patch attacks + physical-world adversarial. (5) Prompt Injection Defence: (a) Input Sanitisation + filtering; (b) Output Sanitisation + filtering; (c) System Prompt Hardening - delimiter + privilege separation; (d) Tool Use Sandboxing + permission system; (e) Anthropic Constitutional AI + RLHF safety training; (f) Indirect Prompt Injection awareness - URLs + tools + documents; (g) Multi-Step Attack chain analysis; (h) Red-Team continuous evaluation; (i) AISI prompt injection benchmarks. (6) Data Poisoning Defence: (a) Training Data Provenance + verification; (b) Anomaly Detection in training data; (c) Activation Clustering + Spectral Signatures for backdoor detection; (d) Certified Robust Training (e.g. randomized smoothing); (e) Differential Privacy training (limits influence of any single sample); (f) Trusted data sources + signed datasets; (g) Reproducible training pipeline + audit; (h) AISI poisoning benchmarks. (7) Model Extraction + IP Protection: (a) API rate limiting + query monitoring; (b) Watermarking model outputs; (c) Differential Privacy at inference; (d) Output noise injection; (e) Tier-based access control; (f) Membership inference defence - differential privacy + regularisation; (g) Patent + Trade Secret protection; (h) Cryptographic model protection - TEE deployment. (8) Supply Chain Security: (a) Foundation Model provenance - SBOM-like AI BOM; (b) Pre-trained model verification - hash + signature; (c) Open Source AI library security - vulnerability monitoring; (d) Dataset provenance + licensing verification; (e) Vendor + Sub-processor due diligence; (f) Continuous monitoring of dependencies; (g) Sigstore + SLSA for AI emerging. (9) AISI Red-Team Activities: (a) AISI conducting voluntary red-team exercises with frontier AI developers; (b) Dangerous capability evaluation - biosecurity + cybersecurity + autonomous + persuasion; (c) Methodology development + benchmark creation; (d) International coordination via AISI Network; (e) Industry voluntary participation; (f) Public + private red-team reports. (10) AI Incident + Vulnerability Response: (a) AI-Specific Incident Response Plan; (b) Coordinated Vulnerability Disclosure (CVD) for AI; (c) Bug Bounty programs including AI safety; (d) AISI + METI voluntary incident reporting; (e) Sector regulator notification where required; (f) Customer + affected stakeholder notification; (g) Lessons Learned + Industry Sharing. (11) National Cybersecurity Integration: (a) National Center of Incident Readiness and Strategy for Cybersecurity (NISC); (b) JPCERT/CC + Japan Vulnerability Notes (JVN); (c) Critical Infrastructure cybersecurity guidelines (METI); (d) Sector CERTs (Financial + Telecommunications + Healthcare); (e) International cooperation - FIRST + APCERT; (f) Coordination with AI Bill cybersecurity provisions. (12) International Standards: (a) ISO/IEC 27001 Information Security Management; (b) ISO/IEC 27002 Code of Practice; (c) ISO/IEC 42001 AI Management System; (d) NIST AI RMF; (e) MITRE ATLAS Adversarial Threat Landscape for AI Systems; (f) OWASP ML Top 10 + LLM Top 10 + AI Exchange; (g) OWASP Generative AI Security project. Coordinates with NISC National Center of Incident Readiness + JPCERT/CC + Japan Vulnerability Notes (JVN) + ISO/IEC 27001 + 27002 + 42001 + 23894 + NIST AI RMF + MITRE ATLAS + OWASP ML Top 10 + OWASP LLM Top 10 + OWASP AI Exchange + Anthropic + OpenAI + Google DeepMind security teams + AISI Network + Hiroshima AI Process + Critical Infrastructure cybersecurity guidelines + Cybersecurity Basic Act 2014 + FSA Cybersecurity Guidelines + Banking Act + sector CERTs. Japan AI Guidelines Security + Adversarial applies.
The graph holds this control, the 0 it maps to, and the evidence behind each claim, over MCP and REST.