Proof-backed AI agent for checking suspicious job posts, recruiter messages, and apply links, now live with case study and pilot intake.
#ai-safety
81 posts
MCP servers expose tools with no information about what they actually do at runtime. mcpsafetywarden sits between your agent and any MCP ser...
MCP & Claude Code security scanner — threat-models plugins, MCP servers, hooks, skills & connectors with an LLM before you trust them. Catch...
Glass Box Framework — runtime constitutional verification for AI answers. Trust Cards with claim-level reasoning chains, formal ECS scoring,...
Runtime artifact existence & freshness verification for AI agent completion claims — a lightweight, zero-LLM MCP gate that source-binds 'don...
AI Red Teaming / AI Safety に関する日本語リソースのキュレーションリスト
A safer MySQL CLI for AI coding agents: connection profiles, SSH tunnels, and automatic sensitive-data masking before query output reaches C...
🛡️ A curated list of resources on agent skills security: attacks, defenses, frameworks, and benchmarks for securing AI agent tool use and sk...
Open-source prompt injection detector — 5 layers, 91.7% F1, ~27ms, offline, Apache 2.0
Deterministic policy language for AI agents. Z3 + TLA+ dual-engine formal verification. Runtime enforcement <1ms.
Pre-execution policy engine for AI agents. Every tool call checked before execution.
All AIs are sycophants.