← All tags

#benchmark

31 posts

Claude Skill 1

Task Compass Skill: deterministic, auditable task routing for AI agents and OpenClaw

Python MIT Updated 2w ago
Claude Skill 2

Workbench for Agent Skills — lint, test, benchmark, and gate-install SKILL.md skills for Claude Code, Codex, Cursor, Gemini, and more

TypeScript MIT Updated 1w ago
Claude Skill 75

[COLM'26] SkillLearnBench is the first benchmark for evaluating continual learning methods that automatically generate agent skills.

Python MIT Updated 2w ago
MCP Server 35

MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols

Python MIT Updated 4mos ago
MCP Server 8

Private, local-first AI assistant for Windows - use your own model, keep durable memory, and approve every sensitive tool.

C# Apache-2.0 Updated 1w ago
MCP Server 6

Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, lay...

Python MIT Updated 2mos ago
MCP Server 11

Open benchmark for AI coding agents on SWE-bench Verified. Compare resolution rates, cost, and unique wins.

Shell MIT Updated 2mos ago
Claude Skill 9

Agent control plane for governed AI coding: validate changes, enforce policy gates, track findings, proofs, and evals based on your habits.

Elixir NOASSERTION Updated 1w ago
MCP Server 4

Portable memory layer for AI agents -*pre-release

Rust Apache-2.0 Updated 2w ago
MCP Server 24

Real-time trustworthiness evaluation and safety interception for AI agents. Semantic analysis, safe alternative suggestions, multi-step atta...

Python NOASSERTION Updated 1mo ago