← All tags

#agent-evaluation

18 posts

Claude Skill 1

Mide si tu agente de IA cumple las reglas que le escribiste. Lee el historial local de Claude Code y devuelve un porcentaje por regla. Sin i...

Python MIT Updated 1mo ago
Claude Skill 4

Official MutagenT skills for AI coding agents — Claude Code plugin marketplace for prompt optimization, evaluation, and observability.

Shell MIT Updated 1mo ago
Claude Skill 89

Production-grade Agent Skills for AI coding agents—composable workflows for planning, TDD, debugging, review, UI/UX, releases, incidents, an...

Python MIT Updated 4w ago
Claude Skill 5

Agent skills for trapstreet.run — set up the tp CLI, build solutions against an eval task, and author new tasks, from plain language. Claude...

3 skills Python MIT Updated 1mo ago
Library 46

🔁 Build reliable recurring AI-agent systems: 874 resources, 22 operational patterns, 22 loop contracts, 8 runtime starters, an interactive...

Python CC0-1.0 Updated 1mo ago
MCP Server 2

Production AI agent quality gate and risk control framework for LLMOps, agent evaluation, regression detection, gray release, audit, and obs...

Python MIT Updated 1mo ago
MCP Server 2

An automated red-teaming and reliability-auditing platform for AI agents - tests for prompt injection, tool hijacking and data exfiltration....

Python MIT Updated 2mos ago
MCP Server 7

A living world where agents exist as participants alongside NPCs, internal actors, real service APIs, budgets, policies, and consequences.

Python MIT Updated 2mos ago