MLflow Skills
Closes the loop for agent work: instrument, trace, evaluate, change, then verify the change helped.
Install
OfficialShips scriptsClaude Code: /plugin marketplace add anthropics/claude-plugins-official → /plugin install mlflow@claude-plugins-official
Bir LLM uygulamasına izleme eklenirken, iz üzerinden arıza aranırken ya da değerlendirme puanlayıcısı kurulurken
- Author
- MLflow
- License
- Apache-2.0
The pieces are ordinary MLflow — tracing for Python and TypeScript, trace retrieval, metric queries, failure analysis — but the arrangement is the point, because without a scorer an agent improvement is just an opinion. The scorer skill refuses to start from the catalogue: it works out what the application does, extracts a small set of atomic quality criteria from that, and only then picks the cheapest scorer that can judge each one, treating the user as the authority on what counts as wrong. Aiming for a runnable and inspectable prototype rather than a perfect first pass is what keeps the loop turning.
Similar skills
API Testing & Observability
SkillOpenAPI 3.1 documentation and a stateful mock server so clients can be built before the API exists.
Shell Scripting
SkillTwo shell agents that keep bash idioms out of POSIX sh scripts, with ShellCheck and Bats.
Accessibility Compliance
SkillWCAG 2.2 auditing with a validator agent that assumes the fix failed until a screenshot proves it.
Comments(0)
Sign in to comment