Token导航 LogoToken导航TokenDH.com
研究检索external-servicegithub未标认证来源可访问许可证需确认审计提醒

exploring-llm-evaluationsexploring LLM evaluations 搜索

Agent Skill

exploring-llm-evaluations 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

682

周安装

29

GitHub Stars

31

下载量

239
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:exploring-llm-evaluations(exploring LLM evaluations 搜索)
来源仓库:https://github.com/posthog/skills
仓库路径:skills/exploring-llm-evaluations
安装命令:
npx skills add https://github.com/posthog/skills --skill exploring-llm-evaluations
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/posthog/skills --skill exploring-llm-evaluations

简介

exploring-llm-evaluations 管理 deterministic rule 与 LLM judge 双轨评估体系。

  • 支持格式校验、长度控制与主观质量评分混合使用。
  • 适用于生成内容质量管控与迭代优化闭环建设。
  • hog 规则优先于 llm_judge,以降低成本并提升可复现性。
  • 适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。

SKILL.md

Exploring LLM evaluations

PostHog evaluations score $ai_generation events. Each evaluation is one of two types, both first-class:

  • hog — deterministic Hog code that returns true/false (and optionally N/A). Best for objective rule-based checks: format validation (JSON parses, schema matches), length limits, keyword presence/absence, regex patterns, structural assertions, latency thresholds, cost guards. Cheap, fast, reproducible — no LLM call per run. Prefer this when the criterion can be expressed as code.
  • llm_judge — an LLM scores generations against a prompt you write. Best for subjective or fuzzy checks: tone, helpfulness, hallucination detection, off-topic drift, instruction-following. Costs an LLM call per run and requires AI data processing approval at the org level.

Results from both types land in ClickHouse as $ai_evaluation events with the same schema, so the read/query/summary workflows are identical regardless of evaluator type — the only thing that changes is whether $ai_evaluation_reasoning was written by Hog code or by an LLM.

This skill covers the full lifecycle: list/inspect/manage evaluation configs (Hog or LLM judge), run them on specific generations, query individual results, and get an AI-generated summary of pass/fail/N/A patterns across many runs.

Tools

ToolPurpose
posthog:evaluations-getList/search evaluation configs (filter by name, enabled flag)
posthog:evaluation-getGet a single evaluation config by UUID
posthog:evaluation-createCreate a new llm_judge or hog evaluation
posthog:evaluation-updateUpdate an existing evaluation (name, prompt, enabled, …)
posthog:evaluation-deleteSoft-delete an evaluation
posthog:evaluation-runRun an evaluation against a specific $ai_generation event
posthog:evaluation-test-hogDry-run Hog source against recent generations (no save)
posthog:llm-analytics-evaluation-summary-createAI-powered summary of pass/fail/N/A patterns across runs
posthog:execute-sqlAd-hoc HogQL over $ai_evaluation events
posthog:query-llm-traceDrill into the underlying generation that an evaluation scored

The first seven evaluation-* tools are hand-coded; llm-analytics-evaluation-summary-create is generated from products/llm_analytics/mcp/tools.yaml.

Event schema

Every run of an evaluation emits an $ai_evaluation event. Key properties:

PropertyMeaning
$ai_evaluation_idUUID of the evaluation config
$ai_evaluation_nameHuman-readable name
$ai_target_event_idUUID of the $ai_generation event being scored
$ai_trace_idParent trace ID (for jumping to the trace UI)
$ai_evaluation_resulttrue = pass, false = fail
$ai_evaluation_reasoningFree-text explanation (set by the LLM judge or Hog code)
$ai_evaluation_applicablefalse when the evaluator decided the generation is N/A

When $ai_evaluation_applicable = false, the run counts as N/A regardless of $ai_evaluation_result. For evaluations that don't support N/A, this property may be null — treat null as "applicable".

Workflow: investigate why an evaluation is failing

Works the same way for llm_judge and hog evaluations — the differences only matter when you eventually go to fix the evaluator (edit the prompt vs. edit the Hog source).

Step 1 — Find the evaluation

posthog:evaluations-get
{ "search": "hallucination", "enabled": true }

Look at the returned id, name, evaluation_type, and either:

  • evaluation_config.prompt for an llm_judge
  • evaluation_config.source for a hog evaluator

The Hog source is the ground truth for why a hog evaluator passes or fails — read it before assuming the failure is in the generation.

Step 2 — Get the AI-generated summary

posthog:llm-analytics-evaluation-summary-create
{
  "evaluation_id": "<uuid>",
  "filter": "fail"
}

Returns:

  • overall_assessment — natural-language summary
  • fail_patterns — grouped patterns with title, description, frequency, and example_generation_ids
  • pass_patterns and na_patterns — same shape, populated when filter includes them
  • recommendations — actionable next steps
  • statisticstotal_analyzed, pass_count, fail_count, na_count

The endpoint analyses the most recent ~250 runs (EVALUATION_SUMMARY_MAX_RUNS). Results are cached for one hour per (evaluation_id, filter, set_of_generation_ids). Pass force_refresh: true to recompute.

Compare filters in two calls to spot what's distinctive about failures vs passes:

posthog:llm-analytics-evaluation-summary-create
{ "evaluation_id": "<uuid>", "filter": "pass" }

Then diff the pass_patterns against the fail_patterns from Step 2.

Step 3 — Drill into example failing runs

Each pattern surfaces example_generation_ids. Pull the underlying trace for the most representative example:

posthog:query-llm-trace
{ "traceId": "<trace_id>", "dateRange": {"date_from": "-30d"} }

(If you only have a generation ID, query for it via execute-sql first to find the parent trace ID — see below.)

Step 4 — Verify the pattern with raw SQL

The summary is LLM-generated and should be verified. Use execute-sql to count and spot-check:

posthog:execute-sql
SELECT
    properties.$ai_target_event_id AS generation_id,
    properties.$ai_trace_id AS trace_id,
    properties.$ai_evaluation_reasoning AS reasoning,
    timestamp
FROM events
WHERE event = '$ai_evaluation'
    AND properties.$ai_evaluation_id = '<evaluation_uuid>'
    AND properties.$ai_evaluation_result = false
    AND (
        properties.$ai_evaluation_applicable IS NULL
        OR properties.$ai_evaluation_applicable != false
    )
    AND timestamp >= now() - INTERVAL 7 DAY
ORDER BY timestamp DESC
LIMIT 25

The N/A guard (IS NULL OR!= false) is important — it matches the same logic the backend uses to bucket runs.

Workflow: run an evaluation against a specific generation

Use this when the user pastes a trace/generation URL and asks "what would evaluation X say about this?".

posthog:evaluation-run
{
  "evaluationId": "<eval_uuid>",
  "target_event_id": "<generation_event_uuid>",
  "timestamp": "2026-04-01T19:39:20Z",
  "event": "$ai_generation"
}

The timestamp is required for an efficient ClickHouse lookup of the target event. Pass distinct_id if you have it — it speeds up the lookup further.

Workflow: build and test a new evaluator

Hog evaluator (deterministic, code-based)

Reach for this first when the criterion is rule-based — it's cheaper, faster, and reproducible. Prototype with evaluation-test-hog (no save):

posthog:evaluation-test-hog
{
  "source": "return event.properties.$ai_output_choices[1].content contains 'sorry';",
  "sample_count": 5,
  "allows_na": false
}

The handler returns the boolean result for each of the most recent N $ai_generation events. Iterate on the source until it behaves as expected, then promote it via evaluation-create:

posthog:evaluation-create
{
  "name": "Output is valid JSON",
  "description": "Fails when the assistant message can't be parsed as JSON",
  "evaluation_type": "hog",
  "evaluation_config": {
    "source": "let raw := event.properties.$ai_output_choices[1].content; try { jsonParseStr(raw); return true; } catch { return false; }"
  },
  "output_type": "boolean",
  "enabled": true
}

Hog evaluators have full access to the event and its properties — common patterns include schema validation, length/token limits, regex matches, and tool-call shape checks. Because they're deterministic, results are reproducible across reruns and trivially diff-able.

LLM-judge evaluator (subjective, prompt-based)

Use this when the criterion is fuzzy and a code rule would be brittle (tone, factuality, helpfulness, on-topic-ness). There's no equivalent of evaluation-test-hog for LLM judges — the typical loop is to create the evaluator with enabled: false, run it manually against a handful of representative generations via evaluation-run, inspect the results, refine the prompt with evaluation-update, and then flip enabled: true when you're satisfied:

posthog:evaluation-create
{
  "name": "Response stays on-topic",
  "description": "LLM judge — fails if the assistant changes topic from the user's question",
  "evaluation_type": "llm_judge",
  "evaluation_config": {
    "prompt": "You are evaluating whether the assistant's reply stays on-topic relative to the user's most recent question. Return true if it does, false if the assistant changed the subject. Return N/A if the user did not actually ask a question."
  },
  "output_type": "boolean",
  "output_config": { "allows_na": true },
  "model_configuration": {
    "provider": "openai",
    "model": "gpt-5-mini"
  },
  "enabled": false
}

Then dry-run against a known-good and a known-bad generation:

posthog:evaluation-run
{
  "evaluationId": "<new_eval_uuid>",
  "target_event_id": "<generation_uuid>",
  "timestamp": "2026-04-01T19:39:20Z"
}

LLM judges require organisation AI data processing approval. Hog evaluators do not.

Workflow: manage the evaluation lifecycle

ActionTool
Add a Hog evaluatorevaluation-create with evaluation_type: "hog" and evaluation_config.source
Add an LLM-judge evaluatorevaluation-create with evaluation_type: "llm_judge", evaluation_config.prompt, and a model_configuration
Tweak the source or promptevaluation-update (edits evaluation_config.source for Hog, evaluation_config.prompt for LLM judge)
Toggle N/A handlingevaluation-update with output_config.allows_na
Disable temporarilyevaluation-update with enabled: false
Removeevaluation-delete (soft-delete via PATCH {deleted: true})

llm_judge evaluations require AI data processing approval at the org level (is_ai_data_processing_approved). The same gate applies to llm-analytics-evaluation-summary-create. Hog evaluations do not require this gate — they run as plain code on the ingestion pipeline.

When to use Hog vs LLM judge

Reach for Hog by default. Switch to LLM judge only when the criterion can't be expressed as code.

Use Hog when…Use LLM judge when…
The check is structural (JSON parses, schema matches)The check is about meaning (on-topic, helpful, factual)
You need a deterministic, reproducible resultA small amount of judgement variability is acceptable
The criterion is cheap to computeThe criterion requires reading and understanding text
You can't get AI data processing approvalYou have approval and the criterion is genuinely fuzzy
You need to enforce a hard limit (length, cost, etc.)You need to rate a quality dimension
You want sub-millisecond evaluationA few hundred milliseconds + LLM cost are acceptable

A common pattern is to layer them: a Hog evaluator gates obvious format/length violations cheaply, and an LLM-judge evaluator only fires on the generations that pass the Hog gate (via conditions).

Investigation patterns

The summarisation tool works the same way regardless of whether the evaluator is hog or llm_judge — it analyses the resulting $ai_evaluation events, not the evaluator itself. The fix path differs (edit Hog source vs. edit prompt) but the diagnosis is identical.

"Why is evaluation X suddenly failing more?"

  1. evaluations-get — confirm the evaluation is still enabled and unchanged (compare evaluation_config.source or evaluation_config.prompt to the version you expect)
  2. llm-analytics-evaluation-summary-create with filter: "fail" — get the dominant failure patterns and example IDs
  3. SQL count of fails per day to confirm the regression window: SELECT toDate(timestamp) AS day, count() AS fails FROM events WHERE event = '$ai_evaluation' AND properties.$ai_evaluation_id = '<uuid>' AND properties.$ai_evaluation_result = false AND timestamp >= now() - INTERVAL 30 DAY GROUP BY day ORDER BY day
  4. Drill into a representative trace per pattern via query-llm-trace

"Are passes and fails caused by the same root content?"

  1. Generate two summaries: one with filter: "pass", one with filter: "fail"
  2. If pass_patterns and fail_patterns describe similar content:

- For an llm_judge: the prompt or rubric is probably ambiguous — reword evaluation_config.prompt and use evaluation-update - For a hog evaluator: the rule is probably under- or over-matching — read the source via evaluation-get, narrow the predicate, and retest with evaluation-test-hog before pushing the fix via evaluation-update

"Did a Hog evaluator regression after a code change?"

Hog evaluators are reproducible — if the source hasn't changed, identical inputs should yield identical outputs. When fail rates jump for a Hog evaluator:

  1. evaluation-get — note the current source and updated_at
  2. Spot-check the latest failing runs with the SQL query from Step 4 above
  3. Re-run the source against those exact generations using evaluation-test-hog with a modified conditions filter that targets them
  4. If the test results match the live results, the change is in the *generations*, not the evaluator (a model upgrade, prompt change upstream, etc.) — investigate the producer
  5. If they diverge, the evaluator was edited; check git history of the source field via the activity log

"What kinds of generations does this evaluator skip as N/A?"

posthog:llm-analytics-evaluation-summary-create
{ "evaluation_id": "<uuid>", "filter": "na" }

Inspect na_patterns to see whether the N/A logic is doing the right thing. If a pattern in na_patterns looks like something that should have been scored:

  • For an llm_judge: the applicability instruction in the prompt is too broad — narrow it
  • For a hog evaluator with output_config.allows_na: true: the source is returning null (or whatever the N/A signal is) too eagerly — tighten the precondition

"Score this single generation right now"

evaluation-run with the trace's generation ID and timestamp. Useful for spot-checking or wiring evaluations into a larger agent loop.

Constructing UI links

  • Evaluations list: https://app.posthog.com/llm-analytics/evaluations
  • Single evaluation: https://app.posthog.com/llm-analytics/evaluations/<evaluation_id>
  • Underlying generation/trace: see the exploring-llm-traces skill's URL conventions

Always surface the relevant link so the user can verify in the UI.

Tips

  • The summary tool is rate-limited (burst, sustained, daily) and caches results for one hour — repeated calls with the same (evaluation_id, filter) are cheap; use force_refresh: true only when you genuinely need fresh analysis
  • Pass generation_ids: [...] to scope a summary to a specific cohort of runs (max 250)
  • The statistics block in the summary response is computed from raw data, not the LLM — trust those counts even if a pattern's frequency field is qualitative
  • For rich filtering not supported by evaluations-get (e.g. by author or model configuration), fall back to execute-sql against the evaluations Postgres table or the $ai_evaluation ClickHouse events
  • When showing failure patterns to the user, always include 1-2 example trace links so they can validate the pattern visually
  • Hand-coded evaluation-* tools and the codegen llm-analytics-evaluation-summary-create share the same scope object (llm_analytics) — evaluation:read for read tools, evaluation:write for mutating tools, and llm_analytics:write for the summary tool
  • Hog evaluators are reproducible — if you suspect a regression, evaluation-test-hog with the suspect source against the failing generations is the fastest way to bisect whether the change is in the evaluator or in the producer of the generations
  • LLM-judge evaluators are non-deterministic across reruns; expect 1-5% noise even with a fixed prompt and model. If you're chasing a small regression in fail rate, prefer Hog or pin a deterministic provider/seed in the model_configuration

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

36.55%
按下载量换算87

Claude

31.87%
按下载量换算76

Cursor

17.1%
按下载量换算41

Gemini CLI

8.86%
按下载量换算21

安全审计

Gen Agent Trust Hub

可疑

Socket

通过

Snyk

通过

权限和风险

external-service

该 Skill 可能调用第三方服务、云服务或外部模型 API,使用前需要确认账号、额度、数据发送范围和服务条款。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills