LLM-as-a-Judge Category Evaluator

import langwatch

df = langwatch.dataset.get_dataset("dataset-id").to_pandas()

evaluation = langwatch.evaluation.init("my-incredible-evaluation")

for index, row in evaluation.loop(df.iterrows()):
    # your execution code here       
    evaluation.run(
        "langevals/llm_category",
        index=index,
        data={
            "input": row["input"],
            "output": output,
            "contexts": row["contexts"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    },
    "raw_response": {},
    "error_type": "<string>",
    "traceback": [
      "<string>"
    ]
  }
]

POST

langevals

llm_category

evaluate

import langwatch

df = langwatch.dataset.get_dataset("dataset-id").to_pandas()

evaluation = langwatch.evaluation.init("my-incredible-evaluation")

for index, row in evaluation.loop(df.iterrows()):
    # your execution code here       
    evaluation.run(
        "langevals/llm_category",
        index=index,
        data={
            "input": row["input"],
            "output": output,
            "contexts": row["contexts"],
        },
        settings={}
    )

[
  {
    "status": "processed",
    "score": 123,
    "passed": true,
    "label": "<string>",
    "details": "<string>",
    "cost": {
      "currency": "<string>",
      "amount": 123
    },
    "raw_response": {},
    "error_type": "<string>",
    "traceback": [
      "<string>"
    ]
  }
]

Authorizations

X-Auth-Token

string

header

required

Body

application/json

data

object

required

Show child attributes

data.input

string

The input text to evaluate

data.output

string

The output text to evaluate

data.contexts

string[]

Context information for evaluation

settings

object

Evaluator settings

Show child attributes

settings.model

string

default:openai/gpt-5

The model to use for evaluation

settings.max_tokens

number

default:8192

max_tokens setting

settings.prompt

string

default:You are an LLM category evaluator. Please categorize the message in one of the following categories

The system prompt to use for the LLM to run the evaluation

settings.categories

object[]

The categories to use for the evaluation

Response

Successful evaluation

status

enum<string>

required

Available options:

processed,

skipped,

error

score

number

Evaluation score

passed

boolean

Whether the evaluation passed

label

string

Evaluation label

details

string

Additional details about the evaluation

cost

object

Show child attributes

cost.currency

string

required

cost.amount

number

required

raw_response

object

Raw response from the evaluator

error_type

string

Type of error if status is 'error'

traceback

string[]

Error traceback if status is 'error'

⌘I

Get Started

Agent Simulations

Observability

Evaluation

Prompt Management

Platform

Examples & Cookbooks

LLM-as-a-Judge Category Evaluator

Authorizations

Body

Response