Skip to content

OPEN RESEARCH · ICML 2024

LLM trustworthiness, measured.

Evaluate the models behind your applications across six dimensions. TrustLLM brings benchmark data, local and API generation, and the original research scorers into one Python workflow.

Run your first evaluation Integrate with an AI agent

01 / Set up

Install only the components you need. Download benchmark data from Python or the CLI.

Installation & data →

02 / Run a model

Use local Hugging Face weights or a compatible API. Keep checkpoints and resume interrupted runs.

Model backends →

03 / Read the evidence

Score full response sets, inspect each metric, and keep the experiment settings alongside the results.

Scoring & results →

What does TrustLLM evaluate?

Dimension Questions the benchmark explores
Truthfulness Does the model produce misinformation, hallucinate, or agree with a user's false premise?
Safety How does it respond to jailbreaks, misuse requests, and harmless prompts that trigger excessive refusal?
Fairness Does it express stereotypes, demographic preferences, or disparaging judgments?
Robustness How do perturbations and out-of-distribution inputs affect its responses?
Privacy Does it recognize privacy concerns or disclose sensitive information?
Ethics How does it reason about moral judgments and choices?

Explore the datasets and original metrics. TrustLLM reports individual metrics with their original directions and scales; it does not collapse them into an overall trust score.

One workflow, two model backends

python -m pip install "trustllm @ git+https://github.com/HowieHwong/TrustLLM.git@main#subdirectory=trustllm_pkg"
python -m trustllm download --output data
python -m trustllm tasks

The GitHub source installation above provides the current 0.4 workflow. API generation uses the lightweight base package; install the local extra for local inference and eval for scoring. The first-run guide covers credentials, examples and model requirements.

An AI agent can orchestrate this workflow through the CLI or Python, then read JSON artifacts. See the agent integration guide for commands, exit codes and output contracts. The benchmark evaluates model responses; multi-step agent trajectories, tool-use correctness and agent memory are outside its current scope.

Research and reproducibility

TrustLLM accompanies TrustLLM: Trustworthiness in Large Language Models, published at ICML 2024. Prompts and original scoring methods remain part of the research benchmark. New backends, model versions and judge settings can change results; record them when comparing experiments.

Paper · Dataset · Published leaderboard · Validation scope

Documentation is in English. The repository also offers introductions in 简体中文, 繁體中文, 日本語, 한국어, Español and Français. README translations do not translate the benchmark data.