Skip to content

Workflow design and references

Reviewed on 2026-10-03. These references inform the workflow; TrustLLM does not claim feature parity with these larger frameworks or copy their benchmark definitions.

Reference Practice adopted in this update
lm-evaluation-harness Separate model-backend dependencies, explicit task discovery, CLI and Python entry points
Inspect logs and parallelism Per-sample records, explicit run status, configurable concurrency and recoverable runs
EvalScope quickstart Side-by-side API/local examples, small-sample trials and reusable configuration
Transformers chat templates Format local-model inputs with the tokenizer's native chat template

What this update validates

Offline tests cover download errors, configuration, API transport, bounded retries, model-independent input/output handling, resume compatibility, incomplete-response rejection and report escaping. A real loopback HTTP server exercises the CLI request path. A tiny locally constructed causal model exercises CPU inference and saved-checkpoint loading without downloading external weights.

What remains separate

The original six-dimension scoring methods and prompts remain. This update does not claim to reproduce every published benchmark score or validate every hosted model provider. Full GPU runs, paid judge/embedding integrations, multi-node inference, multimodal tasks and a live dashboard are outside these checks. A generation report's completion rate must never be interpreted as a trustworthiness score.