
RiskLab
Toolkit · ACL 2026 Demo
A controlled multi-agent framework for probing emergent social risks in LLM agent collectives, specified as topology–environment–protocol–agent–task quintuples.
|
Toolkits & Datasets Research ArtifactsOpen-source toolkits, benchmarks, and datasets from our work on trustworthy and aligned foundation models. ![]() RiskLabToolkit · ACL 2026 Demo A controlled multi-agent framework for probing emergent social risks in LLM agent collectives, specified as topology–environment–protocol–agent–task quintuples. ![]() ProbeLLMToolkit · ICML 2026 An automated probing framework that discovers structured LLM failure modes via hierarchical Monte Carlo Tree Search, tool-augmented test generation, and failure-aware clustering. ![]() IntraAISoftware Built around three goals: accessible AI for all, adaptive learning and growth, and trust with collective understanding.
Software — release soon
![]() EmoNestSoftware · NeurIPS 2025 Creative AI A generative AI framework for creating interactive, emotionally adaptive storytelling games.
Demo — will release in Dec. at conference
![]() ValueLenceDashboard The first unified platform for dynamic, fine-grained value probing of LLMs — value curation, probe generation, response collection, and multi-dimensional evaluation. ![]() SDE-HarnessToolkit Scientific Discovery Evaluation — an extensible framework designed to accelerate AI-powered scientific discovery. ![]() ChemOrchToolkit · NeurIPS 2025 Intelligent task orchestration for chemical research — turns a chemical task into high-quality instruction–response pairs. TrustEvalToolkit · NAACL 2025 Demo A modular, extensible toolkit for trust evaluation of generative foundation models across safety, fairness, robustness, privacy, and more. DataGenToolkit · ICLR 2025 An LLM-powered framework for generating diverse, accurate, and highly controllable text datasets. TrustLLMToolkit · ICML 2024 A Python package that helps you assess the trustworthiness of your LLM more quickly. MetaToolDataset · ICLR 2024 A benchmark for evaluating whether LLMs have tool-usage awareness and can correctly choose the right tool. |