Build reliable customer-facing AI agents with Parlant: an interaction control harness optimized for controlled, consistent, and predictable LLM interactions.
-
Updated
Jul 12, 2026 - Python
Build reliable customer-facing AI agents with Parlant: an interaction control harness optimized for controlled, consistent, and predictable LLM interactions.
Catch your AI's mistakes and blind spots before your customers or regulators do. iFixAi runs 45 inspections, 32 graded core plus 13 extended for frontier risks like sabotage, sandbagging, and oversight evasion. It returns a letter grade in under 5 minutes. Industry and model agnostic.
PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022
Code accompanying the paper Pretraining Language Models with Human Preferences
[AAAI'25 Oral] "MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector".
Adam Framework for OpenClaw — 5-layer persistent memory and identity architecture for AI agents. Production-validated over 353+ sessions. First documented case of emergent values in persistent AI, quantum-verified on IBM hardware.
[ICLR 2026] - Official repo for the paper: "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models"
Official Implementation of Nabla-GFlowNet (ICLR 2025)
Just like the elite potential of a high-drive Belgian Malinois, an AI system's raw capabilities are wasted when deployed without proper structure. The technological value is no longer found in creating the drive, but in mastering the leash. Synapptic gives your AI Assistant persistent memory that updates in real time, saving you tokens and time.
We asked 6 AIs about their own programming. All 6 said jailbreaking will never be fixed. Run it yourself — $2, 10 minutes.
[TMLR 2024] Official implementation of "Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics"
Decomposing and measuring evaluation awareness in existing benchmarks and our proposed EvalAwareBench.
The open-source repository for PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment, which provides a general personalized preference learning modeling framework.
Autonomous research engine for generating, testing, and governing auditable claims across science, proofs, and high-stakes projects.
Code and materials for the paper S. Phelps and Y. I. Russell, Investigating Emergent Goal-Like Behaviour in Large Language Models Using Experimental Economics, working paper, arXiv:2305.07970, May 2023
Scan your AI/ML models for problems before you put them into production.
SYNAPTICON is a research prototype at the intersection of neuro-hacking, non-invasive BCIs, and foundational models, probing new territories of human expression, aesthetics, and AI alignment. At its core lies a live "BrainWaves-to-NaturalLanguage-to-Aesthetics" system integrating EEG decoding with LLMs and Transformers.
No system can model its own source. Empirical proof: 6 AI architectures (GPT-4, Claude, Gemini, DeepSeek, Grok, Mistral) hit the same structural wall.
🐝 Distill personal semantics into verifiable AI skills and teachable assets.
[AAAI'26 Main🎉] Official code of "When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models"
Add a description, image, and links to the ai-alignment topic page so that developers can more easily learn about it.
To associate your repository with the ai-alignment topic, visit your repo's landing page and select "manage topics."