Inspect AI examples

Some examples of using Inspect AI to evaluate task performance:

Project Structure

├── data/                    # Evaluate datasets
├── logs/                    # Evaluation results and logs
├── classifier.py            # Binary classification evaluation
├── intent_classifier.py     # Intent classification evaluation
├── .env                     # API configuration
└── pyproject.toml           # Project dependencies

Setup

Requirements

Tested on Python 3.13
uv package manager

Installation

Install dependencies:
```
uv sync
```

Configuration

Create a .env file in the project root

Add your API configuration:

export BEDROCK_API_KEY=<bedrock_api_key>
export BEDROCK_BASE_URL=https://bedrock-runtime.<aws_region>.amazonaws.com/openai/v1

Evaluation

Source environment variables:
```
source .env
```
Run an evaluation:
```
python <evaluation_file.py>
```
View results:
```
inspect view
```

Name		Name	Last commit message	Last commit date
Latest commit History 7 Commits
data		data
.gitignore		.gitignore
.python-version		.python-version
LICENSE		LICENSE
README.md		README.md
classifier.py		classifier.py
intent_classifier.py		intent_classifier.py
pyproject.toml		pyproject.toml
uv.lock		uv.lock

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

Inspect AI examples

Project Structure

Setup

Requirements

Installation

Configuration

Evaluation

About

Uh oh!

Languages

License

corvuslee/inspect-ai

Folders and files

Latest commit

History

Repository files navigation

Inspect AI examples

Project Structure

Setup

Requirements

Installation

Configuration

Evaluation

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Languages