AI Platform & Agentic Systems Engineer · Independent Consultant
Agentic AI · AI Platforms · MLOps/LLMOps · AI SRE & AIOps · Kubernetes/OpenShift · HPC
I build autonomous AI that acts on infrastructure — with a human at the gate.
I design and operate the platforms that let AI agents and ML models run in production without anyone losing control of them: approval gates, scoped RBAC, model registries, replay and audit trails. I came to it the unusual way round — six years running the IT and networks of a 1,000+ MW combined-cycle power plant, where there is no staging environment and a bad change measures in megawatts, then a PhD in high-performance computing.
Note
Now: Independent Consultant — AI Platform & MLOps Engineer, with a current engagement for a European AI/HPC infrastructure client. Based in Italy · remote across Europe & internationally · open to full-time roles, contracts and freelance engagements → mskazemi.com/hire
My five flagship projects each cover one stage of the same problem — putting AI into production on real infrastructure without giving up control:
flowchart LR
A["<b>Train · version · serve</b><br/>MLOps on HPC<br/><i>ExaMLOps</i>"] --> B["<b>Investigate · propose · act</b><br/>behind human approval<br/><i>KubeIntellect</i>"]
B --> C["<b>Evaluate permissions</b><br/>policy violation = fail<br/><i>AOBench</i>"]
C --> V["<b>Verify before merge</b><br/>independent verifiers<br/><i>idkmesh</i>"]
V --> D["<b>Replay · diff · audit</b><br/>evidence for every run<br/><i>NovaFabric</i>"]
🧠 KubeIntellect — an AI SRE for Kubernetes, human-governed
My role: creator and first author of the peer-reviewed paper.
| Problem | Diagnosing a live Kubernetes fault means correlating kubectl, Prometheus and Loki by hand under time pressure — and any tool that fixes it automatically is a tool nobody will run in production. |
| What I built | Ask the cluster a question in plain English. It gathers live evidence, works out the root cause, and executes the remediation — pausing for explicit human approval before it changes anything. |
| Proof | Peer-reviewed in the Journal of Grid Computing (2026), 10.1007/s10723-026-09837-6 · live demo · pip install kubeintellect |
The shipping implementation is a LangGraph supervisor with PostgreSQL checkpoints that works from live kubectl, Helm, Prometheus and Loki evidence and routes every mutating action through human approval. The published architecture went further — a code-generator agent that wrote and validated new tools at runtime; it was evaluated in the paper and deliberately dropped when the system was simplified for production.
Python · LangGraph · FastAPI · Kubernetes · PostgreSQL · Prometheus · Loki
🏭 ExaMLOps — production MLOps platform for HPC and supercomputers
My role: architect and lead developer.
| Problem | Sixteen partners on a EuroHPC consortium each needed to train, version, govern and serve models on a Tier-0 supercomputer — with no shared platform to do it on. |
| What I built | An end-to-end MLOps platform: a partner registers a model, the platform trains, versions, governs and serves it, behind a sysadmin approval gate. |
Prefect · MLflow · Ray Serve · Slurm · FastAPI · React
🛡️ AOBench — evaluation and permission infrastructure for AI agents
My role: creator.
| Problem | Agent benchmarks score whether the answer looked correct. In operations, an agent that reaches the right answer by exceeding its permissions has failed. |
| What I built | A role-aware, permission-enforced benchmark for LLM agents doing real HPC operations work: a policy violation hard-fails the task, however correct the output. |
| Proof | 88 tasks across 10 categories × 5 roles · archived with a DOI · paper under review. |
Python · MCP · Slurm · RBAC
✅ idkmesh — verification infrastructure for AI agents
My role: creator and maintainer.
| Problem | Teams gate AI-agent work behind review panels — LLM judges, CI checks, human reviewers — and count every vote as independent. When their errors are correlated, a panel can be worth far fewer votes than it has members, and bad changes pass. |
| What I built | An open-source toolkit for AI-agent verification: a worker's self-report never counts as acceptance, independent verifiers check every result against versioned task contracts with bound provenance, and idkmesh gate-audit measures how many independent votes a review panel is really worth and audits review gates for correlated errors — with reproducible evidence for every accept/reject. |
| Proof | Measured, not simulated: in a 25-verifier panel — each verifier a program, 72 candidates with ground truth from hidden tests — majority vote was worth an effective 1.00 of 25 independent votes (E017) · Apache-2.0 · installable idkmesh CLI · every claim tracked in a CI-validated Capability Truth Matrix · built in the open, good first issues tagged |
Python · MCP · A2A · LLM-as-a-judge · GitHub Actions
| Project | What it does | Evidence | Stack |
|---|---|---|---|
| NovaFabric | Open-source, self-hosted replay and evidence infrastructure for AI agents and agentic systems — captures agent executions as portable Run Capsules for replay, behavioral/structural diff, lineage, cryptographic provenance, assurance and audit. | Apache-2.0 · beta · novafabric.ai | Python, OpenTelemetry, Kubernetes, SLURM |
| YazSes | Offline-by-default voice dictation — nothing leaves your machine by default. Hold a key, speak, release — speech-to-text runs on your own CPU and the words are typed into whatever window has focus. Works on Wayland, where most dictation tools silently fail. | Apache-2.0 · cross-platform · built in the open by outside contributors, good first issues tagged · measured accuracy, published method | Python, faster-whisper, Linux/macOS/Windows |
| kube-q | CLI and Python SDK for KubeIntellect — pip install kube-q |
Streaming responses, Rich TUI · AGPL-3.0 | Python |
| Project | What it does | Evidence |
|---|---|---|
| GRAAFE | Graph neural network that anticipates compute-node anomalies on exascale HPC — trained offline, served online through a Kubeflow pipeline on live telemetry. | Published, FGCS 2024 · CINECA Marconi100 |
| HazardNet | Thermal-hazard prediction for datacenters, over a year of telemetry from 3,312 nodes of CINECA's Marconi A2. Six-hour horizon, chosen with the facility manager. | Published, FGCS 2024 · 1 GB dataset on Zenodo, CC BY 4.0 |
More open-source work — MCP, RAG, Slurm, Kubernetes and HPC tooling
| Repository | What it is |
|---|---|
| mcp-zenodo | Zenodo MCP server — tool-based LLM integration with the Zenodo open-access research repository via the Model Context Protocol. |
| AI4HPC | LLM + RAG system for querying HPC platform documentation and code — crawling, chunking, embeddings, vector search, LangChain retrieval. |
| ai-agent-systems-course | Hands-on course for building AI agent systems in Python: MCP, A2A, LangGraph/LangChain, multi-agent orchestration with Ollama. |
| EnergetiScope | Predicts the energy consumption of Kubernetes workloads from their specs, before they run — ground-truth labels from Kepler (eBPF) via Prometheus. |
| m100-silicon-binning | Recovers per-die CPU core-harvest maps from out-of-band BMC telemetry — 1,962 POWER9 sockets, reproducible from public M100 ExaData. |
| SLURM_Simulator | Automated multi-node Slurm cluster on Vagrant + VirtualBox for local HPC testing, with MUNGE auth and MySQL job accounting. |
| Vagrant-Kubernetes-ROS2-Deployment | ROS 2 nodes on a multi-node Kubernetes cluster provisioned with Vagrant, monitored with Prometheus and Grafana. |
| Area | Tools |
|---|---|
| Platform & infrastructure | Kubernetes · OpenShift · Helm · Terraform · Docker · Linux · Azure |
| Agentic AI & LLM | LangGraph · MCP · A2A · RAG · vLLM · FastAPI |
| MLOps & pipelines | MLflow · Kubeflow · Prefect · Ray Serve · KServe |
| HPC | Slurm · MPI · OpenMP |
| Observability | Prometheus · Grafana · Loki · OpenTelemetry |
| ML systems | Python · PyTorch · GNNs · TCN/LSTM · anomaly detection · time-series telemetry at datacenter scale |
PhD: Design, Analysis, and Management of High-Performance Computing Systems · University of Bologna (2018–2022)
| Selected peer-reviewed work | Venue | Year |
|---|---|---|
| KubeIntellect: A Modular LLM-Orchestrated Agent Framework for Kubernetes Management | Journal of Grid Computing | 2026 |
| GRAAFE: GRaph Anomaly Anticipation Framework for Exascale HPC Systems | FGCS | 2024 |
| M100 ExaData: A Data Collection Campaign on CINECA's Marconi100 Tier-0 Supercomputer | Nature Scientific Data | 2023 |
Three open datasets, 26 GB in total: M100 ExaData, the HazardNet thermal dataset (first author) and PM100 — free to download, no registration.
Reviewing and programme committees
Reviewer for IEEE TCAD · FGCS · Journal of Grid Computing · SC · ACM CF · DATE · PDP · AsHES. PC member: PDP 2025 · PDP 2026 · AsHES 2026.
Full publication list and current citation counts → Google Scholar · ORCID · dblp
|
For AI Platform · Agentic Systems · MLOps/LLMOps · Kubernetes/OpenShift roles. Based in Italy; available for remote roles across Europe and internationally, with working-hour overlap by arrangement (Europe/Rome, CET/CEST). |
Each engagement starts with a fixed-price audit, so you see the work before committing to a project:
|
Website · About · LinkedIn · GitLab · PyPI · Mastodon · Scholar · ORCID





