TrialNet is a self-improving AI system with two layers:
- TrialNet Core β A neural network built from scratch with NumPy that uses a novel Try-and-Learn engine to remember and correct its mistakes.
- TrialNet LLM β A locally-running Large Language Model (Qwen2.5-1.5B) fine-tuned on Apple Silicon via MLX with a ChromaDB-backed mistake memory and an automated LLM-as-Judge for continuous self-correction.
| Feature | Traditional Models | TrialNet |
|---|---|---|
| Error handling | Forgotten after weight update | Error Memory Bank stores & prioritizes mistakes |
| Learning signal | Loss gradient only | Gradient + targeted mistake replay |
| Self-awareness | None | Mistake Pattern Analyzer discovers failure patterns |
| Weight updates | Gradient descent only | Gradient + Perturbation Explorer |
| LLM Memory | Static weights | ChromaDB RAG + continuous LoRA self-correction |
| Feedback | Manual retraining | Auto-judge scores every response, logs bad ones |
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β TrialNet System β
β β
β ββββββββββββββββββββββββ ββββββββββββββββββββββββββββββ β
β β TrialNet Core β β TrialNet LLM (Mac) β β
β β (NumPy from scratch)β β Qwen2.5-1.5B + MLX LoRA β β
β β β β β β
β β [Traditional SGD] β β [ChromaDB Memory Bank] β β
β β + β β β stores mistakes β β
β β [Try-and-Learn] β β β β β
β β ββββββββββββββββ β β [LLM-as-Judge] β β
β β β Error Memory β β β β scores every response β β
β β β Bank β β β β β β
β β ββββββββββββββββ€ β β [MLX LoRA Self-Correction]β β
β β β Mistake β β β β injects corrections β β
β β β Analyzer β β β into new adapter β β
β β ββββββββββββββββ€ β ββββββββββββββββββββββββββββββ β
β β β Perturbation β β β
β β β Explorer β β β
β β ββββββββββββββββ β β
β ββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
git clone https://github.com/Suraj-v23/trialnet.git
cd trialnet
pip install -r requirements.txtcd mac_llm_trialnet
pip install -r requirements_mac.txt# Hybrid mode β recommended
python train.py --mode hybrid --epochs 15
# Compare all three learning modes
python train.py --mode compare --epochs 15
# With live dashboard at http://localhost:5050
cd dashboard && python server.py
python train.py --mode hybrid --epochs 15 --dashboardpython evaluate.py --model saved_models/hybridcd mac_llm_trialnet
# Step 1 β Fine-tune on hybrid logic + coding curriculum
python 1_mac_finetune.py
# Step 2 β Chat (Auto-judge runs on every response)
python 2_mac_chatbot.py
# Step 3 β Self-correct (run after 10+ logged mistakes)
bash run_self_correction.shtrialnet/
βββ train.py # Core training script
βββ evaluate.py # Evaluation script
βββ requirements.txt # Core dependencies
β
βββ trialnet/ # NumPy neural network library
β βββ model.py # Main TrialNet class
β βββ utils.py # Data loading utilities
β βββ core/ # Neural network fundamentals
β β βββ tensor.py # Custom tensor ops (no PyTorch)
β β βββ layers.py # Dense, Dropout, BatchNorm
β β βββ activations.py # ReLU, Sigmoid, Softmax, etc.
β β βββ losses.py # CrossEntropy, MSE
β βββ learning/ # Learning engines
β βββ traditional.py # SGD, Adam optimizers
β βββ error_memory.py # Error Memory Bank β NOVEL
β βββ mistake_analyzer.py # Pattern discovery β NOVEL
β βββ perturbation.py # Weight exploration β NOVEL
β βββ trial_learner.py # Orchestrator β NOVEL
β
βββ mac_llm_trialnet/ # Apple Silicon LLM pipeline
β βββ 1_mac_finetune.py # Hybrid LoRA fine-tuning
β βββ 2_mac_chatbot.py # Chat + Auto-judge
β βββ 3_mac_self_correct.py # Self-correction loop
β βββ evaluate_mac.py # Regression eval baseline
β βββ run_self_correction.sh # Full pipeline runner
β βββ requirements_mac.txt # MLX + ChromaDB deps
β βββ memory/
β βββ chroma_bank.py # ChromaDB mistake banking
β βββ judge.py # LLM-as-Judge scorer
β
βββ colab_llm_trialnet/ # Google Colab pipeline
β βββ 1_base_finetune.py
β βββ 2_colab_chatbot.py
β βββ 3_self_correction_loop.py
β
βββ dashboard/ # Real-time training dashboard
β βββ server.py # Flask API
β βββ index.html
β βββ style.css
β βββ app.js
β
βββ PROGRESS.md # Build progress log
βββ ROADMAP.md # Planned features
βββ LICENSE # MIT
βββ README.md # This file
Standard neural network β forward pass, loss, backpropagation, gradient descent.
Only the novel Try-and-Learn system β no gradient descent at all:
- Error Memory Bank: Stores mistakes with priority scoring (high-confidence wrong answers get highest priority)
- Perturbation Explorer: Random weight experiments β keep what helps, revert what hurts
- Targeted Replay: Spends more training time on the hardest, most-repeated mistakes
Combines both. Traditional gradients provide the base learning signal; Try-and-Learn provides targeted corrections for stubborn mistakes.
You (user) βββΊ Chat βββΊ LLM Response
β
LLM Judge scores (0β10)
β
βββββββββββ΄βββββββββββ
score β€ 5 (bad) score β₯ 8 (good)
β
Auto-logged to ChromaDB
β
/correct [fix] β you provide the right answer
β
(10 mistakes collected)
β
bash run_self_correction.sh
β
New LoRA adapter created
β
Model permanently updated β
| Component | Technology |
|---|---|
| Core neural network | Pure NumPy (no PyTorch/TensorFlow) |
| LLM backbone | Qwen/Qwen2.5-1.5B-Instruct |
| Apple Silicon inference | MLX + mlx-lm |
| Mistake memory | ChromaDB (vector database) |
| Auto-judge | LLM-as-Judge (same local model) |
| Dashboard | Flask + Chart.js |
This project uses Graphify for AI-assisted code understanding.
To generate the knowledge graph locally:
pip install graphify
graphify .Then open graphify-out/graph.html to explore the codebase visually.
See ROADMAP.md for planned features including DPO training, GRPO alignment, and multimodal capabilities.
Contributions, issues, and feature requests are welcome! Feel free to:
- Fork the repository
- Create a branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
This project is licensed under the MIT License β see LICENSE for details.
Suraj Verma Β· @Suraj-v23
Building AI that learns the way humans do β by remembering and correcting its mistakes.