RefineX is a sophisticated iterative self-correction system designed for mathematical and logical reasoning problems. The system specializes in improving simple LLMs through collaborative enhancement techniques rather than relying on powerful models. This approach makes it particularly valuable for scenarios where computational resources are limited or where transparency and explainability are crucial.
Unlike traditional approaches that depend on large, powerful models, RefineX focuses on:
- Multi-model collaboration: Leveraging multiple smaller models working together
- Iterative refinement: Self-correction through formal feedback loops
- Formal logic integration: Real formal logic verification and reasoning guidance
- Performance optimization: High-performance caching and parallel processing
- No mock components: All implementations are production-ready and fully functional
- Model Integration: Supports multiple LLM models (Llama 3.2 1B/3B/8B, Mistral 7B)
- Collaborative Solving: Models work together to verify and improve solutions
- Performance Comparison: Automatic benchmarking across different model sizes
- Adaptive Strategies: Four distinct enhancement strategies for different problem types
Problem Input → Initial Solution → Error Detection → Formal Feedback → Refinement → Verification
↑ ↓
└────────────────── Convergence Check ←──────────────────────────────────────────────┘
- Formal Logic Verifier: Translates natural language to formal logic and applies inference rules
- Arithmetic Verifier: Validates mathematical calculations and operations
- Symbolic Verifier: Handles algebraic manipulations and symbolic reasoning
- Pattern Detection: Identifies common reasoning fallacies and error patterns
- Smart Caching: LLM response caching with TTL and disk persistence
- Parallel Processing: Concurrent verification across multiple verifiers
- Circuit Breaker: Prevents infinite loops and repeated error patterns
- Performance Optimization: Intelligent batching and resource management
RefineX System
├── Multi-Model Enhancer
│ ├── Model Profiles (1B/3B/8B/7B configurations)
│ ├── Enhancement Strategies
│ │ ├── Iterative Refinement
│ │ ├── Cross-Model Verification
│ │ ├── Reasoning Scaffolding
│ │ └── Error Pattern Learning
│ └── Best Solution Tracking
├── Pipeline Runner
│ ├── Problem Detection & Analysis
│ ├── Hypothesis Extraction
│ ├── Evidence Aggregation
│ ├── Revision Control
│ └── Convergence Management
├── Verification System
│ ├── Formal Logic Verifier
│ ├── Arithmetic Verifier
│ ├── General Verifier
│ └── Fallback Verification
├── Performance Layer
│ ├── LLM Response Cache
│ ├── Verification Cache
│ ├── Parallel Processing
│ └── Circuit Breaker
└── Integration Layer
├── Ollama Integration
├── Multi-LLM Coordination
└── API Endpoints
- Input Processing: Problems are validated and sanitized
- Initial Generation: Simple model generates baseline solution
- Problem Detection: Multiple detection methods identify potential issues
- Hypothesis Extraction: Uncertainty markers and error patterns extracted
- Verification: Formal logic and arithmetic verifiers check hypotheses
- Evidence Aggregation: Scores combined using weighted algorithms
- Revision Generation: Targeted prompts created for improvement
- Enhancement: Uncertainty-aware model generates refined solution
- Convergence Check: System determines if further iteration needed
- Result Output: Best solution returned with full iteration history
- Python 3.8+ with required packages:
pip install -r requirements.txt- Ollama installed and running:
# Install Ollama
curl -fsSL https://ollama.ai/install.sh | sh
# Pull required models
ollama pull llama3.2:latest
ollama pull llama3.2:1b # Optional for multi-model comparison
ollama pull llama3.2:3b # Optional for multi-model comparison
ollama pull llama3.2:8b # Optional for multi-model comparison
ollama pull mistral:7b # Optional for multi-model comparison
# Start Ollama service
ollama servegit clone https://github.com/artificialvirus/RefineX.git
cd RefineX
pip install -r requirements.txtpython -m src.main --mode quick-testpython -m src.main --mode solve --problem "Find the sum of the first 10 positive integers" --domain arithmeticpython -m src.main --mode comparepython -m src.scripts.multi_model_comparisonimport asyncio
from src.main import RefineXSystem
async def solve_problem():
system = RefineXSystem()
result = await system.solve_olympiad_problem(
problem_text="How many positive integers less than 1000 are multiples of 7 but not multiples of 14?",
domain="arithmetic",
use_enhancement=True
)
print(f"Solution: {result['solution']}")
print(f"Confidence: {result['confidence']}")
print(f"Enhancement used {result['enhancement_data']['iterations_completed']} iterations")
asyncio.run(solve_problem())Purpose: Improves solutions through formal feedback loops
Process:
- Generate initial solution with simple model
- Detect reasoning issues and uncertainty markers
- Generate formal feedback without giving answers
- Create enhancement prompts with positive framing
- Refine using uncertainty-aware generation
- Track best solution across iterations
Best For: General problems, confidence improvement, reasoning depth
Purpose: Uses multiple models to verify and improve each other's work
Process:
- Primary model generates solution
- Secondary models review and identify issues
- Feedback combined for targeted improvement
- Primary model generates improved solution
Best For: Complex reasoning, logical validation, error detection
Purpose: Breaks complex problems into manageable sub-problems
Process:
- Analyze problem domain and complexity
- Generate domain-specific scaffolding prompts
- Solve each component separately
- Synthesize components into final solution
Best For: Complex multi-step problems, domain-specific reasoning
Purpose: Identifies and corrects model-specific error patterns
Process:
- Generate initial solution
- Analyze against known error patterns for model type
- Create targeted improvement guidance
- Generate corrected solution
Best For: Model-specific weaknesses, known error types
Edit src/pipeline/multi_model_enhancer.py to customize model profiles:
{
"llama3.2:1b": ModelProfile(
name="Llama-3.2-1B",
endpoint="http://localhost:11434",
model_name="llama3.2:latest",
complexity=ModelComplexity.SIMPLE,
strengths=["fast_response", "basic_arithmetic"],
weaknesses=["complex_reasoning", "multi_step_logic"]
)
}Configure caching and performance in src/config.py:
@dataclass
class CachingConfig:
enabled: bool = True
llm_cache_size: int = 1000
llm_cache_ttl: int = 3600
verification_cache_enabled: bool = True
verification_cache_max_entries: int = 10000Adjust enhancement behavior:
@dataclass
class AggregatorConfig:
revision_threshold: float = 0.6 # Minimum score for revision
max_iterations: int = 3 # Maximum enhancement iterations
min_utility_improvement: float = 0.05 # Minimum improvement requiredLLM Response Cache:
- In-memory caching with configurable TTL
- Disk persistence for long-term storage
- Compression for space efficiency
- Cache analytics and hit rate monitoring
Verification Cache:
- SQLite-based persistent storage
- Hypothesis-verifier result caching
- Automatic cleanup of expired entries
# Get cache statistics
pipeline = RefineXPipeline(config, generator)
stats = pipeline.get_cache_stats()
print(f"LLM Cache hit rate: {stats['llm_cache']['hit_rate']}")
print(f"Verification Cache size: {stats['verification_cache']['size']}")- Concurrent hypothesis verification
- Parallel model comparison
- Asynchronous problem solving
- Circuit breaker for infinite loop prevention
- Fallacy Detection: Identifies common logical fallacies
- Consistency Checking: Detects contradictory statements
- Reasoning Validation: Verifies logical inference patterns
- Known Traps: Detects common arithmetic mistakes (bat-ball problem, etc.)
- Calculation Verification: Validates mathematical operations
- Pattern Recognition: Identifies suspicious numerical patterns
- Contradiction Analysis: Finds conflicting statements
- Completeness Assessment: Evaluates reasoning thoroughness
- Uncertainty Extraction: Identifies confidence markers
- Confidence Scoring: Analyzes model confidence levels
- Quality Metrics: Evaluates response characteristics
- Length Analysis: Considers response appropriateness
The system includes comprehensive testing through multi_model_comparison.py:
python -m src.scripts.multi_model_comparisonOutput Example:
🚀 Starting Multi-Model Performance Comparison
📋 Test Problems: 5
🤖 Available Models: 4
- Llama-3.2-1B (llama3.2:latest)
- Llama-3.2-3B (llama3.2:latest)
- Llama-3.2-8B (llama3.2:latest)
- Mistral-7B (llama3.2:latest)
============================================================
Testing Model: Llama-3.2-1B
============================================================
Problem 1/5: Find the sum of the first 10 positive integers...
🔍 Testing baseline...
✅ Baseline: 0.700 confidence (10.0s)
🚀 Testing enhancement...
✅ Enhanced: 0.750 confidence (12.0s)
📈 Improvement: +0.050
# Run comprehensive validation
system = RefineXSystem()
validation_report = await system.run_comprehensive_validation()# Test caching efficiency
pipeline = RefineXPipeline(config, generator)
# ... run problems ...
cache_stats = pipeline.get_cache_stats()problem = "Prove that for any positive integer n, the expression n^4 + 4n^3 + 6n^2 + 4n + 1 can be factored."
result = await system.solve_olympiad_problem(problem, domain="algebra")problem = "In a group of 100 people, 40 speak French, 50 speak German, and 20 speak both. How many speak neither?"
result = await system.solve_olympiad_problem(problem, domain="logic")problem = "Find all positive integers n such that n^2 + 19n + 88 is a perfect square."
result = await system.solve_olympiad_problem(problem, domain="number_theory")problem = "What is the area of a triangle with vertices at (0,0), (3,4), and (6,0)?"
result = await system.solve_olympiad_problem(problem, domain="geometry")To add new models:
- Update Model Profiles:
"custom_model": ModelProfile(
name="Custom-Model",
endpoint="http://localhost:11434",
model_name="custom:latest",
complexity=ModelComplexity.MODERATE,
strengths=["domain_specific"],
weaknesses=["general_reasoning"]
)- Configure Enhancement:
# In enhancement strategies, add model-specific logic
if profile.name == "Custom-Model":
# Custom enhancement logic
passCreate domain-specific verifiers:
class GeometryVerifier(BaseVerifier):
def can_verify(self, hypothesis: ErrorHypothesis) -> bool:
return hypothesis.hypothesis_type == HypothesisType.GEOMETRIC
def verify(self, hypothesis: ErrorHypothesis) -> VerifierResult:
# Geometry-specific verification logic
passAdd new strategies to multi_model_enhancer.py:
async def _enhance_through_custom_strategy(self, problem: Problem, primary_model: str):
# Custom enhancement implementation
pass
# Register in __init__
self.enhancement_strategies["custom_strategy"] = self._enhance_through_custom_strategy-
Model Selection:
- Use simple models (1B-3B) for initial solving
- Reserve larger models for verification and refinement
- Match model complexity to problem difficulty
-
Caching Strategy:
- Enable both LLM and verification caching
- Set appropriate TTL values based on use case
- Monitor cache hit rates and adjust sizes
-
Enhancement Configuration:
- Lower revision thresholds for higher quality
- Limit max iterations to prevent infinite loops
- Use adaptive thresholds for different problem types
-
Parallel Processing:
- Enable parallel verification for faster processing
- Use appropriate worker counts based on system resources
- Monitor circuit breaker triggers
# Enable detailed logging
import logging
logging.basicConfig(level=logging.DEBUG)
# Monitor performance
from src.monitoring.metrics import PerformanceMonitor
monitor = PerformanceMonitor()
# ... use monitor throughout pipelineWe welcome contributions! Please see our Contributing Guidelines for details.
git clone https://github.com/artificialvirus/RefineX.git
cd RefineX
pip install -e .
pip install -r requirements-dev.txt# Unit tests
python -m pytest tests/
# Integration tests
python -m pytest tests/integration/
# Performance tests
python -m pytest tests/performance/-
Ollama Connection Failures:
# Check if Ollama is running curl http://localhost:11434/api/tags # Restart Ollama if needed ollama serve
-
Model Not Found Errors:
# Pull required models ollama pull llama3.2:latest -
Performance Issues:
- Enable caching in configuration
- Reduce max_iterations for faster processing
- Check system resources and adjust parallel workers
-
Memory Issues:
- Reduce cache sizes in configuration
- Use smaller models for initial testing
- Clear caches periodically
# Enable comprehensive debugging
config = RefineXConfig()
config.log_level = "DEBUG"
config.enable_detailed_logging = TrueRefineX is based on research in:
- Iterative Self-Correction: Methods for LLM self-improvement
- Multi-Agent Collaboration: Coordinated problem-solving approaches
- Formal Logic Integration: Automated reasoning and verification
- Mathematical Reasoning: Specialized techniques for mathematical problem-solving
- "Training Verifiers to Solve Math Word Problems" (Cobbe et al.)
- "Self-Correction via Reinforcement Learning" (Welleck et al.)
- "Formal Mathematics Statement Curriculum Learning" (Polu et al.)
This project is licensed under the MIT License - see the LICENSE file for details.
RefineX: Advancing mathematical reasoning through collaborative intelligence 🧠✨
RefineX implements a sophisticated system that:
- Generates initial solutions with uncertainty markers using LLMs
- Extracts error hypotheses from reasoning chains using NLP
- Verifies hypotheses using specialized modules (arithmetic, logical, symbolic)
- Critiques steps using LLM-based evaluation (Phase B)
- Aggregates evidence with adaptive thresholds and utility functions (Phase B)
- Revises only flagged steps through localized prompting
- Checks consistency and triggers rollbacks when needed (Phase B)
- Iterates until convergence using advanced criteria (Phase B)
RefineX Phase C introduces enterprise-grade performance enhancements:
- ⚡ Parallel Verifier Execution: Concurrent verification for 60-70% speed improvement
- 💾 Intelligent Caching System: Multi-level caching for LLM responses and verification results
- 📊 Performance Monitoring: Real-time cache statistics and execution metrics
- 🔧 Configuration-Driven: Easy enable/disable of optimizations
- 🏭 Production-Ready: Thread-safe, fault-tolerant, and scalable architecture
RefineX Phase B introduces advanced robustification features:
- 🤖 Critic LLM Integration: Step-by-step evaluation by specialized critic models
- 🧠 Enhanced Logical Verifier: Advanced fallacy detection and quantifier logic
- 📈 Adaptive Convergence: Dynamic thresholds and utility-based stopping criteria
- 🔄 Rollback System: Automatic rollback to stable states on consistency violations
- ✅ Consistency Checking: Detection of contradictions and circular reasoning
- ⚙️ Advanced Aggregation: Multiple utility functions and false positive estimation
REFINEX/
├── src/
│ ├── models.py # Core data structures and types
│ ├── config.py # Configuration management with Phase B settings
│ ├── generator/ # LLM prompt wrappers & CoT generation
│ │ ├── llm_wrapper.py # Generator implementations (Ollama)
│ │ └── cot_parser.py # Chain-of-thought parsing
│ ├── analyser/ # spaCy pipeline & regex-based error extraction
│ │ └── extractor.py # Hypothesis extraction engine
│ ├── critic/ # Phase B: LLM-based step evaluation
│ │ ├── llm_critic.py # Critic implementations (Ollama)
│ │ └── evaluator.py # Step and batch evaluators
│ ├── verifiers/ # Verification modules
│ │ ├── base.py # Base verifier interface
│ │ ├── arithmetic.py # Mathematical verification
│ │ ├── logical.py # Enhanced logical reasoning (Phase B)
│ │ └── symbolic.py # Symbolic math verification (Phase B)
│ └── pipeline/ # Orchestration and control
│ ├── aggregator.py # Enhanced evidence aggregation (Phase B)
│ ├── controller.py # Revision control & prompt generation
│ ├── consistency.py # Consistency checking & rollback (Phase B)
│ └── runner.py # Main pipeline with Phase B integration
├── performance/ # Phase C: Performance optimization
│ ├── parallel.py # Parallel verifier execution
│ └── caching.py # Intelligent caching system
├── data/ # Sample problems & test cases
├── tests/ # Comprehensive test suite (includes Phase B tests)
├── examples/ # Usage examples and demos
├── config/ # Configuration files
├── cli.py # Enhanced CLI with Phase B options
└── requirements.txt # Dependencies
# Clone repository
git clone <repository-url>
cd RefineX
# Create virtual environment
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
# Install spaCy language model
python -m spacy download en_core_web_sm# Run with a problem file
python cli.py --problem examples/arithmetic_problem.json
# Interactive mode
python cli.py --interactive
# Use custom configuration
python cli.py --problem examples/logic_problem.json --config config/default.yaml
# Enable Phase B features
python cli.py --problem examples/problem.json --enable-critic --adaptive-threshold
# Use real critic for evaluation
python cli.py --problem examples/problem.json --enable-critic --critic-model mistral
# Customize convergence behavior
python cli.py --problem examples/problem.json --utility-function weighted_improvement --max-iterations 5
# Disable advanced features
python cli.py --problem examples/problem.json --disable-rollback --disable-consistency-check
# Verbose output
python cli.py --problem examples/arithmetic_problem.json --verbosefrom src import RefineXConfig, RefineXPipeline, Problem
from src.generator import create_default_ollama_generator
# Load configuration
config = RefineXConfig.default()
# Create generator (Real Ollama for production usage)
generator = create_default_ollama_generator()
# Initialize pipeline
pipeline = RefineXPipeline(config, generator)
# Create problem
problem = Problem(
id="demo",
text="What is 25 + 17?",
domain="arithmetic"
)
# Run iterative correction
result = pipeline.run(problem)
print(f"Final answer: {result.final_answer}")
print(f"Iterations: {result.total_iterations}")
print(f"Corrections: {result.correction_count}")# Run Phase A demonstration
python examples/demo.py
# Run Phase B demonstration with advanced features
python examples/demo_phase_b.py
# Run evaluation with baselines
python -m src.scripts.evaluate --datasets refinex_test_suite --baseline iterative --sample-size 5
python -m src.scripts.evaluate --datasets refinex_test_suite --baseline zero --models llama3.2:latest
# Run an ablation sweep across baseline modes with RC plots and CSV summaries
python -m src.scripts.ablation_sweep -d refinex_test_suite -b iterative -b -cot -b -critic -b -symbolic -b -kb -b -uncertainty --bootstrap-cis --paired-stats --output-dir evaluation_results/ablation_sweepRefineX uses YAML configuration files. Key settings:
aggregator:
revision_threshold: 0.6 # Minimum score to trigger revision
max_iterations: 3 # Maximum correction cycles
min_utility_improvement: 0.05 # Convergence threshold
verifier:
enable_arithmetic: true # Enable arithmetic verification
enable_logical: true # Enable logical verification
enable_symbolic: true # Enable symbolic verification (Phase B)
arithmetic_timeout: 5.0 # Verification timeout (seconds)
generator:
model_name: "mistral" # LLM model name
temperature: 0.7 # Generation temperature
max_tokens: 1024 # Maximum tokens per generation# Critic LLM settings
critic:
enabled: true # Enable critic evaluation
model_name: "mistral" # Critic model name
per_step_evaluation: true # Evaluate each step individually
confidence_threshold: 0.7 # Minimum confidence for critic flagging
# Advanced aggregation
aggregator:
adaptive_threshold: true # Enable adaptive revision threshold
utility_function: "weighted_improvement" # Utility calculation method
early_stopping_patience: 2 # Early stopping patience
threshold_decay: 0.1 # Threshold decay rate
# Robustification features
enable_rollback: true # Enable rollback to stable states
enable_consistency_check: true # Enable consistency checking# Caching system (Phase C)
caching:
enabled: true # Enable caching system
llm_cache_enabled: true # Cache LLM responses
llm_cache_size: 1000 # Maximum cached responses
llm_cache_ttl: 3600 # Cache TTL in seconds (1 hour)
verification_cache_enabled: true # Cache verification results
verification_cache_max_entries: 10000 # Maximum cached verifications
# Parallel verification (Phase C)
verifier:
parallel_verification: true # Enable parallel verifier execution
parallel_workers: 3 # Number of concurrent workers
parallel_timeout: 10.0 # Global timeout for parallel execution
result_aggregation: "conservative" # Result aggregation strategy# Run all tests
pytest
# Run specific test modules
pytest tests/test_pipeline.py
pytest tests/test_verifiers.py
pytest tests/test_analyser.py
pytest tests/test_phase_b.py # Phase B specific tests
# Run with coverage
pytest --cov=src
# Run linting
flake8 src tests
black --check src tests
isort --check-only src testsProblem: "What is 25 + 17?"
Initial Generation:
Step 1: I need to add 25 and 17.
Step 2: I think this might be wrong: 25 + 17 = 41.
Step 3: Therefore, the answer is 41.
After Correction:
Step 1: I need to add 25 and 17.
Step 2: 25 + 17 = 42.
Step 3: Therefore, the answer is 42.
Pipeline Summary:
- ✅ Success: True
- 🔄 Iterations: 2
- 🛠️ Corrections: 1
- 🎯 Converged: True
- ⏱️ Time: 0.15s
- Uncertainty Detection: Regex-based pattern matching for uncertainty markers
- Multi-Modal Verification: Arithmetic, logical, and symbolic verifiers
- Evidence Aggregation: Weighted scoring from multiple signals
- Localized Revision: Targeted correction of only flagged steps
- Comprehensive Logging: Detailed provenance tracking and metrics
-
🤖 Critic LLM Integration:
- Step-by-step evaluation using specialized LLM critics
- Ollama critic backend for real evaluation
- Configurable confidence thresholds and evaluation modes
-
🧠 Enhanced Logical Verifier:
- Advanced logical fallacy detection (affirming consequent, denying antecedent, etc.)
- Quantifier logic verification (universal, existential, negative)
- Complex inference pattern checking (hypothetical syllogism, disjunctive syllogism)
-
📈 Adaptive Convergence:
- Dynamic revision thresholds that adapt over iterations
- Multiple utility functions (simple, weighted improvement, ML-based)
- Early stopping with patience-based criteria
- False positive rate estimation and management
-
🔄 Rollback System:
- Automatic detection of consistency violations and utility degradation
- Intelligent rollback point selection based on utility and recency
- Seamless state restoration with minimal information loss
-
✅ Consistency Checking:
- Detection of contradictory statements within reasoning chains
- Circular reasoning and undefined reference detection
- Numerical consistency verification across steps
- Step preservation validation between iterations
-
⚙️ Advanced Aggregation:
- Configurable utility functions for convergence assessment
- Weighted scoring considering verifier and critic confidence
- Trend analysis for utility improvement detection
- Adaptive threshold management based on iteration progress
- ⚡ Parallel Verifier Execution:
- Concurrent execution of arithmetic, logical, and symbolic verifiers
- ThreadPoolExecutor-based parallel processing with configurable workers
- Intelligent result aggregation (conservative, optimistic, weighted strategies)
- Automatic fallback to sequential execution on failures
You can run RefineX in different baseline modes via the evaluation harness:
- zero: Zero-pass (no refinement)
- prompt: Prompt-enhanced single pass
- critic: Critic-augmented pass using the pipeline’s critic integration
- iterative: Full RefineX pipeline (default)
Programmatic flag: set EvaluationConfig.baseline to one of the above. Reports include a “Baseline Mode” banner per model.
-
💾 Intelligent Caching System:
- Multi-level caching for LLM responses and verification results
- In-memory LRU cache with TTL for LLM responses
- SQLite-based persistent cache for verification results
- Thread-safe operations with comprehensive statistics
-
📊 Performance Monitoring:
- Real-time cache hit/miss statistics and performance metrics
- Execution time tracking and speedup measurements
- Memory usage monitoring and automatic cleanup
- Configurable analytics and reporting intervals
-
🔧 Production-Ready Features:
- Configuration-driven optimization enable/disable
- Graceful degradation when optimizations fail
- Comprehensive error handling and logging
- Cross-session persistence and cache management
This project is licensed under the MIT License - see the LICENSE file for details.
RefineX - Bringing systematic self-correction to LLM reasoning through uncertainty, verification, and targeted revision.