Skip to content

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

0 watching

Forks

Repository files navigation

ReviewLens AI

Hybrid code review with explicit evidence and attribution.

ReviewLens is a hybrid code-review platform combining deterministic static analysis with context-aware LLM review using Amazon Bedrock.

Java Spring Boot React PostgreSQL Amazon AWS Docker

Live Demo · API Health · Report a Bug


Why ReviewLens?

Traditional linters are fast but often lack context. General-purpose AI can provide context, but its output is only as useful as the evidence it receives. ReviewLens brings both approaches together in one review pipeline:

  • Find issues deterministically with extensible, rule-based static analysis.
  • Review source independently with Bedrock reasoning over selected source regions, imports, dependency/config context, and optional static signals—even when static analysis finds nothing.
  • Validate AI findings against supplied file paths, line ranges, and exact source evidence before returning them.
  • Attribute each finding as Static Analysis, AI Review, or Static + AI.
  • Track work asynchronously while large repositories are cloned and analyzed in the background.
  • Persist results in PostgreSQL and attempt to archive the complete JSON report in Amazon S3. Archive failure does not discard database results.
  • Keep each review private to its creator through GitHub OAuth 2.0 session ownership.

Live Demo

Try the hosted application at reviewlens-frontend-avas.onrender.com.

  1. Sign in with GitHub.
  2. Choose a repository from your account.
  3. Select Review Repository.
  4. Follow the live status until the report is ready.

The application is hosted on Render. A cold service may take a short time to wake up on the first request.

How It Works

flowchart LR
    U["Developer"] -->|"GitHub OAuth"| UI["React client"]
    UI -->|"Create review"| API["Spring Boot API"]
    API --> DB[("PostgreSQL")]
    API --> W["Async analysis worker"]
    W -->|"Clone"| GH["GitHub repository"]
    W --> SA["Static analysis engine"]
    W --> SRC["Selected source files"]
    SA --> CTX["Bounded context builder"]
    SRC --> CTX
    CTX --> AI["Bedrock source review + bounded cross-file pass"]
    AI --> V["Strict JSON and source-evidence validation"]
    V --> M["Conservative deduplication + attribution"]
    SA --> M
    M --> DB
    M --> S3["Amazon S3 report (best effort)"]
    DB -->|"Status & findings"| API
    S3 -->|"Full JSON report"| API
Loading

Each review moves through a clear lifecycle:

Queued → Cloning → Analyzing → Completed
                              ↘ Failed

The static engine scans Java, JavaScript/JSX, TypeScript/TSX, and Python files for TODO/FIXME markers, common debug output, likely hardcoded secrets, and files over 500 lines. These are deterministic patterns, not LLM findings.

Bedrock separately reviews selected source code for correctness, security, concurrency, resource management, error handling, API usage, data consistency, and performance issues. A lightweight second pass can compare source regions from different files. AI findings are validated against source code before being returned. Validation establishes schema and source grounding; it cannot prove that the model's reasoning is correct.

Review context and limits

  • The scanner supports the extensions listed above. This does not imply compiler-level analysis or measured LLM accuracy for each language.
  • The current input is a repository snapshot, not a pull-request diff. File selection favors application/controller/service/auth/worker paths and files with static signals, while reducing test/fixture priority. PR/change-aware selection is not implemented.
  • Up to 12 candidate files are read, with at most 8 chunks selected in file rounds so a long file does not consume the entire budget. Files over 256 KB, external symlinks, build output, and dependencies are excluded from AI context.
  • Each chunk retains repository, exact relative path, language, original inclusive line range, source, and optional static signals. Chunks use up to 120 lines / 6,000 characters; declaration/block boundaries are preferred. Oversized functions use whole-line windows with 12-line overlap. Oversized individual lines are skipped, never sliced into artificial code.
  • Repository context includes a bounded inventory/import list and small root package.json, pom.xml, pyproject.toml, or tsconfig.json files (only complete files up to 1,200 bytes). Its total limit is 4,000 characters.
  • The optional cross-file pass uses up to 3 distinct already-selected source regions / 12,000 source characters plus up to 6 finding references. It runs only after at least one successful file pass. This is a limited opportunity for cross-layer reasoning, not full dependency-graph analysis.
  • System and serialized user prompts together are capped at 24,000 characters per request; oversized requests are skipped. There are at most 8 file requests plus 1 cross-file request, each with a 3,000 output-token cap. Character limits are not exact token/cost estimates.
  • Common embedded credential patterns are redacted before prompting; .env and local AWS configuration are never loaded as review context. Repository text is treated as untrusted data. Do not assume that regex redaction detects every possible secret format.

Findings, validation, and attribution

Bedrock must return only {"findings":[...]} with source, category, severity, confidence, filePath, startLine, endLine, title, explanation, evidence, and suggestion on each finding. Invalid JSON, unknown/missing fields, nonexistent or unsupplied paths, invalid ranges, nonmatching evidence, and confidence below 0.8 are rejected. Paths and line numbers are never guessed or repaired. Evidence must match the cited lines in a chunk supplied to that exact model call; redaction placeholders cannot serve as evidence. Confidence is the model's self-assessment, not a calibrated probability.

GET /reviews/{id}/findings and new S3 reports include:

Result source detectedBy UI label
Deterministic rule STATIC ["STATIC"] Static Analysis
Validated Bedrock finding AI ["AI"] AI Review
Same issue found by both AI (retained richer finding) ["STATIC", "AI"] Static + AI

Deduplication is intentionally conservative: overlapping hardcoded-secret signals can merge with equivalent AI credential findings. Repeated AI findings need overlapping locations, identical evidence, the same category, and substantially similar titles. Location alone never merges unrelated issues; some paraphrased duplicates can remain. Static findings need not have AI confidence/evidence. lineNumber and message remain as backward-compatible aliases for startLine and explanation.

The AI metadata endpoint includes status and coverage (eligible/selected files, selected/completed chunks, cross-file pass, context limits, and discarded finding counts). Its summary is generated from these counters, not unvalidated model prose. Status is COMPLETED, PARTIAL, UNAVAILABLE, or NO_CONTEXT. COMPLETED means the selected calls completed, not that every repository file was reviewed. A review can complete with static results when AI is unavailable. No findings is not proof of safe code.

Flyway V6 backfills all historical findings as STATIC and labels historical AI narratives LEGACY_UNVALIDATED. Existing archived S3 objects are not rewritten; old raw reports remain legacy narratives. All result endpoints retain the existing owner check before database or S3 access.

Feature Highlights

Capability What it provides
Hybrid code review Deterministic signals plus independent, source-grounded LLM findings
GitHub integration OAuth login and repository discovery; reviews currently support public repositories only
Background processing Non-blocking analysis with observable review status
Structured findings Source attribution, location, severity, explanation, and AI evidence/confidence
AI review Selected-source reasoning, bounded cross-file review, strict validation, and coverage reporting
Durable reports Findings in PostgreSQL; complete attributed JSON reports archived to S3 when available
Production foundation Flyway migrations, health endpoints, Docker, and cloud deployment

Tech Stack

Layer Technology
Frontend React 19, Vite 8, JavaScript, CSS
Backend Java 21, Spring Boot 4.1, Spring MVC
Security Spring Security, GitHub OAuth 2.0
Data PostgreSQL, Spring Data JPA, Hibernate, Flyway
AI Amazon Bedrock Runtime
Storage Amazon S3
Infrastructure Docker, AWS RDS, Render
Build & quality Maven Wrapper, npm, oxlint

Project Structure

ReviewLens-AI/
├── frontend/                         # React web client
│   ├── public/
│   └── src/
│       ├── api/                      # Backend API clients
│       ├── components/               # Repository and review UI
│       └── pages/
├── src/main/java/com/reviewlens/
│   ├── config/                       # Security, AWS, and Bedrock clients
│   ├── controller/                   # REST endpoints
│   ├── dispatcher/                   # Async review dispatch
│   ├── entity/                       # JPA domain model
│   ├── repository/                   # Data access
│   ├── service/                      # GitHub, analysis, AI, and S3 logic
│   └── worker/                       # Background analysis pipeline
├── src/main/resources/db/migration/ # Versioned Flyway migrations
├── docker-compose.yml
├── Dockerfile
└── pom.xml

Getting Started

Prerequisites

  • Java 21+
  • Node.js 20+ and npm
  • PostgreSQL 16+
  • A GitHub OAuth App
  • Existing AWS credential-provider configuration for live Bedrock/S3 use; not required for automated tests

1. Clone the repository

git clone https://github.com/starstarrr/ReviewLens-AI.git
cd ReviewLens-AI

2. Start PostgreSQL

The included Compose configuration starts PostgreSQL and the backend together. For frontend development, you can start only the database:

docker compose up -d postgres

This creates a local database on localhost:5432 with the database, username, and password all set to reviewlens.

3. Configure the backend

Export the required environment variables before starting Spring Boot:

export DB_URL="jdbc:postgresql://localhost:5432/reviewlens"
export DB_USERNAME="reviewlens"
export DB_PASSWORD="reviewlens"

export GITHUB_CLIENT_ID="your-github-client-id"
export GITHUB_CLIENT_SECRET="your-github-client-secret"
export GITHUB_OAUTH_CALLBACK_URL="http://localhost:8080/login/oauth2/code/github"

export AWS_REGION="us-west-2"
export AWS_S3_BUCKET_NAME="your-report-bucket"
export BEDROCK_MODEL_ID="your-bedrock-model-id"

export FRONTEND_URL="http://localhost:5173"

AWS clients use the existing SDK DefaultCredentialsProvider chain. Use your existing authorized local profile or environment/role configuration; do not put AWS keys in this repository. No AWS Console setup is required to run mocked tests. Set AI_REVIEW_ENABLED=false for explicit offline AI mode, which returns UNAVAILABLE and never fabricates findings.

For local authentication, configure your GitHub OAuth callback URL as:

http://localhost:8080/login/oauth2/code/github

Private repositories are listed but cannot currently be reviewed because the isolated clone step does not receive GitHub credentials.

4. Run the backend

./mvnw spring-boot:run

The API is available at http://localhost:8080; Flyway applies the database schema automatically on startup.

5. Run the frontend

cd frontend
npm install
npm run dev

Open http://localhost:5173.

Configuration Reference

Variable Required Description Example
DB_URL Yes PostgreSQL JDBC URL jdbc:postgresql://localhost:5432/reviewlens
DB_USERNAME Yes Database user reviewlens
DB_PASSWORD Yes Database password reviewlens
GITHUB_CLIENT_ID Yes GitHub OAuth client ID —
GITHUB_CLIENT_SECRET Yes GitHub OAuth client secret —
GITHUB_OAUTH_CALLBACK_URL No OAuth callback URL; defaults from the request base URL http://localhost:8080/login/oauth2/code/github
AWS_REGION Yes AWS region for Bedrock and S3 us-west-2
AWS_BEDROCK_REGION No Bedrock-specific region override value of AWS_REGION
AWS_S3_REGION No S3-specific region override value of AWS_REGION
AWS_S3_BUCKET_NAME Yes Destination report bucket reviewlens-reports
BEDROCK_MODEL_ID No Existing Bedrock inference profile or model ID us.anthropic.claude-sonnet-4-6
AI_REVIEW_ENABLED No Enable Bedrock review; default true false for offline mode
AI_REPOSITORY_PASS_ENABLED No Enable the extra bounded cross-file pass; default true true
FRONTEND_URL No Allowed frontend origin and OAuth redirect target http://localhost:5173
CLONE_TIMEOUT No Maximum time allowed for git clone PT2M
CLONE_MAX_BYTES No Maximum cloned repository size in bytes 268435456
CLONE_MAX_CONCURRENT No Maximum concurrent clone processes 2

AWS authentication is resolved through the standard SDK credential chain. Keep credentials outside source/config files and logs.

API Overview

Most endpoints require an authenticated GitHub session.

Method Endpoint Description
GET /api/me Return the authenticated GitHub user
GET /github/repositories List repositories available to the user
GET /github/repos/{owner}/{repo} Fetch repository metadata
POST /reviews Start a new asynchronous review
GET /reviews/{id} Read review state and metadata
GET /reviews/{id}/findings List attributed static, AI, and merged findings
GET /reviews/{id}/summary Return severity totals
GET /reviews/{id}/ai-review Return AI status and coverage (legacy narrative fields remain for compatibility)
GET /reviews/{id}/report Download the archived report from S3, if available

Validation

./mvnw clean verify
cd frontend
npm ci
npm run lint
npm test
npm run build

All Bedrock tests use mocked SDK clients; no automated test resolves AWS credentials or calls AWS. Backend coverage includes input source/context, malformed/truncated JSON, fabricated locations/evidence, confidence bounds, provenance, conservative deduplication, partial/failure handling, report persistence, and per-user ownership checks.

Flyway tests require an explicitly supplied loopback-only PostgreSQL database. They test both an empty schema and a V5 → V6 upgrade, verify the historical attribution/owner, and exercise new constraints. They create and drop only randomly named test schemas. CI runs them using its PostgreSQL service. To include them locally:

MIGRATION_TEST_DB_URL=jdbc:postgresql://localhost:5432/reviewlens \
MIGRATION_TEST_DB_USER=reviewlens \
MIGRATION_TEST_DB_PASSWORD=reviewlens \
./mvnw clean verify

Without MIGRATION_TEST_DB_URL, the two database tests are explicitly skipped. The application datasource (including any remote database) is never used by those tests. The existing /test/bedrock diagnostic endpoint is disabled by default and is not invoked by tests.

Extending Static Analysis

Static checks implement the AnalysisRule interface and are discovered by Spring. Add a new rule under service/analysis/rules, register it as a Spring component, and return findings with a severity, file location, and description. The analysis engine will include it in subsequent reviews automatically.

Roadmap

  • Pull request and diff-aware reviews
  • Multi-language static-analysis rules
  • Repository review history dashboard
  • Queue-backed workers with Amazon SQS
  • Team workspaces and shared reports
  • Repository-aware AI chat
  • Multi-model provider support
  • CloudWatch observability and distributed tracing

Author

Built by Xingran Ma — Computer Science at the University of British Columbia, with a focus on backend engineering, cloud systems, and applied AI.


If ReviewLens helped you, consider giving the project a ⭐.

About

No description, website, or topics provided.

Resources

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages