Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
98 changes: 98 additions & 0 deletions crates/neuromesh-context/src/gold.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1107,10 +1107,108 @@ pub fn localization_order(
.take(k.max(out.len()))
.collect();
}
// The report's own pointers: files it names (a traceback's frames,
// deepest first, then paths and unique file names in prose), at double
// weight. SWE-bench dev (225): Acc@1 0.249 → 0.298, @3 0.484 → 0.516;
// both halves of dev agree (prototype `scripts/research/mention_prior.py`).
if long {
let named: Vec<String> = named_files(graph, prompt)
.into_iter()
.filter(|p| !crate::selector::is_noise_path(std::path::Path::new(p)))
.collect();
if !named.is_empty() {
const K: f32 = 60.0;
const NAMED_WEIGHT: f32 = 2.0;
let mut fused: std::collections::HashMap<String, f32> =
std::collections::HashMap::new();
for (list, w) in [(&out, 1.0), (&named, NAMED_WEIGHT)] {
for (i, p) in list.iter().enumerate() {
*fused.entry(p.clone()).or_insert(0.0) += w / (K + i as f32 + 1.0);
}
}
let mut order: Vec<(String, f32)> = fused.into_iter().collect();
order.sort_by(|a, b| b.1.total_cmp(&a.1).then_with(|| a.0.cmp(&b.0)));
out = order.into_iter().map(|(p, _)| p).collect();
}
}
out.extend(aside);
out
}

/// Indexed files a report names, best pointer first: traceback frames
/// (`File "…/pkg/mod.py", line 12`) from the deepest up, then any other path
/// or file name in order of mention. A path resolves by its longest suffix
/// that names exactly one indexed file (site-packages prefixes fall away); a
/// bare file name only when one non-test file has it.
fn named_files(graph: &neuromesh_graph::NeuralProjectGraph, prompt: &str) -> Vec<String> {
use std::sync::OnceLock;
static FRAME: OnceLock<regex::Regex> = OnceLock::new();
static PATH: OnceLock<regex::Regex> = OnceLock::new();
let frame = FRAME.get_or_init(|| regex::Regex::new(r#"File "([^"\n]+)", line \d+"#).unwrap());
let path =
PATH.get_or_init(|| regex::Regex::new(r"[\w./\\-]+\.[A-Za-z][A-Za-z0-9]{0,5}\b").unwrap());
let files: Vec<String> = graph
.file_node_paths()
.into_iter()
.map(|(_, p)| p.to_string_lossy().replace('\\', "/"))
.collect();
let mut by_name: std::collections::HashMap<&str, Vec<&String>> =
std::collections::HashMap::new();
for f in &files {
if !neuromesh_core::source_path::is_test_path(std::path::Path::new(f)) {
by_name
.entry(f.rsplit('/').next().unwrap_or(f))
.or_default()
.push(f);
}
}
let resolve = |token: &str| -> Option<String> {
let t = token.replace('\\', "/");
let t = t.trim_start_matches("./");
let parts: Vec<&str> = t.split('/').filter(|s| !s.is_empty()).collect();
if parts.is_empty() {
return None;
}
for i in 0..parts.len() {
let suffix = parts[i..].join("/");
let hits: Vec<&String> = files
.iter()
.filter(|f| **f == suffix || f.ends_with(&format!("/{suffix}")))
.collect();
match hits.len() {
1 => return Some(hits[0].clone()),
0 => continue,
_ if i + 1 < parts.len() => continue,
_ => return None,
}
}
match by_name.get(parts[parts.len() - 1]) {
Some(c) if c.len() == 1 => Some(c[0].clone()),
_ => None,
}
};
let mut out: Vec<String> = Vec::new();
let frames: Vec<&str> = frame
.captures_iter(prompt)
.filter_map(|c| c.get(1).map(|m| m.as_str()))
.collect();
for tok in frames.iter().rev() {
if let Some(f) = resolve(tok) {
if !out.contains(&f) {
out.push(f);
}
}
}
for m in path.find_iter(prompt) {
if let Some(f) = resolve(m.as_str()) {
if !out.contains(&f) {
out.push(f);
}
}
}
out
}

#[cfg(test)]
mod tests {
use super::*;
Expand Down
45 changes: 45 additions & 0 deletions docs/planning/13-roadmap-paper.fa.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# نقشه‌ی راه تا مقاله (۲۰۲۶-۱۰-۰۹)

هدف: یک سیستم localization محلی (CPU، متن‌باز) که در Acc@k با روش‌های LLM-محور رقابت کند، با
روش holdout صادقانه و جدول ablation کامل. هر عدد از `docs/research/contributions-log.md`.

## وضعیت (اندازه‌گیری‌شده)

| بنچمارک | ما | بهترین بدون LLM منتشرشده | بهترین با LLM |
|---|---|---|---|
| SWE-bench Lite (۲۷۶ holdout) Acc@1/3/5 | 0.486 / 0.710 / 0.750 (CPU، 2.8s) | CodeRankEmbed 0.526 / 0.777 / 0.847 (GPU) | LocAgent+Claude 0.777 / 0.920 / 0.942 |
| SWE-bench dev (۲۲۵) Acc@5 | 0.600 | — | — |
| پرسش زبان ساده، ۴ holdout، R@3 | 0.500 / 0.917 / 0.875 / 0.833 | BM25 0.500 / 0.750 / 0.875 / 0.583 | — |

## الهام از رقبا — چه برداشتیم، چه می‌ماند

| رقیب | ایده‌ی کلیدی | وضعیت نزد ما |
|---|---|---|
| LocAgent | BM25 روی محتوای entity (تابع/کلاس) | ✅ `chunk_rank` — +0.17 Acc@5 روی dev |
| LocAgent | پیمایش گراف با تصمیم LLM | ❌ رأی کور گراف رد شد؛ فقط با LLM (فاز D2) |
| Agentless | skeleton فشرده‌ی فایل‌ها | ✅ fold/skeleton در packet |
| Agentless | انتخاب سلسله‌مراتبی فایل → تابع → خط با LLM | ⏳ فاز D1 |
| Agentless | چند نمونه + رأی‌گیری (sampling + voting) | ⏳ فاز D3 |
| CodeRankEmbed / Jina | بازیابی چگال کد | ❌ روی CPU غیرعملی (۲.۶ chunk/s)؛ فقط روی short-list (فاز D4) |
| Continue / Cody | reranker پس از بازیابی | ❌ روی holdout تازه ضرر (cobra 0.792 → 0.708) |

## فازها (به ترتیب، با معیار پذیرش)

| فاز | کار | ورودی | معیار پذیرش | موازی‌پذیر |
|---|---|---|---|---|
| **A** ✅ | ترکیب لیست برای پرسش ساده، holdoutهای ۳ و ۴، اصلاح باگ ارزیابی | — | بدون افت در ۱۴ ست | — |
| **B** | Verified (۵۰۰) با v1.2.0، یک‌باره | `verified.json` | جدول با CI | ✅ (IO) |
| **C** | ablation روی Lite holdout: بدون بهداشت متن / بدون chunk_rank / فقط packet | ۳ باینری | هر جزء سهم مثبت جدا | ✅ |
| **D1** | مرحله‌ی LLM سبک Agentless روی ۱۵ نامزد + skeleton، مدل محلی Qwen2.5-Coder-3B (llama.cpp، CPU) | `llm_localize.py` | dev-fast: Acc@1 ≥ +0.10 | پس از دانلود مدل |
| **D2** | سبک LocAgent: LLM یک بار گسترش گراف را برای ۳ نامزد برتر می‌خواهد (callers/imports) | MCP trace | dev-fast: Acc@5 ≥ +0.05 | — |
| **D3** | چند نمونه با دمای > 0 و رأی (Agentless) | — | پایداری + Acc@1 | — |
| **D4** | embedding فقط روی short-list (۲۰ فایل) — dense بدون هزینه‌ی کل ریپو | jina-code | dev-fast: Acc@5 | — |
| **E** | holdout نهایی Lite با بهترین ترکیب، یک‌باره؛ توکن در هر issue در برابر Agentless | — | جدول اصلی مقاله | — |
| **F** | مقایسه‌ی رودررو روی همان ۲۷۴ instance مقاله‌ی LocAgent | لیست instanceها | — | — |
| **G** | holdout ≥ ۵۰ پرسش + annotator دوم | ⚠️ تیم | توافق بین annotatorها | — |
| **H** | پیش‌نویس مقاله | log | — | ✅ همیشه |

## موانعی که فقط پارسا می‌تواند باز کند

- کلید API (مثلاً Baseten که در فاز C قبلی کار کرد) برای مقایسه‌ی LLM قوی در D1/E.
- یک نفر دوم برای gold مستقل (فاز G).
36 changes: 36 additions & 0 deletions docs/research/contributions-log.md
Original file line number Diff line number Diff line change
Expand Up @@ -294,3 +294,39 @@ unchanged; packet recall/precision on the new sets: cobra 0.833/0.357, axios 0.6
affected: cobra BM25 R@3 is 0.875, not the 0.958 first printed. Re-run on every published set
(holdout-2/c/ml2/lang, ripgrep, click, this repo): identical numbers. SWE-bench Lite has no root-level
gold file (dev: 4 of 225, Verified: 1 of 500), and the Rust gold harness matches bare names exactly.

### 8.7 Head-to-head on LocAgent's exact Lite subset (274)

LocAgent §5.1 keeps the 274 Lite instances whose patch modifies an existing function. Rebuilt with
`scripts/research/locagent_subset.py` (removed line or insertion point inside a function span of
the base file, Python `ast`): **exactly 274 of 300**. Same v1.2.0 results file, scored on that set:

| method (274, same instances) | LLM | Acc@1 | Acc@3 | Acc@5 |
|---|---|---|---|---|
| BM25 (LocAgent's) | – | 0.387 | 0.518 | 0.617 |
| BM25 (ours) | – | 0.299 | 0.522 | 0.606 |
| Jina-Code-v2 | – | 0.434 | 0.712 | 0.803 |
| **engine v1.2.0 (CPU)** | – | **0.474** [0.42,0.53] | **0.708** [0.65,0.76] | **0.752** [0.70,0.80] |
| CodeRankEmbed | – | 0.526 | 0.777 | 0.847 |
| Agentless + GPT-4o | ✓ | 0.672 | 0.745 | 0.745 |
| Agentless + Claude-3.5 | ✓ | 0.726 | 0.792 | 0.796 |
| LocAgent + Qwen2.5-7B (fine-tuned) | ✓ | 0.708 | 0.847 | 0.883 |
| LocAgent + Claude-3.5 | ✓ | 0.777 | 0.920 | 0.942 |

Caveat: all 24 session-16 dev-class instances are among the 274 (measured) (flask/requests/seaborn/
xarray/pylint); without them (250 strict holdout instances) the engine scores 0.488 / 0.716 /
0.756 and BM25 0.304 / 0.512 / 0.592 — the looked-at instances do not flatter the result.

### 8.8 Files the report names (C13) and quoted error messages (session 17)

Inspired by what agents do first (open the frame of a traceback, grep the error message). On
SWE-bench dev, 50 single-file issues name the gold file in their text, yet only 28 had it first.

- **Named files** (`named_files` in `gold.rs`): traceback frames deepest first, then paths and
unique file names in prose, resolved against the index by longest unique path suffix; a third
RRF vote at weight 2 for reports. Prototype (`mention_prior.py`) on dev 225: Acc@1/3/5/10
0.249/0.484/0.600/0.680 → **0.298/0.516/0.609/0.689**; both halves agree (dev-fast 0.254 → 0.271
@1, dev-rest 0.247 → 0.307 @1; weight 1/2/4 swept, 2 kept for @3). Engine port on dev-fast
reproduces the prototype exactly (0.271/0.525/0.576/0.695). Fourteen sets unchanged.
- **Error-message grep** (`error_grep_prior.py`): fires on 15 of 225 dev issues, fixes 2 (+0.009
@1). Small and positive; not shipped.
117 changes: 117 additions & 0 deletions docs/research/paper-draft.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
# Paper draft (living) — *Where to look: LLM-free, CPU-only code localisation for coding agents, measured on holdouts*

Status: skeleton with measured numbers (session 17). Every number traces to
`contributions-log.md`; TODO marks what is not measured yet.

## Abstract (draft)

Coding agents spend most of their budget finding the code a task needs. Strong localisers today
either run an LLM agent over the repository (Agentless, LocAgent, SWE-agent) or a GPU-scale dense
retriever (CodeRankEmbed). We ask how far a local engine gets with neither: a per-project code
graph, field-weighted lexical ranking, a definition-level ranking of long reports, and explicit
report hygiene, all on a laptop CPU. On SWE-bench Lite, run once after tuning only on the SWE-bench
dev split, the engine reaches file-level Acc@1/3/5 of 0.486/0.710/0.750 on 276 held-out issues
(plain BM25 0.301/0.507/0.587) in 2.8 s per issue including a cold index — above a code embedding
model at Acc@1 and level with Agentless+GPT-4o at Acc@5. We report every component's effect,
negative results (dense retrieval on CPU, cross-encoder reranking, blind graph expansion), and a
holdout protocol that caught two results that looked-at sets had suggested. TODO: LLM stage (D1),
Verified, ablations.

## 1. Introduction

- Problem: context for coding agents; cost of reading; localisation as the first step.
- Gap: LLM-heavy or GPU-heavy localisers; local-first tools (Aider RepoMap, Continue) are not
measured on standard localisation benchmarks; tuning on the evaluation set is common.
- Contributions:
1. A holdout protocol for retrieval engines (gold locked before runs; dev vs holdout labelled per
set; one run per holdout) and evidence it matters (§6).
2. An LLM-free engine: prompt-evidence seeding, BM25F with a comment field, definition-level
ranking for reports, report hygiene, localisation list (C2–C12 in the log).
3. Results on SWE-bench Lite/dev/Verified and on four plain-language holdouts in four languages.
4. Negative results with numbers.

## 2. Related work

Agentless (hierarchical LLM localisation), LocAgent (graph + LLM agent), SWE-agent / OpenHands /
MoatlessTools (agentic search), CodeRankEmbed / Jina code embeddings (dense), BM25, Aider RepoMap
(PageRank over tags), Continue/Cody (retrieve + rerank), BM25F (Robertson et al.), hybrid fusion
(Bruch et al. 2023), RRF (Cormack et al. 2009).

## 3. System

3.1 Index: per-project graph (files, definitions, calls, imports), built in 26 s for django
(3.5k files) after the link-resolution fixes (C10).
3.2 Questions: prompt anchors → seeds → activation → packet; BM25F over path/symbols/comments/body.
3.3 Reports (≥60 words): hygiene (template scaffolding, links, checklists, headings), code-like
tokens, definition-level BM25 with title ×3, RRF with the packet order → `where_to_look`.
3.4 Short questions: RRF of packet order and whole-question ranking.

## 4. Evaluation protocol

- SWE-bench dev (225) = tuning; dev-fast (59) for iteration; Lite test (300) = holdout run once;
the 24 instances looked at in an earlier session excluded from the strict row (276).
- Metric: file-level Acc@k (all gold files in top k), bootstrap 95% CIs.
- Plain-language: four holdouts (ripgrep/Rust, click/Python, cobra/Go, axios/JS), 12 q each, gold
locked by commit before any run.

## 5. Results

### 5.1 SWE-bench Lite (holdout, 276)

| method | LLM | Acc@1 | Acc@3 | Acc@5 |
|---|---|---|---|---|
| BM25 (ours) | – | 0.301 | 0.507 | 0.587 |
| **engine** | – | **0.486** | **0.710** | **0.750** |
| Jina-Code-v2 † | – | 0.434 | 0.712 | 0.803 |
| CodeRankEmbed † | – | 0.526 | 0.777 | 0.847 |
| Agentless + GPT-4o † | ✓ | 0.672 | 0.745 | 0.745 |
| LocAgent + Claude-3.5 † | ✓ | 0.777 | 0.920 | 0.942 |
| engine + local LLM (D1) | ✓ (3B, CPU) | TODO | TODO | TODO |

† LocAgent Table 4, their 274-instance subset (head-to-head on the same subset: TODO).

### 5.2 SWE-bench dev (225) and Verified (500, TODO)

dev: BM25 0.160/0.347/0.427/0.538 → engine 0.249/0.484/0.600/0.680 (Acc@1/3/5/10).

### 5.3 Ablation (Lite holdout) — TODO (runs in progress)

| removed | Acc@1 | Acc@3 | Acc@5 |
|---|---|---|---|
| nothing | 0.486 | 0.710 | 0.750 |
| report hygiene + code tokens | TODO | | |
| definition-level ranking | TODO | | |
| localisation list (packet order only) | 0.366 | 0.536 | 0.558 |

### 5.4 Plain-language holdouts (R@3)

| set | BM25 | engine |
|---|---|---|
| ripgrep (Rust) | 0.500 | 0.500 |
| click (Python) | 0.750 | 0.917 |
| cobra (Go) | 0.875 | 0.875 |
| axios (JS, fresh) | 0.583 | 0.833 |

### 5.5 Cost

p50 2.8 s / p90 11.4 s per issue on a 6-core laptop CPU, cold index included; packet p50 11.1k
tokens; no network, no GPU.

## 6. Negative results and what the holdouts caught

| idea | effect |
|---|---|
| dense chunk retrieval on CPU | 2.6 chunks/s — hours per large repo |
| cross-encoder rerank (v2) on reports | Acc@3 0.492 → 0.373 |
| cross-encoder on short questions | +0.375/+0.208 on two looked-at sets, **−0.084 on fresh cobra** |
| blind graph-neighbour vote | Acc@1 0.254 → 0.136 |
| body term frequency in BM25F | Acc@5 0.544 → 0.509 |
| MiniLM fusion | lowered every plain-language set |
| evaluation: suffix match without segment boundary | inflated a BM25 baseline (0.958 vs 0.875) |

## 7. Threats to validity

Different instance subsets vs published numbers; single annotator for plain-language gold; 12
questions per plain-language holdout; Windows-only timing; blobless clones / network.

## 8. Conclusion — TODO
Loading
Loading