Skip to content

feat!: hosted server with Event Registry login, mentions, aggregates and a single search tool - #1

Merged
eriknovak merged 18 commits into
EventRegistry:mainfrom
eriknovak:main
Oct 9, 2026
Merged

eriknovak merged 18 commits into
EventRegistry:mainfrom
eriknovak:main

Conversation

@eriknovak

@eriknovak eriknovak commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds a hosted server with Event Registry login, three new data surfaces (aggregates, breaking events, mentions), and replaces the three search tools with one search tool whose queries are composed server-side. The redesign was driven by an internal benchmark aimed at validating every tool, cutting NewsAPI.ai token spend and the MCP instruction/schema context, while keeping answer quality.

Changes

  • Hosted server (src/http.ts): stateless Streamable HTTP entry point for https://mcp.newsapi.ai/mcp. The MCP client runs the OAuth login against auth.id.eventregistry.org; the server verifies the JWT (iss, aud, exp, JWKS via OpenID discovery), serves protected-resource metadata and runs each request on the caller's bearer token. Ships as a Docker image with GET /healthz; configured via MCP_PUBLIC_URL, MCP_AUTH_ISSUER, PORT. Deployment guide in docs/deployment.md.
  • Local server login (src/oauth.ts): without NEWSAPI_KEY, npx newsapi-mcp logs in through the browser (authorization code + PKCE on loopback ports 51337–51339), keeps tokens in the OS credential store with a file fallback, refreshes with rotation and retries once on 401. login/logout subcommands manage the stored login; NEWSAPI_KEY still selects API-key mode. Tool errors explain an unlinked account or an expired login per auth mode.
  • Hosted results marked as source material: results carry audience: ["assistant"], sit inside <source_material> tags and tell the model to report key points with links rather than paste raw results. No client currently honours the audience annotation, so it is a hint only.
  • Aggregate result types: article and event searches accept resultType (timeAggr, sourceAggr, authorAggr, keywordAggr, locAggr, conceptAggr, categoryAggr, sentimentAggr, langAggr) rendered as label/count rows.
  • get_breaking_events wraps event/getBreakingEvents with count, page and minimum score.
  • Mentions (kind: "mentions"): sentences tagged with an event type (acquisition, layoffs, launch, recall…) with entities, sentiment, fact level and article link; mention-only filters (eventTypeUri, industryUri, sdgUri, sasbUri, esgUri, factLevel…) and eventTypeAggr. suggest(type: "eventTypes") resolves event type names.
  • Single search tool (breaking): search_articles / search_events / search_mentions become search({kind}) with generic page, count, sortBy, options.sortByAsc. Kind-only params are tagged in their descriptions and rejected with a corrective 400 for other kinds. Static session-start cost drops from ~15.8k to ~10.5k tokens.
  • Server-side query composition (src/query.ts): flat filters given alongside an advanced query are merged (leaf filters ANDed, ignore* → $not, duplicate/sentiment/rank flags → $filter); arrays and $not lists normalised; boolean keyword strings ("Tesla AND (recall OR lawsuit) NOT Musk") switch to keywordSearchMode: exact automatically; query-shaped 400s return rewrite guidance.
  • Cost guards: a 31-day window is applied when no date filter is given (an unbounded search costs 65×); searches costing more than 1 API token say why under the result; default count is 50 (scans ask for 100); the response cap is 50k chars so oversized results are truncated instead of dropped by the client.
  • Output: scan rows (articleBodyLen: 0) are one compact row per article (# | uri | date | source | title), with get_article_details supplying text and URL; rows print YYYY-MM-DD HH:MM and date-sorted article pages are re-sorted by publish time (the API sorts by crawl time). Rarely used filters live under one options object.
  • Housekeeping: suggest cache removed (no measurable benefit), lock file bumped to clear npm audit, README/CHANGELOG/skill/guide updated, CONTEXT.md added.

Breaking changes: the three search tools are gone (use search with kind); articlesPage/Count/SortBy etc. become page/count/sortBy; sentiment, source rank, author/location/source-group filters, secondary ignore* filters, date-mention filters, dataType and *SortByAsc are accepted only inside options; searches without a date filter cover the last 31 days.

Test plan

  • npm test → 11 files, 396 tests passed (unit, server integration via InMemoryTransport, hosted server with locally signed JWTs, loopback login flow).
  • Internal benchmark (15 questions, live API, headless claude -p client with Sonnet and Haiku): NewsAPI tokens for the question set fell from 105 to 47 (Sonnet) and 133 to 56 (Haiku), session-start context from ~15.8k to ~10.5k tokens, answer quality flat to slightly up; single runs, indicative only.

Decisions

  • Login goes through Event Registry (hosted: client-run OAuth; local: loopback PKCE) rather than API keys.
  • The hosted server is stateless, fresh McpServer per request.
  • Hosted results are marked as source material; the audience annotation is kept as a hint although no client honours it.
  • 31-day default window kept despite changing behaviour for intentional archive searches: the cost cliff is 65× and a note tells the model how to widen.
  • Three search tools merged into one search with kind, no aliases for the old names.
  • A natural-language → query tool (Sentry pattern) was not pursued; composition plus boolean strings covered the benchmark cases without an extra LLM call.

Hosted server (src/http.ts): stateless Streamable HTTP entry point for
https://mcp.newsapi.ai/mcp. The MCP client runs the OAuth login against
auth.id.eventregistry.org; the server verifies the JWT (iss, aud, exp, JWKS
from OpenID discovery, Hydra scp scopes), serves protected-resource metadata,
and runs each request on the caller's own bearer token. Ships as a Docker
image with GET /healthz.

Local server login (src/oauth.ts): without NEWSAPI_KEY the npm package runs
the authorization-code + PKCE flow itself as a public client on a loopback
redirect (ports 51337-51339), keeps tokens in the OS credential store with a
file fallback, refreshes one at a time with rotation, and renews once and
retries on 401. `login`/`logout` subcommands manage the stored login;
NEWSAPI_KEY keeps API-key mode.

Tool errors explain an unlinked Event Registry account and an expired login
per auth mode. The suggest cache is removed: the endpoints are cheap and it
bought no speed. ADRs 0001 (login through Event Registry) and 0002
(stateless hosted server), CONTEXT.md and docs/agents are added.
feat: add hosted server and Event Registry login
The hosted server cannot hide tool results from the user, so it labels them for the model only and tells the model to report key points with links instead of pasting article text or raw results (ADR-0003). Results carry audience ["assistant"], sit inside <source_material> tags and end with a reminder; tool descriptions, server instructions and the guide carry the rules. The local server is unchanged.
feat(hosted): mark tool results as source material for the model
Claude Code, Claude Desktop, Claude.ai and Codex show every tool result
block to the user regardless of annotations.audience; Cursor likely does
too. ADR-0003 now carries the per-client findings and the annotation is
marked as a hint only.
Bumps transitive packages (express chain, MCP SDK, hono, vitest, eslint
tooling) to patched versions. No package.json range changes.
search_articles and search_events accept resultType: the list by default,
or one aggregate (timeAggr, sourceAggr, authorAggr, keywordAggr, locAggr,
conceptAggr, categoryAggr, sentimentAggr, langAggr for articles) that
summarises every match in one call. Aggregate requests send the filters
only and render as numbered label/count rows; sentimentAggr buckets are
labelled as ranges and sourceAggr reads its nested counts.frequency.

get_breaking_events wraps event/getBreakingEvents with count, page and
minimum score, filtered like events with the breaking score kept.

Docs, server instructions, the guide resource and the news skill
allow-list name the new tool and when an aggregate replaces a scan.
The spec lives in .scratch/api-aggregates-breaking.
Specs and issues under .scratch/ are the local issue tracker, not part of
the published project.
feat: add aggregate result types and breaking events tool
Add search_mentions on the getMentions endpoint: sentences tagged with an
event type (acquisition, layoffs, launch, recall...) with the entities
involved, sentiment, fact level and a link to the article. Supports the
shared content filters the endpoint accepts, the mention-only filters
(eventTypeUri, industryUri, sdgUri, sasbUri, esgUri, factLevel, sentence
index range, showDuplicates), mentions* paging and sorting, includeFields
groups (slots, categories, frameworks, metadata, full) and the aggregate
resultType values including eventTypeAggr.

Route through eventType/mention with action getMentions: the REST path on
the documentation page is rejected by the API.

Add suggest(type: "eventTypes") to resolve event type names to URIs, a
shared pagination footer for the list formatters, and document the new
tool in the README, server instructions, guide resource and news skill.
feat: add search_mentions tool and event type lookup
Search tools now build the request through a shared query module:
flat filters given alongside an advanced `query` are merged into it
(leaf filters ANDed, ignore* as $not, duplicate/sentiment/rank flags
into $filter) instead of being rejected by the API; arrays and $not
lists are normalised; a 31-day window is applied when no date is
given, since an unbounded search costs 65x; boolean keyword strings
("Tesla AND (recall OR lawsuit) NOT Musk") switch to exact mode
automatically. The query param is documented as a grammar with an
example, and query-shaped 400s return rewrite guidance.

Scan results (articleBodyLen: 0) render one compact row per article
without the URL; get_article_details supplies text and URL. Server
notes (defaults applied, body truncation) print under each result.

Rarely used filters move under one `options` object per search tool
and the shared descriptions are shortened, cutting the static schema
below its previous size; the schema builder supports nested objects.

BREAKING CHANGE: sentiment, source rank, author/location/source-group
filters, secondary ignore* filters, date-mention filters, dataType and
the *SortByAsc flags are accepted only inside `options`, and searches
without a date filter cover the last 31 days.
Replace search_articles, search_events and search_mentions with a single
search tool taking kind: "articles" | "events" | "mentions". The shared
filters, the query grammar and the options object are now loaded once,
cutting the session-start cost from about 14k to about 10.5k tokens.
Paging is generic (page, count, sortBy, options.sortByAsc) and mapped to
the API names per kind; kind-only params are tagged in their descriptions
and rejected with a corrective 400 for other kinds, as are a foreign
sortBy or aggregate.

Keep NewsAPI spend from rising with the smaller schema: searches that
cost more than one API token say why under the result (events cost 5,
dates reaching back more than 31 days cost 5-10x), the instructions map
recent periods to forceMaxDataTimeWindow and stop reworded repeats of a
scan that already returned enough rows, without limiting how many calls
a model may make. Default count drops to 50 for articles and mentions
(scans still ask for 100) and the response cap to 50k chars, so an
oversized result is truncated instead of dropped by the client.

BREAKING CHANGE: the tools search_articles, search_events and
search_mentions are gone; call search with kind. articlesPage/Count/
SortBy, eventsPage/Count/SortBy and mentionsPage/Count/SortBy are now
page, count and sortBy.
feat!: compose queries server-side and merge the search tools into one
Article, scan, detail and mention rows print "YYYY-MM-DD HH:MM" from
the publish time instead of the day alone, so same-day items stay
orderable for "latest news" questions. Events keep their day-only date.

The API's "date" sort orders by crawl time, which can trail publication
by minutes to hours; a date-sorted article page is now re-sorted by
publish time server-side, honouring sortByAsc. Other sorts are untouched.
@eriknovak eriknovak added the enhancement New feature or request label Oct 9, 2026
The decision records and the agent workflow notes are internal; the
reasons they held now live in the code comments and docs that cite them.
@eriknovak
eriknovak merged commit a8aef3f into EventRegistry:main Oct 9, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant