You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Let a client set the facet bucket limit per facet, and report the total number of values #802
The number of buckets a facet returns is a deployment-wide cap, invisible in the GraphQL contract and not adjustable per request. A client receives ten creators and has no way to learn there were 195, or to ask for more.
Today
maxFacetValues on the Typesense query compiler is a single value for every facet in every query; “left unset, Typesense defaults to 10 – too few for high-cardinality facets” (packages/search-typesense/src/query-compiler.ts:69). LOL sets nothing, so it serves ten.
The facet fields in the generated schema take no arguments – creator: [IRIBucket!]! – so a UI cannot say how many it will show, and cannot fetch more for one facet without changing the deployment for all of them.
Nothing reports truncation. Typesense returns facet_counts[].stats.total_values alongside the capped list, but it is never surfaced.
Raising the cap is the wrong lever: a UI that shows ten and a deployment that allows a thousand makes every sidebar pay for a thousand.
Proposal
Make the limit a request parameter bounded by a deployment ceiling, and report the total:
typeCreativeWorkFacets {
creator(limit: Int = 10, query: String): IRIFacet!material(limit: Int = 10, query: String): IRIFacet!
}
typeIRIFacet {
buckets: [IRIBucket!]!totalValues: Int # distinct values for this facet under the query; may be an estimate
}
limit per facet field, default 10. A UI that shows ten asks for ten; a “more” panel asks for a hundred on that one facet in a follow-up query – cheap, because only selected facets are computed. maxFacetValues becomes the ceiling a client may ask for; exceeding it is a caller error in the existing out-of-range style (Out-of-range perPage/page returns "Unexpected error." instead of the documented clear error #715).
totalValues from stats.total_values. Typesense returns it regardless of max_facet_values; it is exact under facet_strategy: exhaustive and approximate under top_values (the docs say so), and facet sampling makes it an estimate on large result sets – so declare it nullable and “may be an estimate”, the way the NDE generic API specification declares its facet totalItems.
limit: 100 returns the hundred buckets with the highest counts. It is a superset of the ten a limit: 10 request returned, and the client drops the ones it already shows. Typesense has no facet offset (Solr’s facet.offset has no counterpart), so there is no second page to ask for, and the schema must not suggest one: no offset, no after, no pageInfo. Document the argument as top N by count.
That does mean a follow-up query repeats the search: the same query and where, with perPage: 0 and only the one facet selected. This is not overhead to design away. Facet counts exist only relative to a result set, so the search context has to travel with every facet request – which is exactly what the NDE specification’s standalone facet resource cannot do. The repeat costs one facet-only engine search on one field, and no documents are fetched.
Should a client ever need to walk thousands of values – #533 measured keyword at about 838 distinct values in the Dataset Register and expects tens of thousands eventually – the answer is a search over the referenced collection (terms, persons) with a where on the ids it finds, not paging inside the facet. That loses the counts within the result set, and that is the honest trade.
query may later move to its own root field
Algolia serves facet-value typeahead as a separate call, searchForFacetValues, that takes the facet name, the typed text and the full search parameters as context. The reason is operational: a typeahead fires on every keystroke and wants shorter latency and its own cache, while the sidebar’s facets are computed once per page and batched.
Putting query on the facet field first is still the right first step: no new root query, and the search context is already there. If the typeahead later turns out to need different caching or latency from the batched facet query, split it into a root field such as creativeWorkFacetValues(field:, query:, where:). The {buckets, totalValues} wrapper and the argument names survive that move unchanged; only where the client sends the request changes.
Why the object, not an argument on the list
Adding limit to creator: [IRIBucket!]! would be non-breaking, but a bare list has nowhere to carry totalValues. Wrapping the facet in {buckets, totalValues} is a breaking change to the facet types – a ! commit and a minor bump on a 0.x package – and it is the change #533 would force anyway. Make it once.
Implementation notes
facet-batch.ts: the DataLoader key becomes (field, limit, query) instead of field.
groupFacetQueries keeps grouping by effective where (skip-own-filter unchanged); a group requests max_facet_values equal to the largest limit in it and truncates per field on the way out, since the Typesense cap is per search, not per facet.
A facet with a query gets its own search in the batch: facet_query targets one field.
totalValues is read from facet_counts[].stats.total_values.
No new engine capability; the REST surface (@lde/search-api-rest, when built) exposes the same three as query parameters.
Context
Found while reviewing the NDE generic API specification against LDE and LOL. That specification has the right shape for high-cardinality facets – a per-facet size on the request, a paged facet resource, and a totalItems declared as possibly an estimate – and this is the half of its facet design worth taking. (Its own facet design has a different problem: the standalone facet resource cannot receive the search context; reported there.) Presentation-layer request for the same thing on the Valeros pilot API: netwerk-digitaal-erfgoed/prototypes-data-layers#11.
The number of buckets a facet returns is a deployment-wide cap, invisible in the GraphQL contract and not adjustable per request. A client receives ten creators and has no way to learn there were 195, or to ask for more.
Today
maxFacetValueson the Typesense query compiler is a single value for every facet in every query; “left unset, Typesense defaults to 10 – too few for high-cardinality facets” (packages/search-typesense/src/query-compiler.ts:69). LOL sets nothing, so it serves ten.creator: [IRIBucket!]!– so a UI cannot say how many it will show, and cannot fetch more for one facet without changing the deployment for all of them.facet_counts[].stats.total_valuesalongside the capped list, but it is never surfaced.Raising the cap is the wrong lever: a UI that shows ten and a deployment that allows a thousand makes every sidebar pay for a thousand.
Proposal
Make the limit a request parameter bounded by a deployment ceiling, and report the total:
limitper facet field, default 10. A UI that shows ten asks for ten; a “more” panel asks for a hundred on that one facet in a follow-up query – cheap, because only selected facets are computed.maxFacetValuesbecomes the ceiling a client may ask for; exceeding it is a caller error in the existing out-of-range style (Out-of-range perPage/page returns "Unexpected error." instead of the documented clear error #715).totalValuesfromstats.total_values. Typesense returns it regardless ofmax_facet_values; it is exact underfacet_strategy: exhaustiveand approximate undertop_values(the docs say so), and facet sampling makes it an estimate on large result sets – so declare it nullable and “may be an estimate”, the way the NDE generic API specification declares its facettotalItems.queryon the same field is Search-within-a-facet (facet_query typeahead) for high-cardinality facets beyond the maxFacetValues cap #533’s facet-value typeahead (facet_query), landing where it belongs: browse withlimit, find withquery, one field.limitis a cap, not a pagelimit: 100returns the hundred buckets with the highest counts. It is a superset of the ten alimit: 10request returned, and the client drops the ones it already shows. Typesense has no facet offset (Solr’sfacet.offsethas no counterpart), so there is no second page to ask for, and the schema must not suggest one: nooffset, noafter, nopageInfo. Document the argument as top N by count.That does mean a follow-up query repeats the search: the same
queryandwhere, withperPage: 0and only the one facet selected. This is not overhead to design away. Facet counts exist only relative to a result set, so the search context has to travel with every facet request – which is exactly what the NDE specification’s standalone facet resource cannot do. The repeat costs one facet-only engine search on one field, and no documents are fetched.Should a client ever need to walk thousands of values – #533 measured
keywordat about 838 distinct values in the Dataset Register and expects tens of thousands eventually – the answer is a search over the referenced collection (terms,persons) with awhereon the ids it finds, not paging inside the facet. That loses the counts within the result set, and that is the honest trade.querymay later move to its own root fieldAlgolia serves facet-value typeahead as a separate call,
searchForFacetValues, that takes the facet name, the typed text and the full search parameters as context. The reason is operational: a typeahead fires on every keystroke and wants shorter latency and its own cache, while the sidebar’s facets are computed once per page and batched.Putting
queryon the facet field first is still the right first step: no new root query, and the search context is already there. If the typeahead later turns out to need different caching or latency from the batched facet query, split it into a root field such ascreativeWorkFacetValues(field:, query:, where:). The{buckets, totalValues}wrapper and the argument names survive that move unchanged; only where the client sends the request changes.Why the object, not an argument on the list
Adding
limittocreator: [IRIBucket!]!would be non-breaking, but a bare list has nowhere to carrytotalValues. Wrapping the facet in{buckets, totalValues}is a breaking change to the facet types – a!commit and a minor bump on a 0.x package – and it is the change #533 would force anyway. Make it once.Implementation notes
facet-batch.ts: the DataLoader key becomes(field, limit, query)instead offield.groupFacetQuerieskeeps grouping by effectivewhere(skip-own-filter unchanged); a group requestsmax_facet_valuesequal to the largestlimitin it and truncates per field on the way out, since the Typesense cap is per search, not per facet.querygets its own search in the batch:facet_querytargets one field.totalValuesis read fromfacet_counts[].stats.total_values.@lde/search-api-rest, when built) exposes the same three as query parameters.Context
Found while reviewing the NDE generic API specification against LDE and LOL. That specification has the right shape for high-cardinality facets – a per-facet
sizeon the request, a paged facet resource, and atotalItemsdeclared as possibly an estimate – and this is the half of its facet design worth taking. (Its own facet design has a different problem: the standalone facet resource cannot receive the search context; reported there.) Presentation-layer request for the same thing on the Valeros pilot API: netwerk-digitaal-erfgoed/prototypes-data-layers#11.