DATA & INDEXING
Apache Solr and LipiCore
Solr is the indexing foundation. LipiCore Engine is what your applications talk to. Here is what each layer does, and why you never call Solr directly.
What Solr is
Apache Solr is an open-source search server. It takes documents, indexes them with Apache Lucene - the library that most search engines, open and commercial, are built on - and answers queries over HTTP. It has been in production use since the mid-2000s, it is maintained by the Apache Software Foundation, and it is designed for exactly the problem this site keeps returning to: large volumes of data and complex queries that a relational database answers slowly.
Lucene gives it the structures: the inverted index for text, doc values for sorting and counting, per-field analysis so Jackets and jacket meet. Solr wraps them in a server: a schema, an HTTP API, collections, facets, suggesters, and clustering for scale and availability. What is an index server? explains those structures from first principles.
What Solr gives you
Everything that makes search and counting fast. Full-text search with relevance scoring, fuzzy matching for typos, suggesters for typeahead, boosting by field. Filters on any indexed field. Field, range, stat and nested facets, computed on the filtered set. Sorting and paging at depth. And an operational model - collections, replicas, shards - that scales from one server to a cluster. It is a complete, proven indexing foundation, and nothing here is reinvented on top of it.
What a product adds on top
Solr is a search server, not a business-data product. It does not know what an order is, where your orders live, how often they change, or what your application screen needs back. A product built on it supplies that layer.
| Capability | Apache Solr provides | LipiCore Engine adds |
|---|---|---|
| Indexing | Inverted index, doc values, per-field analysis | Collections modelled on your entities, one index and schedule each |
| Getting data in | An update API that accepts documents | Push through the Engine API, or pull from your database on a schedule; full or incremental imports |
| Schema | Field types and analyzers | Mapping from your source fields to search, filter, range and stat fields |
| Querying | A rich but Solr-specific query language | Query translation: one REST request naming the collection and what you need |
| Facets | Field, range, stat and nested facets | Returned beside the records in the same call |
| Search quality | Relevance, fuzzy matching, suggesters, boosting | Configured per collection: weighted fields, typeahead, spell correction from your data |
| Operations | Clustering and replication (SolrCloud) | Deployed and run inside your infrastructure, upgrades handled |
| Who talks to it | Whoever learns Solr | Your applications, through REST APIs; and Insights and the AI Assistant |
Collections modelled on your entities. Products, orders, customers, invoices, each with its own index and its own refresh schedule, mapped from the fields in your source systems.
Getting data in and keeping it fresh. Push through an API or pull from the database on a schedule; a full import or only new and changed records; nothing written back. Solr accepts documents; the product decides which documents, when, and how to detect change. Incremental indexing covers that part.
Query translation. Solr's query language is powerful and its own. A product exposes one plain REST request - the collection, the text, the filters, the facets - and translates it, so an application developer never writes a Solr query and never has to learn how facets are expressed in one.
One response for a screen. Records, counts, ranges and totals together, shaped for the screen that asked.
Why your applications never call Solr
Three reasons, and they compound. Coupling: a screen written against Solr's API is written against Solr's version, schema and quirks, and every upgrade is an application change. Expertise: Solr is a specialist system, and a team should not need a Solr engineer to add a filter to a listing page. Safety: the product's API exposes the operations screens need and nothing else, which is also what makes it a sensible thing to put an assistant or an MCP server on top of. Solr stays behind the product, inside your infrastructure, and the REST API is the only door.
There is also a quieter reason. Solr's API is general: it will do whatever it is asked, including expensive queries and schema changes, to anyone who can reach it. A product's API is specific: this collection, these filters, these facets, this page. Keeping applications on the specific API is what keeps the general one private, and it is the same principle that makes an assistant or an MCP server safe to put on top: they, too, only see the specific API.
What this means for your team
You get the engine without the engineering. The indexing foundation is open source, mature and widely understood, which matters when a security review asks what is underneath. The product layer is what your developers touch, and it looks like any other REST API in your stack. If you already know Solr, nothing is hidden from you; if you do not, nothing requires you to.