Features of Apache Solr

Apache Solr offers a wide range of features to make search experiences powerful, flexible and user-friendly. Depending on the data model, configuration and integration, search results can be precisely controlled, filtered, analysed and expanded. The following overview shows typical search features that can be implemented using Solr.

Key functional areas of Apache Solr


Basic Search Functions

Full-text search

Searches the entire content of documents, records or web pages and finds matches even if a term appears only in the body text.

Field search

Search specifically within certain fields, for example, only in the title, in categories or in metadata.

Phrase Search

Finds exact sequences of words and is useful when search terms need to appear in a specific order.

Boolean search (AND / OR / NOT)

Allows search queries using operators such as AND, OR or NOT to combine or exclude specific terms.

Wildcard and placeholder search

Supports placeholders such as * or ? to search for the beginnings of words, variants or incomplete terms.

Fuzzy Search

It also finds terms that are spelled similarly and helps with typos or slightly different spellings.

Category search

Enables searches within defined ranges of values, for example by price, date or number ranges.

Proximity Search

Finds terms that appear within a certain distance of one another, even if they do not appear directly next to each other as an exact phrase.

Search for exact terms

Enables a search for exactly one specific term or value, without any linguistic expansions or interpretations.

Multi-field search

Searches multiple fields simultaneously – for example, title, description and metadata – and can apply different weights to them.

Relevance & Ranking

Relevance-based sorting

Automatically ranks results according to how well their content matches the search query.

Sorting by fields

Alternatively, results can also be sorted by fields such as date, alphabetical order, price or popularity.

Boosting

Certain documents or attributes can be given greater weight so that more relevant results appear higher up in the list.

Query-Time Boosting

Allows you to adjust the weighting of fields or criteria directly within the search query, without modifying the index itself.

Function Queries

Use calculated values in the search, for example to weight results according to recency, distance or popularity.

Re-Ranking

Adjust the ranking retrospectively for the top results to further optimise particularly relevant results.

Learning to Rank (LTR)

Enables the use of self-learning ranking models to improve the order of search results based on data.

Content Elevation / Query Elevation

Specific content can be prominently featured for defined search queries.

Decay-/Distance-Boosting

Weights search results based on how close a value is to a desired target value, for example in terms of recency, price or geographical distance.

Personalised ranking

Tailor the order of search results to individual user characteristics, interests or past behaviour.

Suggestions & Fault Tolerance

Auto Suggest / Search Suggestions

Displays relevant search suggestions as you type, helping you find the result you’re looking for more quickly.

Autocomplete

It suggests search terms as users type and helps them formulate their search query.

Did you mean / Spell check

Detects potential typos and suggests appropriate corrections or alternative search terms.

Handling of Synonyms

Take into account different terms with the same or similar meanings, so that more relevant results are found.

Stemming / Lemmatisation

Reduces words to their base forms and takes language-specific rules into account during the search.

Stopword handling

Include or exclude very common, low-significance words such as ‘and’, ‘or’ or ‘the’, so that the search focuses more closely on relevant terms.

Normalisation / Tokenisation / Language-specific analysis

Process search terms and content linguistically, for example by standardising spellings, breaking them down into individual components and taking language-specific rules into account.

Presentation & Navigation within Results

Results Highlighting

Highlights the search terms directly within the results and displays relevant text snippets, so that users can see more quickly why a particular result was found.

Faceted search

Add filtering options to search results – such as by category, language, content type or date – so that users can narrow down their results.

Multi-select facets

Allow users to select multiple filter values within a single facet at the same time – for example, multiple categories or locations – to narrow down search results more flexibly.

Range facets

Group values into ranges – for example, prices, time periods or orders of magnitude – so that search results can be filtered by meaningful ranges.

Pivot facets

Combine several filters hierarchically – for example, first by country and then by city – so that search results can be filtered in a structured way across multiple levels.

Heatmap facets

Process raster-based geographical data and highlight the regions with a particularly high number of hits, so that spatial distributions can be quickly identified.

Facets with statistics

Expand the facets to include metrics such as count, average, minimum and maximum, so that search results can not only be filtered but also analysed in terms of content.

Drill-down / Drill-sideways

They enable search results to be narrowed down step by step using filters, whilst also showing which further options are available in other facets.

Pagination

Split large sets of results into individual pages so that search results remain clear and organised, and so that not all results have to be loaded at once.

Result clustering

Group similar search results by topic or content so that users can get a quicker overview of large numbers of results and identify relevant connections more easily.

Grouping

Groups search results that share common characteristics – for example, variants of a product or multiple results from the same source – to make the results list clearer.

Duplicate detection / De-duplication

It identifies duplicate or near-identical content and merges or hides it so that search results are presented more clearly and are more relevant to users.

Collapse & Expand

It initially groups similar results into a single main result and, if required, allows you to expand further related results to keep the results list concise.

Diversification of Results

Ensures greater variety in the results list by preventing similar results from appearing one after the other and by presenting different content, sources or types in a more balanced way.

Analyses, Key Figures & Evaluation

Number of hits

Shows how many results were found for a search query, so that users can quickly gauge the size of the results set.

Statistical aggregations

Calculate key metrics across search results – such as totals, averages or distributions – to enable further analysis of the data in addition to the search itself.

Faceted analyses

Show how search results are distributed across different filter criteria, such as category, language or type, thereby revealing patterns within the set of results.

Analytics / JSON Facet API

Enables complex analyses and nested faceted queries in a structured format, allowing search results to be analysed flexibly and presented efficiently.

Term Vectors

Store information about which terms appear in a document and how often, so that features such as highlighting, similarity searches or text analysis can be specifically supported.

Term / Phrase Frequencies

To identify how frequently individual terms or entire phrases appear in documents or sets of search results, in order to better assess relevance, patterns or key themes.

Significant Terms / More Like This

Identify particularly significant terms in a document or set of search results and help to find similar content based on shared content characteristics.

Similarity Search / Similar Documents

It finds documents with similar content or patterns of terms and is suitable for displaying related results, thematically relevant content or recommendations.

Geographical search

Geo-search

Finds results based on geographical data – for example, near a location, within an area or across a region – to enable location-based searches.

Radius search

Searches for results within a specified radius of a location, for example, all results within 10 kilometres of a specific point.

Bounding Box Search

Searches for results within a rectangular geographical area defined by coordinates, to find specific content within a section of the map.

Sort by distance

Sorts geographical search results by their proximity to a specific location, so that the nearest results are displayed first.

Geographical filters

Narrow down search results using geographical criteria, such as country, region, radius or defined areas.

Geo-faceting

Enhances geographical searches with spatial groupings or counts to show how results are distributed across regions, areas or distances.

Heatmaps / spatial visualisation

They visualise the geographical distribution of hits and show at a glance which regions have a particularly high or low number of results.

Documents, content & data sources

Search in documents / attachments

It also searches the content of files such as PDFs, Word documents and presentations, so that information in attachments can be found just as easily as regular content.

Text extraction from files

Extracts content from files such as PDFs, Office documents or other formats, so that the text within them can be used for searching, indexing and analysis.

Metadata extraction

Retrieves additional information from files or data records – such as title, author, date or format – so that this information can be used for searching, filtering and sorting.

Indexing of structured data

Processes clearly structured data fields – for example, from databases, JSON or XML – so that content can be searched, filtered and analysed in a targeted manner.

Indexing of unstructured data

It captures content without a fixed data structure – such as continuous text or documents – and processes it in such a way that it can still be searched and analysed.

Indexing from databases

Imports content directly from databases and makes it available for searching, filtering and analysis in Solr.

Indexing from XML / JSON / CSV

Imports content from structured files such as XML, JSON or CSV and prepares it for searching, filtering, sorting and analysis.

Incremental indexing

Updates only new or changed content in the index, rather than re-reading all the data from scratch, thereby saving time and resources.

Near Real Time Search

Makes new or updated content searchable shortly after it has been indexed, so that updates appear in search results almost in real time.

Nested Documents / Child Documents

Store related content in a hierarchical structure – for example, a main document with associated sub-entries – so that complex data can be searched and analysed together.

Block Join Queries

Enable search queries across hierarchically linked documents, for example between main and sub-documents, so that relationships within complex data structures can be utilised in a targeted manner.

Comprehension of language and content

Multilingual search

Supports searches across multiple languages, taking into account language-specific rules to ensure that results are found accurately, even within international content.

Language-dependent analysers

Process content according to the language using appropriate rules – for example, for hyphenation, base forms or special characters – so that search results are more linguistically accurate.

Synonyms by language

Expand search queries to include language-specific variations in meaning, so that related terms or different phrasing also yield relevant results.

Transliteration

Converts characters from different writing systems into a comparable form so that search queries return relevant results even when the transcription differs.

Phonetic search

It finds terms even when they sound similar but are spelled differently, thus helping with name searches or when entries are prone to errors.

Semantic Search

Takes into account not only exact search terms, but also their meaning and contextual relationships, in order to provide more relevant results.

Vector Search / Dense Vector Search

Finds content based on similarity of content rather than just exact terms, by comparing texts or data as vectors.

Hybrid Search (lexical + semantic)

Combines classic keyword search with semantic similarity search to find both exact matches and results that are relevant in terms of content.

Security & Access

Access Restrictions / Access Control Filtering

Filter search results by permissions or roles so that users only see content they are actually authorised to view.

Multi-tenancy

Enables separate management and searching for multiple organisations, departments or clients within a shared Solr environment.

Document-based rights filters

Check permissions directly at the level of individual documents to ensure that only authorised content is displayed, even within a shared index.

Query-level security filters

Add authorisation rules to search queries so that unauthorised content is filtered out at the query stage.

Performance, Scalability & Operations

Distributed Search

Runs search queries in parallel across multiple Solr instances or data ranges and consolidates the results into a single hit list.

Sharding

Distributes large volumes of data across several sub-indexes to ensure that search and indexing remain high-performing and scalable, even with high volumes.

Replication

Maintains identical copies of an index across multiple instances to enhance reliability and distribute search queries across multiple systems.

Cloud operations with SolrCloud

Enables distributed and fault-tolerant operation of Solr across multiple nodes, allowing for centralised control of scaling, load balancing and high availability.

Load Balancing

Distributes search queries across multiple instances so that the load is processed more evenly and the search remains stable even during periods of high traffic.

Caching

Caches frequently used data or query results so that search queries can be answered more quickly and the load on systems is reduced.

Query Routing

Forwards search queries directly to the appropriate Solr instances or data ranges so that queries can be processed more efficiently and quickly.

Failover

Ensures that, in the event of an instance failure, the system automatically switches to available systems so that the search remains accessible.

Scalable index and query processing

Enables large volumes of data and a high number of concurrent search queries to be processed efficiently by distributing the load and processing across multiple systems.

API & Integration

REST-/HTTP-API

Enables access to search, indexing and management functions via standardised HTTP requests, thereby facilitating the integration of external systems.

JSON Request API

It allows search queries and parameters to be formulated in a structured JSON format, making complex queries clearer and easier to integrate.

XML / CSV / Binary Responses

Provides search results in various output formats so that they can be further processed, exported or used directly by systems, depending on the specific use case.

Streaming Expressions

Enable the processing and analysis of large volumes of data directly in Solr, for example for aggregations, joins or multi-step analyses.

SQL Query Support

Enables search and analysis queries using SQL-like syntax, making structured data more familiar and accessible to a wider range of users.

Exporting large numbers of search results

Enables the reliable retrieval of a very large number of search results without having to retrieve them page by page, for example for analysis, further processing or data exports.

Integration with CMS, online shops, portals, DAM, PIM and DMS

It can be integrated with various source systems and platforms, enabling content from different applications to be searched and analysed centrally.

Integration of external data sources

Integrates content from external systems, databases or interfaces so that it can be indexed together and made available for search.