Features of Apache Solr
Apache Solr offers a wide range of features to make search experiences powerful, flexible and user-friendly. Depending on the data model, configuration and integration, search results can be precisely controlled, filtered, analysed and expanded. The following overview shows typical search features that can be implemented using Solr.
Key functional areas of Apache Solr
Basic Search Functions
Full-text search
Searches the entire content of documents, records or web pages and finds matches even if a term appears only in the body text.
Field search
Search specifically within certain fields, for example, only in the title, in categories or in metadata.
Phrase Search
Finds exact sequences of words and is useful when search terms need to appear in a specific order.
Boolean search (AND / OR / NOT)
Allows search queries using operators such as AND, OR or NOT to combine or exclude specific terms.
Wildcard and placeholder search
Supports placeholders such as * or ? to search for the beginnings of words, variants or incomplete terms.
Fuzzy Search
It also finds terms that are spelled similarly and helps with typos or slightly different spellings.
Category search
Enables searches within defined ranges of values, for example by price, date or number ranges.
Proximity Search
Finds terms that appear within a certain distance of one another, even if they do not appear directly next to each other as an exact phrase.
Search for exact terms
Enables a search for exactly one specific term or value, without any linguistic expansions or interpretations.
Multi-field search
Searches multiple fields simultaneously – for example, title, description and metadata – and can apply different weights to them.
Relevance & Ranking
Relevance-based sorting
Automatically ranks results according to how well their content matches the search query.
Sorting by fields
Alternatively, results can also be sorted by fields such as date, alphabetical order, price or popularity.
Boosting
Certain documents or attributes can be given greater weight so that more relevant results appear higher up in the list.
Query-Time Boosting
Allows you to adjust the weighting of fields or criteria directly within the search query, without modifying the index itself.
Function Queries
Use calculated values in the search, for example to weight results according to recency, distance or popularity.
Re-Ranking
Adjust the ranking retrospectively for the top results to further optimise particularly relevant results.
Learning to Rank (LTR)
Enables the use of self-learning ranking models to improve the order of search results based on data.
Content Elevation / Query Elevation
Specific content can be prominently featured for defined search queries.
Decay-/Distance-Boosting
Weights search results based on how close a value is to a desired target value, for example in terms of recency, price or geographical distance.
Personalised ranking
Tailor the order of search results to individual user characteristics, interests or past behaviour.
Suggestions & Fault Tolerance
Auto Suggest / Search Suggestions
Displays relevant search suggestions as you type, helping you find the result you’re looking for more quickly.
Autocomplete
It suggests search terms as users type and helps them formulate their search query.
Did you mean / Spell check
Detects potential typos and suggests appropriate corrections or alternative search terms.
Handling of Synonyms
Take into account different terms with the same or similar meanings, so that more relevant results are found.
Stemming / Lemmatisation
Reduces words to their base forms and takes language-specific rules into account during the search.
Stopword handling
Include or exclude very common, low-significance words such as ‘and’, ‘or’ or ‘the’, so that the search focuses more closely on relevant terms.
Normalisation / Tokenisation / Language-specific analysis
Process search terms and content linguistically, for example by standardising spellings, breaking them down into individual components and taking language-specific rules into account.
Presentation & Navigation within Results
Results Highlighting
Highlights the search terms directly within the results and displays relevant text snippets, so that users can see more quickly why a particular result was found.
Faceted search
Add filtering options to search results – such as by category, language, content type or date – so that users can narrow down their results.
Multi-select facets
Allow users to select multiple filter values within a single facet at the same time – for example, multiple categories or locations – to narrow down search results more flexibly.
Range facets
Group values into ranges – for example, prices, time periods or orders of magnitude – so that search results can be filtered by meaningful ranges.
Pivot facets
Combine several filters hierarchically – for example, first by country and then by city – so that search results can be filtered in a structured way across multiple levels.
Heatmap facets
Process raster-based geographical data and highlight the regions with a particularly high number of hits, so that spatial distributions can be quickly identified.
Facets with statistics
Expand the facets to include metrics such as count, average, minimum and maximum, so that search results can not only be filtered but also analysed in terms of content.
Drill-down / Drill-sideways
They enable search results to be narrowed down step by step using filters, whilst also showing which further options are available in other facets.
Pagination
Split large sets of results into individual pages so that search results remain clear and organised, and so that not all results have to be loaded at once.
Result clustering
Group similar search results by topic or content so that users can get a quicker overview of large numbers of results and identify relevant connections more easily.
Grouping
Groups search results that share common characteristics – for example, variants of a product or multiple results from the same source – to make the results list clearer.
Duplicate detection / De-duplication
It identifies duplicate or near-identical content and merges or hides it so that search results are presented more clearly and are more relevant to users.
Collapse & Expand
It initially groups similar results into a single main result and, if required, allows you to expand further related results to keep the results list concise.
Diversification of Results
Ensures greater variety in the results list by preventing similar results from appearing one after the other and by presenting different content, sources or types in a more balanced way.
Analyses, Key Figures & Evaluation
Number of hits
Shows how many results were found for a search query, so that users can quickly gauge the size of the results set.
Statistical aggregations
Calculate key metrics across search results – such as totals, averages or distributions – to enable further analysis of the data in addition to the search itself.
Faceted analyses
Show how search results are distributed across different filter criteria, such as category, language or type, thereby revealing patterns within the set of results.
Analytics / JSON Facet API
Enables complex analyses and nested faceted queries in a structured format, allowing search results to be analysed flexibly and presented efficiently.
Term Vectors
Store information about which terms appear in a document and how often, so that features such as highlighting, similarity searches or text analysis can be specifically supported.
Term / Phrase Frequencies
To identify how frequently individual terms or entire phrases appear in documents or sets of search results, in order to better assess relevance, patterns or key themes.
Significant Terms / More Like This
Identify particularly significant terms in a document or set of search results and help to find similar content based on shared content characteristics.
Similarity Search / Similar Documents
It finds documents with similar content or patterns of terms and is suitable for displaying related results, thematically relevant content or recommendations.
Geographical search
Geo-search
Finds results based on geographical data – for example, near a location, within an area or across a region – to enable location-based searches.
Radius search
Searches for results within a specified radius of a location, for example, all results within 10 kilometres of a specific point.
Bounding Box Search
Searches for results within a rectangular geographical area defined by coordinates, to find specific content within a section of the map.
Sort by distance
Sorts geographical search results by their proximity to a specific location, so that the nearest results are displayed first.
Geographical filters
Narrow down search results using geographical criteria, such as country, region, radius or defined areas.
Geo-faceting
Enhances geographical searches with spatial groupings or counts to show how results are distributed across regions, areas or distances.
Heatmaps / spatial visualisation
They visualise the geographical distribution of hits and show at a glance which regions have a particularly high or low number of results.
Documents, content & data sources
Search in documents / attachments
It also searches the content of files such as PDFs, Word documents and presentations, so that information in attachments can be found just as easily as regular content.
Text extraction from files
Extracts content from files such as PDFs, Office documents or other formats, so that the text within them can be used for searching, indexing and analysis.
Metadata extraction
Retrieves additional information from files or data records – such as title, author, date or format – so that this information can be used for searching, filtering and sorting.
Indexing of structured data
Processes clearly structured data fields – for example, from databases, JSON or XML – so that content can be searched, filtered and analysed in a targeted manner.
Indexing of unstructured data
It captures content without a fixed data structure – such as continuous text or documents – and processes it in such a way that it can still be searched and analysed.
Indexing from databases
Imports content directly from databases and makes it available for searching, filtering and analysis in Solr.
Indexing from XML / JSON / CSV
Imports content from structured files such as XML, JSON or CSV and prepares it for searching, filtering, sorting and analysis.
Incremental indexing
Updates only new or changed content in the index, rather than re-reading all the data from scratch, thereby saving time and resources.
Near Real Time Search
Makes new or updated content searchable shortly after it has been indexed, so that updates appear in search results almost in real time.
Nested Documents / Child Documents
Store related content in a hierarchical structure – for example, a main document with associated sub-entries – so that complex data can be searched and analysed together.
Block Join Queries
Enable search queries across hierarchically linked documents, for example between main and sub-documents, so that relationships within complex data structures can be utilised in a targeted manner.
Comprehension of language and content
Multilingual search
Supports searches across multiple languages, taking into account language-specific rules to ensure that results are found accurately, even within international content.
Language-dependent analysers
Process content according to the language using appropriate rules – for example, for hyphenation, base forms or special characters – so that search results are more linguistically accurate.
Synonyms by language
Expand search queries to include language-specific variations in meaning, so that related terms or different phrasing also yield relevant results.
Transliteration
Converts characters from different writing systems into a comparable form so that search queries return relevant results even when the transcription differs.
Phonetic search
It finds terms even when they sound similar but are spelled differently, thus helping with name searches or when entries are prone to errors.
Semantic Search
Takes into account not only exact search terms, but also their meaning and contextual relationships, in order to provide more relevant results.
Vector Search / Dense Vector Search
Finds content based on similarity of content rather than just exact terms, by comparing texts or data as vectors.
Hybrid Search (lexical + semantic)
Combines classic keyword search with semantic similarity search to find both exact matches and results that are relevant in terms of content.
Security & Access
Access Restrictions / Access Control Filtering
Filter search results by permissions or roles so that users only see content they are actually authorised to view.
Multi-tenancy
Enables separate management and searching for multiple organisations, departments or clients within a shared Solr environment.
Document-based rights filters
Check permissions directly at the level of individual documents to ensure that only authorised content is displayed, even within a shared index.
Query-level security filters
Add authorisation rules to search queries so that unauthorised content is filtered out at the query stage.
Performance, Scalability & Operations
Distributed Search
Runs search queries in parallel across multiple Solr instances or data ranges and consolidates the results into a single hit list.
Sharding
Distributes large volumes of data across several sub-indexes to ensure that search and indexing remain high-performing and scalable, even with high volumes.
Replication
Maintains identical copies of an index across multiple instances to enhance reliability and distribute search queries across multiple systems.
Cloud operations with SolrCloud
Enables distributed and fault-tolerant operation of Solr across multiple nodes, allowing for centralised control of scaling, load balancing and high availability.
Load Balancing
Distributes search queries across multiple instances so that the load is processed more evenly and the search remains stable even during periods of high traffic.
Caching
Caches frequently used data or query results so that search queries can be answered more quickly and the load on systems is reduced.
Query Routing
Forwards search queries directly to the appropriate Solr instances or data ranges so that queries can be processed more efficiently and quickly.
Failover
Ensures that, in the event of an instance failure, the system automatically switches to available systems so that the search remains accessible.
Scalable index and query processing
Enables large volumes of data and a high number of concurrent search queries to be processed efficiently by distributing the load and processing across multiple systems.
API & Integration
REST-/HTTP-API
Enables access to search, indexing and management functions via standardised HTTP requests, thereby facilitating the integration of external systems.
JSON Request API
It allows search queries and parameters to be formulated in a structured JSON format, making complex queries clearer and easier to integrate.
XML / CSV / Binary Responses
Provides search results in various output formats so that they can be further processed, exported or used directly by systems, depending on the specific use case.
Streaming Expressions
Enable the processing and analysis of large volumes of data directly in Solr, for example for aggregations, joins or multi-step analyses.
SQL Query Support
Enables search and analysis queries using SQL-like syntax, making structured data more familiar and accessible to a wider range of users.
Exporting large numbers of search results
Enables the reliable retrieval of a very large number of search results without having to retrieve them page by page, for example for analysis, further processing or data exports.
Integration with CMS, online shops, portals, DAM, PIM and DMS
It can be integrated with various source systems and platforms, enabling content from different applications to be searched and analysed centrally.
Integration of external data sources
Integrates content from external systems, databases or interfaces so that it can be indexed together and made available for search.