The JSON Facet API can now change the domain for facet commands, essentially doing a block join and moving from parents to children, or children to parents before calculating the facet data.
For example, if you indexed chapters with pages as nested child documents, then you could map from chapters to pages before faceting by adding the following parameter to the facet command:
domain : { blockChildren : "type:chapter" }
Or if you started with pages, you could map to chapters with
domain : { blockParent : "type:chapter" }
Note that in both cases, we provide the parent filter (how parent documents are defined) of “type:chapter” regardless of which direction we are mapping.
See this Nested Objects tutorial for complete examples of combining faceting and block join / nested documents.
Major improvements in performance of the new Facet Module / JSON Facet API. See the facet performance benchmarks for more details and benchmark results.
The MoreLikeThis QParser mlt now supports all options provided by the MLT Handler. The query parser is much more versatile than the handler as it works in cloud mode as well as anywhere a normal query can be specified.
Example (on techproducts index):
q={!mlt qf=name mintf=1 mindf=1}SP2514N
More documentation on the mlt parser can be found in the Solr Ref Guide
Solr’s pseudo-join query parser has a new optional attribute score that can be used specify the scores produced on the resulting documents. It’s value can be min, max,avg,or total.
Smile is a binary data interchange format that is very close to Solr’s own “javabin” (encoded sizes are very close). Adding wt=smile to a request will cause the response to come back in this format.
In addition to many other improvements in the security framework, Solr now includes an AuthenticationPlugin implementing HTTP Basic Auth that stores credentials securely in ZooKeeper. This is a simple way to require a username and password for anyone accessing Solr’s admin screen or APIs.
Solr 5 has a completely re-written faceted search and analytics module with a structured JSON API to control the faceting and analytics commands. NOTE: Some examples use syntax only supported in later Solr 5 releases, or even Solr 6. Download a recent Solr release or snapshot to try them out.
The structured nature of nested sub-facets are more naturally expressed in a nested structure like JSON rather than the flat structure that normal query parameters provide.
Goals of the new Faceting Module:
First class JSON support
Easier programmatic construction of complex nested facet commands
Support a much more canonical response format that is easier for clients to parse
First class analytics support
Ability to sort facet buckets by any calculated metric
Support a cleaner way to do distributed faceting
Support better integration with other search features
Of course if you prefer to use Solr’s existing faceting capabilities, that’s fine too. You can even use both at once if you want!
UPDATE: The JSON Facet API is now part of the JSON Request API, so a complete request may be expressed in JSON.
Some of the ease-of-use enhancements over traditional Solr faceting come from the inherent nested structure of JSON. As an example, here is the faceting command for two different range facets using Solr’s flat legacy API:
&facet=true
&facet.range={!key=age_ranges}age
&f.age.facet.range.start=0
&f.age.facet.range.end=100
&f.age.facet.range.gap=10
&facet.range={!key=price_ranges}price
&f.price.facet.range.start=0
&f.price.facet.range.end=1000
&f.price.facet.range.gap=50
And here is the equivalent faceting command in the new JSON Faceting API:
{
age_ranges: {
type : range
field : age,
start : 0,
end : 100,
gap : 10
}
,
price_ranges: {
type : range
field : price,
start : 0,
end : 1000,
gap : 50
}
}
These aren’t even nested facets, but already one can see how much nicer the JSON API looks. With deeply nested sub-facets and statistics, the clarity of the inherently nested JSON API only grows.
A number of JSON extensions have been implemented to further increase the clarity and ease of constructing a JSON faceting command by hand. For example:
{ // this is a single-line comment, which can help add clarity to large JSON commands
/* traditional C-style comments are also supported */
x : "avg(price)" , // Simple strings can occur unquoted
y : 'unique(manu)'// Strings can also use single quotes (easier to embed in another String)
Nicely indented JSON is very easy to understand. If you get a large piece of non-indented JSON somehow, and are trying to make sense of it, you can cut and paste into one of the online validators: http://jsonlint.comhttp://jsonformatter.curiousconcept.com Both of these validators will indent your JSON, even when it contains extensions unsupported by them (such as comments or bare strings).
There are two types of facets, one that breaks up the domain into multiple buckets, and aggregations / facet functions that provide information about the set of documents belonging to each bucket.
Faceting can be nested! Any bucket produced by faceting can further be broken down into multiple buckets by a sub-facet.
Statistics are now fully integrated into faceting. Since we start off with a single facet bucket with a domain defined by the main query and filters, we can even ask for statistics for this top level bucket, before breaking up into further buckets via faceting. Example:
json.facet={
x : "avg(price)", // the average of the price field will appear under "x"
y : "unique(manufacturer)"// the number of unique manufacturers will appear under "y"
}
See facet functions for a complete list of the available aggregation functions.
The general form of the JSON facet commands are: <facet_name> : { <facet_type> : <facet_parameter(s)> } Example: top_authors : { terms : { field : authors, limit : 5 } }
After Solr 5.2, a flatter structure with a “type” field may also be used: <facet_name> : { "type" : <facet_type> , <other_facet_parameter(s)> } Example: top_authors : { type : terms, field : authors, limit : 5 }
The results will appear in the response under the facet name specified. Facet commands are specified using json.facet request parameters.
The terms facet, or field facet, produces buckets from the unique values of a field. The field needs to be indexed or have docValues.
The simplest form of the terms facet
{
top_genres : { terms : genre_field }
}
An expanded form allows for more parameters:
{
top_genres : {
type : terms,
field : genre_field,
limit : 3,
mincount : 2
}
}
Example response:
"top_genres":{
"buckets":[
{
"val":"Science Fiction",
"count":143},
{
"val":"Fantasy",
"count":122},
{
"val":"Biography",
"count":28}
]
}
Parameters:
field - The field name to facet over.
offset - Used for paging, this skips the first N buckets. Defaults to 0.
limit - Limits the number of buckets returned. Defaults to 10.
mincount - Only return buckets with a count of at least this number. Defaults to 1.
sort - Specifies how to sort the buckets produced. “count” specifies document count, “index” sorts by the index (natural) order of the bucket value. One can also sort by any facet function / statistic that occurs in the bucket. The default is “count desc”. This parameter may also be specified in JSON like sort:{count:desc}. The sort order may either be “asc” or “desc”
missing - A boolean that specifies if a special “missing” bucket should be returned that is defined by documents without a value in the field. Defaults to false.
numBuckets - A boolean. If true, adds “numBuckets” to the response, an integer representing the number of buckets for the facet (as opposed to the number of buckets returned). Defaults to false.
allBuckets - A boolean. If true, adds an “allBuckets” bucket to the response, representing the union of all of the buckets. For multi-valued fields, this is different than a bucket for all of the documents in the domain since a single document can belong to multiple buckets. Defaults to false.
prefix - Only produce buckets for terms starting with the specified prefix.
method - Provides an execution hint for how to facet the field.
method:uif - Stands for UninvertedField, a method of faceting indexed, multi-valued fields using top-level data structures that optimize for performance over NRT capabilities.
method:dv - Stands for DocValues, a method of faceting indexed, multi-valued fields using per-segment data structures. This method mirrors faceting on real docValues fields but works by building on-heap docValues on the fly from the index when docValues aren’t available. This method is better for a quickly changing index.
method:stream - This method creates each individual facet bucket (including any sub-facets) on-the-fly while streaming the response back to the requester. Currently only supports sorting by index order.
The range facet produces multiple range buckets over numeric fields or date fields.
Range facet example:
{
prices : {
type : range,
field : price,
start : 0,
end : 100,
gap : 20
}
}
Example response:
"prices":{
"buckets":[
{
"val":0.0, // the bucket value represents the start of each range. This bucket covers 0-20
"count":5},
{
"val":20.0,
"count":3},
{
"val":40.0,
"count":2},
{
"val":60.0,
"count":1},
{
"val":80.0,
"count":1}
]
}
To ease migration, these parameter names, values, and semantics were taken directly from the old-style (non JSON) Solr range faceting.
Parameters:
field - The numeric field or date field to produce range buckets from
mincount - Minimum document count for the bucket to be included in the response. Defaults to 0.
start - Lower bound of the ranges
end - Upper bound of the ranges
gap - Size of each range bucket produced
hardend - A boolean, which if true means that the last bucket will end at “end” even if it is less than “gap” wide. If false, the last bucket will be “gap” wide, which may extend past “end”.
other - This param indicates that in addition to the counts for each range constraint between facet.range.start and facet.range.end, counts should also be computed for…
"before" all records with field values lower then lower bound of the first range
"after" all records with field values greater then the upper bound of the last range
"between" all records with field values between the start and end bounds of all ranges
"none" compute none of this information
"all" shortcut for before, between, and after
include - By default, the ranges used to compute range faceting between facet.range.start and facet.range.end are inclusive of their lower bounds and exclusive of the upper bounds. The “before” range is exclusive and the “after” range is inclusive. This default, equivalent to lower below, will not result in double counting at the boundaries. This behavior can be modified by the facet.range.include param, which can be any combination of the following options…
"lower" all gap based ranges include their lower bound
"upper" all gap based ranges include their upper bound
"edge" the first and last gap ranges include their edge bounds (ie: lower for the first one, upper for the last one) even if the corresponding upper/lower option is not specified
"outer" the “before” and “after” ranges will be inclusive of their bounds, even if the first or last ranges already include those boundaries.
Traditional faceted search (also called guided navigation) involves counting search results that belong to categories (also called facet constraints). The new facet functions in Solr extends normal faceting by allowing additional aggregations on document fields themselves. Combined with the new Sub-facet feature, this provides powerful new realtime analytics capabilities. Also see the page about the new JSON Facet API.
Faceting involves breaking up the domain into multiple buckets and providing information about each bucket. There are multiple aggregation functions / statistics that can be used:
Aggregation
Example
Effect
sum
sum(sales)
summation of numeric values
avg
avg(popularity)
average of numeric values
sumsq
sumsq(rent)
sum of squares
min
min(salary)
minimum value
max
max(mul(price,popularity))
maximum value
unique
unique(state)
number of unique values (count distinct)
hll
hll(state)
number of unique values using the HyperLogLog algorithm
percentile
percentile(salary,50,75,99,99.9)
calculates percentiles
stddev
stddev(salary)
calculates standard deviation (Solr6.6+)
variance
variance(salary)
calculates variance (Solr 6.6+)
Numeric aggregation functions such as avg can be on any numeric field, or on another function of multiple numeric fields.
See Count Distinct in Solr for more information on distributed cardinality estimation / calcDistinct.
The faceting domain starts with the set of documents that match the main query and filters. We can ask for statistics over this whole set of documents:
http://localhost:8983/solr/query?q=*:*&
json.facet={x:'avg(price)'}
And the response will contain a facets section:
[...]
"facets":{
"count":32,
"x":164.10218846797943
}
[...]
If we want to break up the domain into buckets and then calculate a function per bucket, we simply add a nested facet command to the facet parameters. For example (using curl this time):
$curl http://localhost:8983/solr/query -d 'q=*:*&
json.facet={
categories:{
type : terms, // terms facet creates a bucket for each indexed term (or value) in the field
field : cat,
facet:{
x : "avg(price)",
y : "sum(price)"
}
}
}
'
The response will contain the two stats we asked for in each category bucket.
The default sort for a field or terms facet is by bucket count descending. We can optionally sort ascending or descending by any facet function that appears in each bucket. For example, if we wanted to find the top buckets by average price, then we would add sort:"x desc" to the previous facet request:
Subfacets (also called Nested Facets) is a more generalized form of Solr’s current pivot faceting that allows adding additional facets for every bucket produced by a parent facet.
Subfacet advantages over pivot faceting:
Subfacets work with facet functions (statistics), enabling powerful real-time analytics
Can add a subfacet to any facet type (field, query, range)
A subfacet can be of any type (field/terms, query, range)
A given facet can have multiple subfacets
Just like top-level facets, each subfacet can have it’s own configuration (i.e. offset, limit, sort, stats)
Subfacets are part of the new Facet Module, and are naturally expressed in the JSON Facet API. Every facet command is actually a sub-facet since there is an implicit top-level facet bucket (the domain) defined by the documents matching the main query and filters. Simply add a facet section to the parameters of any existing facet command.
For example, a terms facet on the “genre” field looks like:
top_genres:{
type: terms,
field: genre,
limit: 5
}
Now if we wanted to add a subfacet to find the top 4 authors for each genre bucket:
Assume we want to do the following complex faceting request:
Facet on the “genre” field and find the top buckets
For ever “genre” bucket generated above, find the top 7 authors
For ever “genre” bucket, create a bucket of high popularity items (defined by popularity 8 - 10) and call it “highpop”
For ever “highpop” bucket generated above, find the top 5 publishers
In short, this request finds the top authors for each genre and finds the the top publishers for high popularity books in each genre. Using the JSON Facet API, the full request (using curl) would look like the following:
type: terms, // nested terms facet under the nested query facet
field: publisher,
limit: 5
}
}
}
}
}
}
'
An example response would look like the following:
[...]
"facets":{
"top_genres":{
"buckets":[{
"val":"Fantasy",
"count":5432,
"top_authors":{ // these are the top authors in the "Fantasy" genre
"buckets":[{
"val":"Mercedes Lackey",
"count":121},
{
"val":"Piers Anthony",
"count":98}]}},
"highpop":{ // bucket for books in the "Fantasy" genre with popularity between 8 and 10
"count":876
"publishers":{ // top publishers in this bucket (highpop fantasy)
"buckets":[{
"val":"Bantam Books",
"count":346},
{
"val":"Tor",
"count":217}]}},
{
"val":"Science Fiction", // the next genre bucket
"count":4188,
[...]
All the reporting and sorting was done using document count (i.e. number of books). If instead, we wanted to find top authors by total revenue (assuming we had a “sales” field), then we could simply change the author facet from the previous example as follows:
Facet functions and Subfacets are in Solr 5.1 and later, but the syntax used on this page requires Solr 5.3 or later. Download the latest release and give it a spin!
The new facet module has a native JSON Facet API, first-class support for statistics and analytics via facet functions (aggregations), and supports unlimited nesting of facets within other facets via sub-facets.
One can calculate statistics such as averages, number of unique values (distinct values), and percentiles over each facet bucket (groups of documents), and even sort facet buckets by any calculated metrics.
Parameter substitution is now done across the entire query request. It supports default values, multiple levels of indirection, and it even works within the body of a JSON request. This can also be viewed as a powerful form of request templates.
Example:
q=price:[ ${low} TO ${high} ]
&low=100
&high=200
Parameters can also be passed in the params block of a JSON request.
Syntax within the standard lucene/solr query parser for constant score queries quit the general form of ^=<constant_score>. Think of a query boost with ^ replaced with ^=. Example:
There is a new general purpose parallel computing framework for SolrCloud. The Streaming API is (currently) a Java API that can do streaming aggregations (like sum and average) and streaming transformations (like group-by and join).
The admin UI can show segment info such as size, number of docs, and number of deletions for each segment in the index. For the “demo” collection, simply point your browser at http://localhost:8983/solr/#/demo/segments Or click on the “Segments Info” link in the admin UI after you select the core/collection you are interested in.
Many additional configuration items can now be managed via the Config API. This includes managing named components such as requestHandler, queryParser, queryResponseWriter, valueSourceParser, transformer, and queryConverter.
Changes do not directly change solrconfig.xml, but instead are reflected in configoverlay.json which override settings in solrconfig.xml.
Upload config sets to zookeeper with CloudSolrClient
Named config sets (schema.xml, solrconfig,xml, etc) are referenced by name when creating new collections in SolrCloud. These config sets may now be uploaded and downloaded via SolrJ to and from the local filesystem. The following methods were added to CloudSolrClient:
There is a new API to add a jar to a collection’s classpath (as well as update and delete a jar). Components that depend on such a jar should have a new attribute called runtimeLib set to true since a separate classloader is used for these jars.
Caches using the LRUCache implementation can specify a new parameter maxRamMB that will evict based on RAM use rather than number of elements in the cache. Least recently used items are evicted until the RAM use is brought under the limit. RAM use calculations do not currently cover the cache keys, so using this for the query cache and caching large queries can still lead to greater memory use than expected.
Multi-select faceting is a powerful faceting style that allows users to see and select multiple facet constraints (facet values) for a facet. For example, one may want to select multiple price ranges or multiple colors they are interested in.
The new Facet Analytics Module / JSON Facet API now supports multi-select faceting via filter exclusions. A new excludeTags parameter will disregard any top-level filters with matching tags.
Both the older Stats component and the new Facet Analytics Module have added support for HyperLogLog based statistical cardinality estimate. For the JSON Facet API, a new hllfacet function was added as an alternative to the existing faster (but less accurate for high cardinality) unique function. Example:
json.facet={ numProducts : "hll(product_id)" }
See Solr Count Distinct functionality for examples that calculate the number of distinct values in a given field per facet bucket.
Add a new “facet.range.method” parameter to let users choose how to do range faceting between an implementation based on filters (previous algorithm, using “facet.range.method=filter”) or DocValues (“facet.range.method=dv”). Input parameters and output of both methods are the same.
If you have a field value that consists of well formed XML or JSON, you can return those raw values in the appropriate response writer. Example: ?fl=id,name,json_s:[json],xml_s:[xml]
This new SolrCloud feature allows the specification of rules which govern placement of replicas in the cluster. Rules are specified during collection creation and persisted in zookeeper.
The percentile aggregation function was just added to the new Solr Facet Module. This allows one to calculate one or more percentiles for each facet bucket (i.e. each group of documents produced by faceting), and even sort facet buckets by any given percentile.
The percentile aggregation even works with distributed search! The algorithm used is Ted Dunnings “t-digest”, which gives good approximations with relatively little memory consumption.
We can also sort by a percentile statistic. If you request more than one percentile value, the sort will be on the first value in the list requested. Let’s find the top states by 99.9th percentile salary:
JSON strings are normally encapsulated by double quotes. It’s often desirable to use single quotes if for example you are embedding some JSON in another double quoted string in a program.
Allowing trailing commas or extra commas can make it easier to produce JSON that doesn’t throw a parse exception. One use-case is templating JSON. Given the following template,
Large string values can optionally be handled in a streaming fashion a piece at a time. Noggit will only construct a single String object in memory if asked. This allows for stream processing with very little memory overhead.
{
"big_string" : "A very large string... pretend its's 1GB in size... we can process it and send it on without reading it all into memory at once!"
Noggit can also handle multiple JSON values streamed over a single connection and simply catenated together. Primitive values should of course be separated by whitespace to avoid ambiguity.
{first_object:10}
['another array object']['yet another object']
{more:objects}{another:object}
['who knows how many json values will be streamed by the writer...']
Noggit can parse huge JSON messages with minimal overhead.
A single byte of state needed per nested object or array. This is needed to keep track of the type of enclosing entity.
A user can optionally provide an input buffer for Noggit to use when parsing from a Reader, allowing re-use across different parsers and thus lower memory consumption and garbage collection activity.
Streaming values: very large values (such as strings) can be obtained in chunks, thus the whole value never needs to reside in memory at once.
Lucene/Solr trunk (the future 6.0 release) is now on Java8, while version 5.x is still on Java7. Linux and Windows allows one to install a JDK any place in the filesystem, and I use the convention of installing in /opt/jdk7 and /opt/jdk8. Things are a little more difficult on Mac OS-X however, as you can’t chose the install location. Luckily there is a command called java_home to show you where a JDK is installed.
Here’s a snippet from my .profile to help manage working with different java versions:
Terminal window
OS=`uname`
case"$OS"in
CYGWIN*)
OS=cygwin
OPT=c:/opt
;;
*)
OPT=/opt
;;
esac
set-java () {
exportJAVA_HOME="$*"
if [ $OS="cygwin" ]; then
exportPATH="`cygpath$JAVA_HOME/bin`:$PATH"
else
exportPATH="$JAVA_HOME/bin:$PATH"
fi
}
if [ $OS="Darwin" ]; then
JAVA7=`/usr/libexec/java_home-v1.7`
JAVA8=`/usr/libexec/java_home-v1.8`
else
JAVA7=$OPT/jdk7
JAVA8=$OPT/jdk8
fi
set-java$JAVA8
Now, if I switch from working on trunk to working on Lucene 5 or Solr 5, I can easily switch the default JDK for a single terminal via the set-java shell function.
Terminal window
/opt/heliosearch$java-version
javaversion"1.8.0_25"
Java(TM) SE Runtime Environment (build1.8.0_25-b17)
Solr 4.10 and Heliosearch .07 have added a terms query (or terms filter) to more efficiently match many terms in a single field. A large number of terms are often useful for things like access control lists or security filters. Previously, the only way to do this was a large boolean query with many clauses, which has unnecessary overhead when scoring is not needed.
Solr’s implementation uses Lucene’s TermFilter class, as does Elasticsearch’s terms filter.
The Heliosearch terms query implementation has some additional features:
prefix compression including off-heap construction
direct creation of off-heap filter for faster execution and less garbage production
For reference, specifying a filter query (fq) in the normal lucene syntax via a boolean query looks like the following (assumes default boolean operator of OR):
Performance of terms queries is shown relative to using a Boolean query in Solr. For example the last column in the first chart represents a 10 term filter that matches 10,000,000 documents (1 million per term). The request execution time is:
381,342 microseconds with a Solr Boolean Querty
122,119 microseconds with a Solr Terms Query
67,075 microseconds with a Heliosearch Terms Query
Benchmark details:
10M document index
64 bit Java 1.8.0_20 Oracle JDK
Windows 8 64 bit, quad-core Intel i5-3570K @ 3.4GHz
Request time was measured externally and includes the entire request time, including the time for the client to send the request and read the response.
The first performance tests were run multiple times and the amount of garbage produced was recorded.
The Heliosearch off-heap optimizations clearly pay dividends here, resulting in much less heap usage, less garbage production (which will mean less garbage collection work), and a smaller process size.