Skip to content

Blog

Solr 5.3 Features

Here’s an overview of some of the new features in Solr 5.3 Also see Solr Download Links and upcoming Features of the next Solr release.

The JSON Facet API can now change the domain for facet commands, essentially doing a block join and moving from parents to children, or children to parents before calculating the facet data.

For example, if you indexed chapters with pages as nested child documents, then you could map from chapters to pages before faceting by adding the following parameter to the facet command:

domain : { blockChildren : "type:chapter" }

Or if you started with pages, you could map to chapters with

domain : { blockParent : "type:chapter" }

Note that in both cases, we provide the parent filter (how parent documents are defined) of “type:chapter” regardless of which direction we are mapping.

See this Nested Objects tutorial for complete examples of combining faceting and block join / nested documents.

 

Major improvements in performance of the new Facet Module / JSON Facet API. See the facet performance benchmarks for more details and benchmark results.  

Just like the JSON Facet API, pivot facets can how nest other facet types such as range and query facets.

Example:

&facet=true
&facet.range={!tag=r1}price
&f.price.facet.range.start=0
&f.price.facet.range.end=100
&f.price.facet.range.gap=10
&facet.query={!tag=q1}popularity:[8 TO 10]
&facet.pivot={!range=r1 query=q1}category

The equivalent in the JSON Facet API would be:

json.facet={
categories : {
type : terms,
field : category,
facet : {
r1 : {
type : range,
start : 0,
end : 100,
gap : 10
},
q1 : { query : "popularity:[8 TO 10]" }
}
}
}

The MoreLikeThis QParser mlt now supports all options provided by the MLT Handler. The query parser is much more versatile than the handler as it works in cloud mode as well as anywhere a normal query can be specified.

Example (on techproducts index):

q={!mlt qf=name mintf=1 mindf=1}SP2514N

More documentation on the mlt parser can be found in the Solr Ref Guide

The new SchemaRequest Java class in SolrJ can be used to make requests to the Schema API.

Also see the Solr Schema API itself in the ref guide.

Scoring mode for query-time join and block join

Section titled “Scoring mode for query-time join and block join”

Solr’s pseudo-join query parser has a new optional attribute score that can be used specify the scores produced on the resulting documents. It’s value can be min, max,avg,or total.

Query-time join example:

q={!join from=author_id to=id score=total}blog_text:awesome

Block join example:

q={!parent of=type:author score=total}blog_text:awesome

See Nested Objects in Solr for more information on nested documents and block join.

Lucene/Solr query syntax (i.e. Solr’s dialect of the lucene syntax) now supports nested C-style comments.

+cat:electronics /* this is a comment */ +name:HDTV

Smile is a binary data interchange format that is very close to Solr’s own “javabin” (encoded sizes are very close). Adding wt=smile to a request will cause the response to come back in this format.

A second parameter has been added to the field function to select the minimum or maximum value of a multi-valued field with docValues.

Example:

sort=field(my_dv_field,max) asc

In addition to many other improvements in the security framework, Solr now includes an AuthenticationPlugin implementing HTTP Basic Auth that stores credentials securely in ZooKeeper. This is a simple way to require a username and password for anyone accessing Solr’s admin screen or APIs.

See the Basic Authentication Plugin section of the Solr ref guide under the Securing Solr section.

JSON Facet API

Related Pages

Solr 5 has a completely re-written faceted search and analytics module with a structured JSON API to control the faceting and analytics commands. NOTE: Some examples use syntax only supported in later Solr 5 releases, or even Solr 6. Download a recent Solr release or snapshot to try them out.

The structured nature of nested sub-facets are more naturally expressed in a nested structure like JSON rather than the flat structure that normal query parameters provide.

Goals of the new Faceting Module:

  • First class JSON support
  • Easier programmatic construction of complex nested facet commands
  • Support a much more canonical response format that is easier for clients to parse
  • First class analytics support
  • Ability to sort facet buckets by any calculated metric
  • Support a cleaner way to do distributed faceting
  • Support better integration with other search features

Of course if you prefer to use Solr’s existing faceting capabilities, that’s fine too. You can even use both at once if you want!

UPDATE: The JSON Facet API is now part of the JSON Request API, so a complete request may be expressed in JSON.

Some of the ease-of-use enhancements over traditional Solr faceting come from the inherent nested structure of JSON. As an example, here is the faceting command for two different range facets using Solr’s flat legacy API:

&facet=true
&facet.range={!key=age_ranges}age
&f.age.facet.range.start=0
&f.age.facet.range.end=100
&f.age.facet.range.gap=10
&facet.range={!key=price_ranges}price
&f.price.facet.range.start=0
&f.price.facet.range.end=1000
&f.price.facet.range.gap=50

And here is the equivalent faceting command in the new JSON Faceting API:

{
age_ranges: {
type : range
field : age,
start : 0,
end : 100,
gap : 10
}
,
price_ranges: {
type : range
field : price,
start : 0,
end : 1000,
gap : 50
}
}

These aren’t even nested facets, but already one can see how much nicer the JSON API looks. With deeply nested sub-facets and statistics, the clarity of the inherently nested JSON API only grows.

A number of JSON extensions have been implemented to further increase the clarity and ease of constructing a JSON faceting command by hand. For example:

{ // this is a single-line comment, which can help add clarity to large JSON commands
/* traditional C-style comments are also supported */
x : "avg(price)" , // Simple strings can occur unquoted
y : 'unique(manu)' // Strings can also use single quotes (easier to embed in another String)
}

Nicely indented JSON is very easy to understand. If you get a large piece of non-indented JSON somehow, and are trying to make sense of it, you can cut and paste into one of the online validators: http://jsonlint.com http://jsonformatter.curiousconcept.com Both of these validators will indent your JSON, even when it contains extensions unsupported by them (such as comments or bare strings).

 

There are two types of facets, one that breaks up the domain into multiple buckets, and aggregations / facet functions that provide information about the set of documents belonging to each bucket.

Faceting can be nested! Any bucket produced by faceting can further be broken down into multiple buckets by a sub-facet.

Statistics are now fully integrated into faceting. Since we start off with a single facet bucket with a domain defined by the main query and filters, we can even ask for statistics for this top level bucket, before breaking up into further buckets via faceting. Example:

json.facet={
x : "avg(price)", // the average of the price field will appear under "x"
y : "unique(manufacturer)" // the number of unique manufacturers will appear under "y"
}

See facet functions for a complete list of the available aggregation functions.

The general form of the JSON facet commands are: <facet_name> : { <facet_type> : <facet_parameter(s)> } Example: top_authors : { terms : { field : authors, limit : 5 } }

After Solr 5.2, a flatter structure with a “type” field may also be used: <facet_name> : { "type" : <facet_type> , <other_facet_parameter(s)> } Example: top_authors : { type : terms, field : authors, limit : 5 }

The results will appear in the response under the facet name specified. Facet commands are specified using json.facet request parameters.

To test out different facet requests by hand, it’s easiest to use “curl” from the command line. Example:

$ curl http://localhost:8983/solr/query -d 'q=*:*&rows=0&
json.facet={
categories:{
type : terms,
field : cat,
sort : { x : desc},
facet:{
x : "avg(price)",
y : "sum(price)"
}
}
}
'

 

The terms facet, or field facet, produces buckets from the unique values of a field. The field needs to be indexed or have docValues.

The simplest form of the terms facet

{
top_genres : { terms : genre_field }
}

An expanded form allows for more parameters:

{
top_genres : {
type : terms,
field : genre_field,
limit : 3,
mincount : 2
}
}

Example response:

"top_genres":{
"buckets":[
{
"val":"Science Fiction",
"count":143},
{
"val":"Fantasy",
"count":122},
{
"val":"Biography",
"count":28}
]
}

Parameters:

  • field - The field name to facet over.

  • offset - Used for paging, this skips the first N buckets. Defaults to 0.

  • limit - Limits the number of buckets returned. Defaults to 10.

  • mincount - Only return buckets with a count of at least this number. Defaults to 1.

  • sort - Specifies how to sort the buckets produced. “count” specifies document count, “index” sorts by the index (natural) order of the bucket value. One can also sort by any facet function / statistic that occurs in the bucket. The default is “count desc”. This parameter may also be specified in JSON like sort:{count:desc}. The sort order may either be “asc” or “desc”

  • missing - A boolean that specifies if a special “missing” bucket should be returned that is defined by documents without a value in the field. Defaults to false.

  • numBuckets - A boolean. If true, adds “numBuckets” to the response, an integer representing the number of buckets for the facet (as opposed to the number of buckets returned). Defaults to false.

  • allBuckets - A boolean. If true, adds an “allBuckets” bucket to the response, representing the union of all of the buckets. For multi-valued fields, this is different than a bucket for all of the documents in the domain since a single document can belong to multiple buckets. Defaults to false.

  • prefix - Only produce buckets for terms starting with the specified prefix.

  • method - Provides an execution hint for how to facet the field.

    • method:uif - Stands for UninvertedField, a method of faceting indexed, multi-valued fields using top-level data structures that optimize for performance over NRT capabilities.
    • method:dv - Stands for DocValues, a method of faceting indexed, multi-valued fields using per-segment data structures. This method mirrors faceting on real docValues fields but works by building on-heap docValues on the fly from the index when docValues aren’t available. This method is better for a quickly changing index.
    • method:stream - This method creates each individual facet bucket (including any sub-facets) on-the-fly while streaming the response back to the requester. Currently only supports sorting by index order.

 

The query facet produces a single bucket that matches the specified query.

An example of the simplest form of the query facet

{
high_popularity : { query : "popularity:[8 TO 10]" }
}

An expanded form allows for more parameters (or sub-facets / facet functions):

{
high_popularity : {
type : query,
q : "popularity:[8 TO 10]",
facet : { average_price : "avg(price)" }
}
}

Example response:

"high_popularity" : {
"count" : 147,
"average_price" : 74.25
}

 

The range facet produces multiple range buckets over numeric fields or date fields.

Range facet example:

{
prices : {
type : range,
field : price,
start : 0,
end : 100,
gap : 20
}
}

Example response:

"prices":{
"buckets":[
{
"val":0.0, // the bucket value represents the start of each range. This bucket covers 0-20
"count":5},
{
"val":20.0,
"count":3},
{
"val":40.0,
"count":2},
{
"val":60.0,
"count":1},
{
"val":80.0,
"count":1}
]
}

To ease migration, these parameter names, values, and semantics were taken directly from the old-style (non JSON) Solr range faceting.

Parameters:

  • field - The numeric field or date field to produce range buckets from

  • mincount - Minimum document count for the bucket to be included in the response. Defaults to 0.

  • start - Lower bound of the ranges

  • end - Upper bound of the ranges

  • gap - Size of each range bucket produced

  • hardend - A boolean, which if true means that the last bucket will end at “end” even if it is less than “gap” wide. If false, the last bucket will be “gap” wide, which may extend past “end”.

  • other - This param indicates that in addition to the counts for each range constraint between facet.range.start and facet.range.end, counts should also be computed for…

  • "before" all records with field values lower then lower bound of the first range

  • "after" all records with field values greater then the upper bound of the last range

  • "between" all records with field values between the start and end bounds of all ranges

  • "none" compute none of this information

  • "all" shortcut for before, between, and after

  • include - By default, the ranges used to compute range faceting between facet.range.start and facet.range.end are inclusive of their lower bounds and exclusive of the upper bounds. The “before” range is exclusive and the “after” range is inclusive. This default, equivalent to lower below, will not result in double counting at the boundaries. This behavior can be modified by the facet.range.include param, which can be any combination of the following options…

  • "lower" all gap based ranges include their lower bound

  • "upper" all gap based ranges include their upper bound

  • "edge" the first and last gap ranges include their edge bounds (ie: lower for the first one, upper for the last one) even if the corresponding upper/lower option is not specified

  • "outer" the “before” and “after” ranges will be inclusive of their bounds, even if the first or last ranges already include those boundaries.

  • "all" shorthand for lower, upper, edge, outer

Parameters that all faceting methods have in common include

Solr Facet Functions and Analytics

Traditional faceted search (also called guided navigation) involves counting search results that belong to categories (also called facet constraints). The new facet functions in Solr extends normal faceting by allowing additional aggregations on document fields themselves. Combined with the new Sub-facet feature, this provides powerful new realtime analytics capabilities. Also see the page about the new JSON Facet API.

Faceting involves breaking up the domain into multiple buckets and providing information about each bucket. There are multiple aggregation functions / statistics that can be used:

Aggregation Example Effect
sum sum(sales) summation of numeric values
avg avg(popularity) average of numeric values
sumsq sumsq(rent) sum of squares
min min(salary) minimum value
max max(mul(price,popularity)) maximum value
unique unique(state) number of unique values (count distinct)
hll hll(state) number of unique values using the HyperLogLog algorithm
percentile percentile(salary,50,75,99,99.9) calculates percentiles
stddev stddev(salary) calculates standard deviation (Solr6.6+)
variance variance(salary) calculates variance (Solr 6.6+)

  Numeric aggregation functions such as avg can be on any numeric field, or on another function of multiple numeric fields.

See Count Distinct in Solr for more information on distributed cardinality estimation / calcDistinct.

 

The faceting domain starts with the set of documents that match the main query and filters. We can ask for statistics over this whole set of documents:

http://localhost:8983/solr/query?q=*:*&
json.facet={x:'avg(price)'}

And the response will contain a facets section:

[...]
"facets":{
"count":32,
"x":164.10218846797943
}
[...]

  If we want to break up the domain into buckets and then calculate a function per bucket, we simply add a nested facet command to the facet parameters. For example (using curl this time):

$ curl http://localhost:8983/solr/query -d 'q=*:*&
json.facet={
categories:{
type : terms, // terms facet creates a bucket for each indexed term (or value) in the field
field : cat,
facet:{
x : "avg(price)",
y : "sum(price)"
}
}
}
'

The response will contain the two stats we asked for in each category bucket.

[...]
"facets":{
"count":32,
"categories":{
"buckets":[
{
"val":"electronics",
"count":12,
"x":231.02666823069254,
"y":2772.3200187683105
},
{
"val":"memory",
"count":3,
"x":86.66333262125652,
"y":259.98999786376953
},
[...]

 

The default sort for a field or terms facet is by bucket count descending. We can optionally sort ascending or descending by any facet function that appears in each bucket. For example, if we wanted to find the top buckets by average price, then we would add sort:"x desc" to the previous facet request:

$ curl http://localhost:8983/solr/query -d 'q=*:*&
json.facet={
categories:{
type : terms,
field : cat,
sort : "x desc", // can also use sort:{x:desc}
facet:{
x : "avg(price)",
y : "sum(price)"
}
}
}
'

 

Facet functions and Subfacets are currently only in Solr 5.1. Download the latest release and give it a spin!

Solr Subfacets

Subfacets (also called Nested Facets) is a more generalized form of Solr’s current pivot faceting that allows adding additional facets for every bucket produced by a parent facet.

Subfacet advantages over pivot faceting:

  • Subfacets work with facet functions (statistics), enabling powerful real-time analytics
  • Can add a subfacet to any facet type (field, query, range)
  • A subfacet can be of any type (field/terms, query, range)
  • A given facet can have multiple subfacets
  • Just like top-level facets, each subfacet can have it’s own configuration (i.e. offset, limit, sort, stats)

Subfacets are part of the new Facet Module, and are naturally expressed in the JSON Facet API. Every facet command is actually a sub-facet since there is an implicit top-level facet bucket (the domain) defined by the documents matching the main query and filters. Simply add a facet section to the parameters of any existing facet command.

For example, a terms facet on the “genre” field looks like:

top_genres:{
type: terms,
field: genre,
limit: 5
}

Now if we wanted to add a subfacet to find the top 4 authors for each genre bucket:

top_genres:{
type: terms,
field: genre,
limit: 5,
facet:{
top_authors:{
type: terms,
field: author,
limit: 4
}
}
}

Assume we want to do the following complex faceting request:

  • Facet on the “genre” field and find the top buckets
  • For ever “genre” bucket generated above, find the top 7 authors
  • For ever “genre” bucket, create a bucket of high popularity items (defined by popularity 8 - 10) and call it “highpop”
  • For ever “highpop” bucket generated above, find the top 5 publishers

In short, this request finds the top authors for each genre and finds the the top publishers for high popularity books in each genre. Using the JSON Facet API, the full request (using curl) would look like the following:

$ curl http://localhost:8983/solr/query -d 'q=*:*&
json.facet=
{
top_genres:{
type: terms,
field: genre,
facet:{
top_authors: {
type : terms, // nested terms facet
field: author,
limit: 7
},
highpop:{
type : query, // nested query facet
q: "popularity:[8 TO 10]", // lucene query string
facet:{
publishers:{
type: terms, // nested terms facet under the nested query facet
field: publisher,
limit: 5
}
}
}
}
}
}
'

An example response would look like the following:

[...]
"facets":{
"top_genres":{
"buckets":[{
"val":"Fantasy",
"count":5432,
"top_authors":{ // these are the top authors in the "Fantasy" genre
"buckets":[{
"val":"Mercedes Lackey",
"count":121},
{
"val":"Piers Anthony",
"count":98}]}},
"highpop":{ // bucket for books in the "Fantasy" genre with popularity between 8 and 10
"count":876
"publishers":{ // top publishers in this bucket (highpop fantasy)
"buckets":[{
"val":"Bantam Books",
"count":346},
{
"val":"Tor",
"count":217}]}},
{
"val":"Science Fiction", // the next genre bucket
"count":4188,
[...]

  All the reporting and sorting was done using document count (i.e. number of books). If instead, we wanted to find top authors by total revenue (assuming we had a “sales” field), then we could simply change the author facet from the previous example as follows:

top_authors:{
type: terms,
field: author,
limit: 7,
sort: "revenue desc",
facet:{
revenue: "sum(sales)"
}
}

 

Facet functions and Subfacets are in Solr 5.1 and later, but the syntax used on this page requires Solr 5.3 or later. Download the latest release and give it a spin!

Solr 5.1 Features

Solr 5.1 has been released! Here’s an overview of how to use some of the new features.

Also see Solr download links and upcoming features of the next Solr release.

The new facet module has a native JSON Facet API, first-class support for statistics and analytics via facet functions (aggregations), and supports unlimited nesting of facets within other facets via sub-facets.

One can calculate statistics such as averages, number of unique values (distinct values), and percentiles over each facet bucket (groups of documents), and even sort facet buckets by any calculated metrics.

A JSON Request API that allows passing a full Solr query request in JSON.

Example:

curl http://localhost:8983/solr/query -d '
{
query : "*:*",
filter : [
"author:brandon",
"genre_s:fantasy"
],
offset : 0,
limit : 5,
fields : ["title","author"], // we could also use the string form "title,author"
sort : "sequence_i desc",
facet : { // the JSON Facet API is nicely integrated as well
avg_price : "avg(price)",
median_price : "percentile(price,50)",
top_authors : {terms : author}
}
}'

Parameter substitution is now done across the entire query request. It supports default values, multiple levels of indirection, and it even works within the body of a JSON request. This can also be viewed as a powerful form of request templates.

Example:

q=price:[ ${low} TO ${high} ]
&low=100
&high=200

Parameters can also be passed in the params block of a JSON request.

Syntax within the standard lucene/solr query parser for constant score queries quit the general form of ^=<constant_score>. Think of a query boost with ^ replaced with ^=. Example:

q=(color:blue color:green)^=2.0 text:shoes

There is a new general purpose parallel computing framework for SolrCloud. The Streaming API is (currently) a Java API that can do streaming aggregations (like sum and average) and streaming transformations (like group-by and join).

The admin UI can show segment info such as size, number of docs, and number of deletions for each segment in the index. For the “demo” collection, simply point your browser at http://localhost:8983/solr/#/demo/segments Or click on the “Segments Info” link in the admin UI after you select the core/collection you are interested in. segments_info

The bulk schema API how has the ability to replace or remove fields, fieldTypes, dynamic fields, and copy fields.

Example of adding a field (this was already possible):

curl http://localhost:8983/solr/demo/schema -d '
{
"add-field":{
"name" : "powerLevel",
"type" : "int",
"indexed" : true,
"stored" : true
}
}'

Now we can replace the field definition:

curl http://localhost:8983/solr/demo/schema -d '
{
"replace-field":{
"name" : "powerLevel",
"type" : "int",
"indexed" : false,
"stored" : true
}
}'

We can verify that Solr now has the updated field definition with

curl http://localhost:8983/solr/demo/schema/fields/powerLevel

And solr returns:

"field":{
"name":"powerLevel",
"type":"int",
"indexed":false,
"stored":true}

And lastly, we can delete the field definition with

curl http://localhost:8983/solr/demo/schema -d '
{
"delete-field":{ "name" : "powerLevel" }
}'

Solr can now execute a two dimensional facet on RPT field types (Spatial Recursive Prefix Tree).

Parameters Example:

q=*:*
&facet=true
&facet.heatmap=location_rpt
&facet.heatmap.geom=["-180 -90" TO "180 90"]
&facet.heatmap.gridLevel=6
&facet.heatmap.distErrPct=0.15
&facet.heatmap.format=ints2D

The facet.heatmap.format=ints2D parameter causes a 2D array of counts to be returned:

{
"counts_ints2D":[[4, 0, 1, 3, ....],[2, 0, 1, 2, ...],...]
}

If facet.heatmap.format=png is passed instead, a basic base64-encoded PNG (picture) will be returned of the heatmap grid.

There is now an explicit API in SolrJ to use Real-time Get

HttpSolrClient client = new HttpSolrClient("http://localhost:8983/solr/demo");
SolrDocument sdoc = client.getById("book1");
System.out.println("I found book " + sdoc);
client.close(); // shut down the client when we are done

StatsComponent Enable/disable individual stats

Section titled “StatsComponent Enable/disable individual stats”

Localparams may now be used to selectively enable or disable specific stats in the StatsComponent. Example: stats.field={!min=true max=true}field_name

Both the new facet module and the stats component gained support for percentiles.

json.facet={ median_age : "percentile(age,50)" }
stats.field={!percentiles='50'}age

Many additional configuration items can now be managed via the Config API. This includes managing named components such as requestHandler, queryParser, queryResponseWriter, valueSourceParser, transformer, and queryConverter.

Changes do not directly change solrconfig.xml, but instead are reflected in configoverlay.json which override settings in solrconfig.xml.

Upload config sets to zookeeper with CloudSolrClient

Section titled “Upload config sets to zookeeper with CloudSolrClient”

Named config sets (schema.xml, solrconfig,xml, etc) are referenced by name when creating new collections in SolrCloud. These config sets may now be uploaded and downloaded via SolrJ to and from the local filesystem. The following methods were added to CloudSolrClient:

public void uploadConfig(Path configPath, String configName);
public void downloadConfig(String configName, Path downloadPath);

There is a new API to add a jar to a collection’s classpath (as well as update and delete a jar). Components that depend on such a jar should have a new attribute called runtimeLib set to true since a separate classloader is used for these jars.

Example of uploading a jar:

curl http://localhost:8983/solr/demo/config -d '{
"add-runtimelib" : {"name": "jarname" , "version":2 }
}'

Example registering a new value source parser using a class in the jar:

curl http://localhost:8983/solr/demo/config -d '{
"create-valuesourceparser" : {
"name": "nvl",
"runtimeLib" : true,
"class" : "solr.org.apache.solr.search.function.NvlValueSourceParser ,
"nvlFloatValue" : 0.0
}
}'

Solr 5.2 Features

Here’s an overview of some of the new features in Solr 5.2 Also see Solr download links and upcoming features of the next Solr release.

Caches using the LRUCache implementation can specify a new parameter maxRamMB that will evict based on RAM use rather than number of elements in the cache. Least recently used items are evicted until the RAM use is brought under the limit. RAM use calculations do not currently cover the cache keys, so using this for the query cache and caching large queries can still lead to greater memory use than expected.

To make a backup, we can send a request to the replication handler:

curl -XPOST "http://localhost:8983/solr/demo/replication?command=backup&name=my_backup100"

This will create a backup of the index in Solr’s data directory (this can be changed via the location parameter) named snapshot.my_backup100

This index snapshot can later be restored with the following command:

curl -XPOST "http://localhost:8983/solr/demo/replication?command=restore&name=my_backup100"

Flatter request structure for the JSON Facet API

Section titled “Flatter request structure for the JSON Facet API”

Here’s an example of a terms facet in Solr 5.1:

top_authors : { terms : {
field : author,
limit : 5,
}}

In the Solr 5.2 JSON Facet API, the “type” can optionally be specified in the same object as the facet arguments:

top_authors : {
type : terms,
field : author,
limit : 5
}

Range facets now support the mincount parameter to screen out range facet buckets that don’t meet a minimum document count.

prices:{
type:range,
field:price,
mincount:1,
start:0, end:100, gap:10
}

The unique facet function now works on numeric and date fields. Example:

json.facet={
num_codes : "unique(error_code)"
}

Multi-select faceting is a powerful faceting style that allows users to see and select multiple facet constraints (facet values) for a facet. For example, one may want to select multiple price ranges or multiple colors they are interested in.

The new Facet Analytics Module / JSON Facet API now supports multi-select faceting via filter exclusions. A new excludeTags parameter will disregard any top-level filters with matching tags.

Here’s a Multi-Select Faceting Example, using the JSON Facet API.

Both the older Stats component and the new Facet Analytics Module have added support for HyperLogLog based statistical cardinality estimate. For the JSON Facet API, a new hll facet function was added as an alternative to the existing faster (but less accurate for high cardinality) unique function. Example:

json.facet={ numProducts : "hll(product_id)" }

See Solr Count Distinct functionality for examples that calculate the number of distinct values in a given field per facet bucket.

“facet.range.method” (traditional query-parameter API)

Section titled ““facet.range.method” (traditional query-parameter API)”

Add a new “facet.range.method” parameter to let users choose how to do range faceting between an implementation based on filters (previous algorithm, using “facet.range.method=filter”) or DocValues (“facet.range.method=dv”). Input parameters and output of both methods are the same.

If you have a field value that consists of well formed XML or JSON, you can return those raw values in the appropriate response writer. Example: ?fl=id,name,json_s:[json],xml_s:[xml]

This new SolrCloud feature allows the specification of rules which govern placement of replicas in the cluster. Rules are specified during collection creation and persisted in zookeeper.

See the blog post from LucidWorks for further details and examples.

Solr Streaming Expressions adds an expression based interface to the Streaming API added in Solr 5.1.

Some examples from include

// merge two distinct searches together on common fields
merge(
search(collection1, q="id:(0 3 4)", fl="id,a_s,a_i,a_f", sort="a_f asc, a_s asc"),
search(collection2, q="id:(1 2)", fl="id,a_s,a_i,a_f", sort="a_f asc, a_s asc"),
on="a_f asc, a_s asc")
// find top 20 unique records of a search
top(
n=20,
unique(
search(collection1, q=*:*, fl="id,a_s,a_i,a_f", sort="a_f desc"),
over="a_f desc"),
sort="a_f desc")

See the Solr Reference Guide for more documentation.

An authentication framework and Kerberose authentication module. See the Security section of the Solr Reference Guide.

Percentiles for Solr Faceting

The percentile aggregation function was just added to the new Solr Facet Module. This allows one to calculate one or more percentiles for each facet bucket (i.e. each group of documents produced by faceting), and even sort facet buckets by any given percentile.

The percentile aggregation even works with distributed search! The algorithm used is Ted Dunnings “t-digest”, which gives good approximations with relatively little memory consumption.

NOTE: requires Solr 5.3 or later.

First, let’s start Solr and create a “demo” collection.

$ bin/solr start
$ bin/solr create -c demo
# HINT: use "bin/solr stop -all" when you're finished.

Now, lets index some salary survey data in CSV format, using dynamic fields:

$ curl http://localhost:8983/solr/demo/update?commitWithin=5000 -H 'Content-type:text/csv' -d '
id,gender_s,loc_s,year_i,job_s,salary_d
mark,M,NJ,2011,clerk,21250
john,M,NY,2011,engineer,42500
mary,F,CT,2015,manager,87299
alice,F,NJ,2013,dentist,75000
mike,M,NY,2012,sales,59500
nancy,F,CT,2014,engineer,110000
greg,M,NJ,2012,manager,74000
cindy,F,NJ,2012,engineer,81000
janet,F,NJ,2015,clerk,30150
joe,M,NY,2014,dentist,74000
luke,M,CT,2015,dentist,78000
zoe,F,NY,2013,manager,89500
eli,M,CT,2011,sales,66000
anna,F,CT,2012,sales,59500
evan,M,NY,2014,clerk,2920
'

Now we can use Solr’s analytics / facet functions to slice and dice our data!

Let’s say we want the 25%, 50%, and 75% percentile salaries across all our jobs:

$ curl http://localhost:8983/solr/demo/query -d 'q=*:*&json.facet={salary_percentiles:"percentile(salary_d,25,50,75)"}'

And at the end of our response, we’ll get our facet results:

[...]
"facets" : {
"count" : 15,
"salary_percentiles" : [51000.0, 74000.0, 79500.0]
}
}

  We can add in other statistics such as the average salary, the number of different jobs, and the number of different states in our salary survey:

$ curl http://localhost:8983/solr/demo/query -d 'q=*:*&
json.facet={
average_salary : "avg(salary_d)",
num_jobs : "unique(job_s)",
num_states : "unique(loc_s)",
salary_percentiles : "percentile(salary_d,25,50,75)"
}'
"facets":{
"count":15,
"average_salary":63374.6,
"num_jobs":5,
"num_states":3,
"salary_percentiles":[51000.0,74000.0,79500.0]
}

  Now let’s take a look at median salary broken out by gender:

$ curl http://localhost:8983/solr/demo/query -d 'q=*:*&
json.facet={
by_gender:{
type:terms
field:gender_s,
facet:{
median_salary:"percentile(salary_d,50)"
}
}
}'
"facets":{
"count":15,
"by_gender":{
"buckets":[
{
"val":"M",
"count":8,
"median_salary":62750.0
},
{
"val":"F",
"count":7,
"median_salary":81000.0
}
]
}
}

  We can also sort by a percentile statistic. If you request more than one percentile value, the sort will be on the first value in the list requested. Let’s find the top states by 99.9th percentile salary:

$ curl http://localhost:8983/solr/demo/query -d 'q=*:*&
json.facet={
rich_states:{
type : terms,
field : loc_s,
sort : {sal:desc}, // specifying the sort as a string, like sort:"sal desc" will also work
facet : {
sal : "percentile(salary_d,99.9)"
}
}
}'
"facets":{
"count":15,
"rich_states":{
"buckets":[{
"val":"CT",
"count":5,
"sal":109909.19600000001},
{
"val":"NY",
"count":5,
"sal":89438.00000000001},
{
"val":"NJ",
"count":5,
"sal":80976.0}]}}

  We can get even more interesting by nesting facets. How about finding the highest earning occupation (99.9th percentile) for every state?

$ curl http://localhost:8983/solr/demo/query -d 'q=*:*&
json.facet={
states:{
type:terms,
field:loc_s,
facet:{
top_jobs:{ // nested terms facet
type : terms,
field : job_s,
sort : "sal desc", // sort will be on first percentile (99.9)
limit : 1, // only show top occupation
facet:{
sal : "percentile(salary_d,99.9,50,10)"
}
}
} // end facet block for the loc_s field
}
}'

The response has been omitted since we don’t have enough data for it to be interesting.

  We can also show how median salary has changed over time for each individual state:

$ curl http://localhost:8983/solr/demo/query -d 'q=*:*&
json.facet={
states:{
type:terms,
field:loc_s,
facet:{
over_time:{ // nested range facet
type : range,
field : year_i,
start : 2011,
end : 2015,
gap : 1,
facet:{
median_salary : "percentile(salary_d,50)"
}
}
} // end facet block for the loc_s field
}
}'
"facets":{
"count":15,
"states":{
"buckets":[{
"val":"CT",
"count":5,
"over_time":{
"buckets":[{
"val":2011,
"count":1,
"median_salary":66000.0},
{
"val":2012,
"count":1,
"median_salary":59500.0},
{
"val":2013,
"count":0},
{
"val":2014,
"count":1,
"median_salary":110000.0}]}},
{
"val":"NJ",
"count":5,
"over_time":{
"buckets":[{
"val":2011,
"count":1,
"median_salary":21250.0},
{
"val":2012,
"count":2,
"median_salary":77500.0},
{
"val":2013,
"count":1,
"median_salary":75000.0},
{
"val":2014,
"count":0}]}},
{
"val":"NY",
"count":5,
"over_time":{
"buckets":[{
"val":2011,
"count":1,
"median_salary":42500.0},
{
"val":2012,
"count":1,
"median_salary":59500.0},
{
"val":2013,
"count":1,
"median_salary":89500.0},
{
"val":2014,
"count":2,
"median_salary":38460.0}]}}]}}

Noggit, the JSON Streaming Parser

Noggit is the world’s fastest streaming JSON parser for Java.

Section titled “Noggit is the world’s fastest streaming JSON parser for Java.”

Noggit is the streaming JSON parser used in Solr. It lives here on github.

Noggit supports a number of extensions to the JSON grammar. All of these extensions are optional and may be disabled.

{ // This is a single line comment
# This is also a single line comment
/* This is a multi-line
* C-style comment.
*/
}
{
first : Yonik,
last : Seeley
}

JSON strings are normally encapsulated by double quotes. It’s often desirable to use single quotes if for example you are embedding some JSON in another double quoted string in a program.

['how', 'now', 'brown', 'cow']

Sometimes one may not know exactly what characters need to be backslash escaped. It can be useful to accept this without throwing an exception.

'This is just a " string'

Allowing trailing commas or extra commas can make it easier to produce JSON that doesn’t throw a parse exception. One use-case is templating JSON. Given the following template,

{
filters:["instock:true", ${FILT1}]
} # Note: templating is not part of JSON or Noggit... but may happen before parsing.

If FILT1 is not defined and is replaced with empty space, this results in the following JSON:

{
filters:["instock:true", ] // this will be parsed as filters:["instock:true"]
}

Noggit ignores all extra commas, not just trailing commas:

[
[,] // equivalent to []
, {,} // equivalent to {}
, [,,3,,,6,,] // equivalent to [3,6]
]

Large string values can optionally be handled in a streaming fashion a piece at a time. Noggit will only construct a single String object in memory if asked. This allows for stream processing with very little memory overhead.

{
"big_string" : "A very large string... pretend its's 1GB in size... we can process it and send it on without reading it all into memory at once!"
}

Noggit can handle huge values that are JSON compliant but may be too large to be parsed into a Java primitive.

{
"big_int" : 1234567890987654321334325343534535342325786237862578625725867258672356711107,
"big_float" : 112412133377778226524562431234215423.23421434645743234564758453322342,
"big_sci" : 2.342669039282149050282364845982748592e-94321
}

Noggit can also handle multiple JSON values streamed over a single connection and simply catenated together. Primitive values should of course be separated by whitespace to avoid ambiguity.

{first_object:10}
['another array object']['yet another object']
{more:objects}{another:object}
['who knows how many json values will be streamed by the writer...']
42
"is this the end?"

Noggit can parse huge JSON messages with minimal overhead.

  • A single byte of state needed per nested object or array. This is needed to keep track of the type of enclosing entity.
  • A user can optionally provide an input buffer for Noggit to use when parsing from a Reader, allowing re-use across different parsers and thus lower memory consumption and garbage collection activity.
  • Streaming values: very large values (such as strings) can be obtained in chunks, thus the whole value never needs to reside in memory at once.

Switching between Java7 and Java8 in Lucene/Solr

Lucene/Solr trunk (the future 6.0 release) is now on Java8, while version 5.x is still on Java7. Linux and Windows allows one to install a JDK any place in the filesystem, and I use the convention of installing in /opt/jdk7 and /opt/jdk8. Things are a little more difficult on Mac OS-X however, as you can’t chose the install location. Luckily there is a command called java_home to show you where a JDK is installed.

Here’s a snippet from my .profile to help manage working with different java versions:

Terminal window
OS=`uname`
case "$OS" in
CYGWIN*)
OS=cygwin
OPT=c:/opt
;;
*)
OPT=/opt
;;
esac
set-java () {
export JAVA_HOME="$*"
if [ $OS = "cygwin" ]; then
export PATH="`cygpath $JAVA_HOME/bin`:$PATH"
else
export PATH="$JAVA_HOME/bin:$PATH"
fi
}
if [ $OS = "Darwin" ]; then
JAVA7=`/usr/libexec/java_home -v 1.7`
JAVA8=`/usr/libexec/java_home -v 1.8`
else
JAVA7=$OPT/jdk7
JAVA8=$OPT/jdk8
fi
set-java $JAVA8

Now, if I switch from working on trunk to working on Lucene 5 or Solr 5, I can easily switch the default JDK for a single terminal via the set-java shell function.

Terminal window
/opt/heliosearch$ java -version
java version "1.8.0_25"
Java(TM) SE Runtime Environment (build 1.8.0_25-b17)
Java HotSpot(TM) 64-Bit Server VM (build 25.25-b02, mixed mode)
/opt/heliosearch$ set-java $JAVA7
/opt/heliosearch$ java -version
java version "1.7.0_71"
Java(TM) SE Runtime Environment (build 1.7.0_71-b14)
Java HotSpot(TM) 64-Bit Server VM (build 24.71-b01, mixed mode)
/opt/heliosearch$

Solr Terms Query for matching many terms

Solr 4.10 and Heliosearch .07 have added a terms query (or terms filter) to more efficiently match many terms in a single field. A large number of terms are often useful for things like access control lists or security filters. Previously, the only way to do this was a large boolean query with many clauses, which has unnecessary overhead when scoring is not needed.

Solr’s implementation uses Lucene’s TermFilter class, as does Elasticsearch’s terms filter.

The Heliosearch terms query implementation has some additional features:

  • prefix compression including off-heap construction
  • direct creation of off-heap filter for faster execution and less garbage production
  • native code bit-setting
  • ability to skip sorting the terms if desired

For reference, specifying a filter query (fq) in the normal lucene syntax via a boolean query looks like the following (assumes default boolean operator of OR):

fq=id:doc334 id:doc125 id:doc777 id:doc321 id:doc253

or in a more compact form, like:

fq=id:(doc334 doc125 doc777 doc321 doc253)

Be aware that going over the limit of 1024 terms in Solr will cause an exception by default. Heliosearch has no such limit.

The corresponding new terms query in both Solr and Heliosearch is:

fq={!terms f=id}doc334,doc125,doc777,doc321,doc253

Performance of terms queries is shown relative to using a Boolean query in Solr. For example the last column in the first chart represents a 10 term filter that matches 10,000,000 documents (1 million per term). The request execution time is:

  • 381,342 microseconds with a Solr Boolean Querty
  • 122,119 microseconds with a Solr Terms Query
  • 67,075 microseconds with a Heliosearch Terms Query

Benchmark details:

  • 10M document index
  • 64 bit Java 1.8.0_20 Oracle JDK
  • Windows 8 64 bit, quad-core Intel i5-3570K @ 3.4GHz
  • Request time was measured externally and includes the entire request time, including the time for the client to send the request and read the response.
  • Solr versions: Apache Solr 4.10.0, Heliosearch 0.07 (based on Solr 4.10)

  The first set of tests consist of 10 term queries that match various number of documents.: terms_perf_10

  The next set of tests deal with 100 term queries that match various number of documents: terms_perf_100

  And finally the last test deals with term queries on the id field (i.e. each term matches a single document): terms_perf_ids

The first performance tests were run multiple times and the amount of garbage produced was recorded. terms_perf_memory

The Heliosearch off-heap optimizations clearly pay dividends here, resulting in much less heap usage, less garbage production (which will mean less garbage collection work), and a smaller process size.