From 0cfb204c04e9ee2eda8b7bcd7ba3bd9fb5bbc1e3 Mon Sep 17 00:00:00 2001 From: Javier Goizueta Date: Tue, 9 Jan 2018 14:49:33 +0100 Subject: [PATCH 1/8] Add MapConfig extension for aggregation --- docs/MapConfig-Aggregation-extension.md | 59 +++++++++++++++++++++++++ 1 file changed, 59 insertions(+) create mode 100644 docs/MapConfig-Aggregation-extension.md diff --git a/docs/MapConfig-Aggregation-extension.md b/docs/MapConfig-Aggregation-extension.md new file mode 100644 index 00000000..ffcd5144 --- /dev/null +++ b/docs/MapConfig-Aggregation-extension.md @@ -0,0 +1,59 @@ +# 1. Purpose + +This specification describes an extension for +[MapConfig 1.7.0](https://github.com/CartoDB/Windshaft/blob/master/doc/MapConfig-1.7.0.md) version. + + +# 2. Changes over specification + +This extension introduces a new layer options for aggregated data tile generation. + +## 2.1 Aggregation options + +The layer options attribute is extended with a new optional `aggregation` attribute. +The value of this attribute can be `false` to explicitly disable aggregation for the layer. + +```javascript +{ + aggregation: { + + // OPTIONAL + // string, defines the placement of aggregated geometries. Can be one of: + // * "point-sample", the default places geometries at a sample point (one of the aggregated geometries) + // * "point-grid" places geometries at the center of the aggregation grid cells + // * "centroid" places geometriea at the average position of the aggregated points + // See https://github.com/CartoDB/Windshaft-cartodb/blob/master/docs/... for more details + placement: "point-sample", + + // OPTIONAL + // object, defines the columns of the aggregated datasets. Each property corresponds to a columns name and + // should contain an object with two properties: "aggregate_function" (one of "sum", "max", "min", "avg", "mode" or "count"), + // and "aggregated_column" (the name of a column of the original layer query or "*") + // A column defined as `"_cdb_features_count": {"aggregate_function": "count", aggregated_column: "*"}` + // is always generated in addition to the defined columns. + // The column names `cartodb_id`, `the_geom`, `the_geom_webmercator` and `_cdb_feature_count` cannot be used + // for aggregated columns, as they correspond to columns always present in the result. + columns: { + "aggregated_column_1": { + "aggregate_function": "sum", + "aggregated_column": "original_column_1" + } + }, + + // OPTIONAL + // Number, defines the cell-size of the spatial aggregation grid as a pixel resolution power of two (1/4, 1/2,... 2, 4, 16) + // to scale from 256x256 pixels; the default is 1 corresponding to 256x256 cells per tile. + resolution: 1, + + // OPTIONAL + // Number, the minimum number of (estimated) rows in the dataset (query results) for aggregation to be applied. + threshold: 500000 + } +} +``` + +# History + +## 1.0.0 + + - Initial version From de8ed2720788511349068d81ce7a69dbf7b49faa Mon Sep 17 00:00:00 2001 From: Javier Goizueta Date: Tue, 9 Jan 2018 14:51:37 +0100 Subject: [PATCH 2/8] Document the tilejon and url metadata. --- docs/anonymous_maps.md | 30 ++++++++++++++++++++++++++++-- 1 file changed, 28 insertions(+), 2 deletions(-) diff --git a/docs/anonymous_maps.md b/docs/anonymous_maps.md index 9ad03487..521584a0 100644 --- a/docs/anonymous_maps.md +++ b/docs/anonymous_maps.md @@ -42,6 +42,13 @@ updated_at | The ISO date of the last time the data involved in the query was up metadata | Includes information about the layers. cdn_url | URLs to fetch the data using the best CDN for your zone. +**Improved response metadata** + +Originally, you needed to concantenate the `layergroupid` with the correct domain and the path for the tiles. +Now, for convenience, the layergroup includes the final URLs in two formats: +1. Leaflet's urlTemplate alike: useful when working with raster tiles or with libraries with an API similar to Leaflet's one. +1. [TileJSON spec](https://github.com/mapbox/tilejson-spec): useful when working with Mapbox GL or any other library that supports TileJSON. + ### Example #### Call @@ -62,11 +69,30 @@ curl 'https://{username}.carto.com/api/v1/map' -H 'Content-Type: application/jso "type": "mapnik", "meta": {} } - ] + ], + "tilejson": { + "raster": { + "tilejson": "2.2.0", + "tiles": [ + "http://a.cdb.com/c01a54877c62831bb51720263f91fb33/{z}/{x}/{y}.png", + "http://b.cdb.com/c01a54877c62831bb51720263f91fb33/{z}/{x}/{y}.png" + ] + } + }, + "url": { + "raster": { + "urlTemplate": "http://{s}.cdb.com/c01a54877c62831bb51720263f91fb33/{z}/{x}/{y}.png", + "subdomains": ["a", "b"] + } + } }, "cdn_url": { "http": "http://cdb.com", - "https": "https://cdb.com" + "https": "https://cdb.com", + "templates": { + "http": { "subdomains": ["a","b"], "url": "http://{s}.cdb.com" }, + "https": { "subdomains": ["a","b"], "url": "https://{s}.example.com" }, + } } } ``` From cef7545c173013f6319de929f60a834d54558a57 Mon Sep 17 00:00:00 2001 From: Javier Goizueta Date: Tue, 9 Jan 2018 14:51:55 +0100 Subject: [PATCH 3/8] Add documentation section for aggregation --- docs/aggregation.md | 202 ++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 202 insertions(+) create mode 100644 docs/aggregation.md diff --git a/docs/aggregation.md b/docs/aggregation.md new file mode 100644 index 00000000..37151a7d --- /dev/null +++ b/docs/aggregation.md @@ -0,0 +1,202 @@ +# Tile Aggregation + +To be able to represent a large amount of data (say, hundred of thousands to millions of points) in a til. This can be useful both for raster tiles (where the aggregation reduces the number of features to be rendered) and vector tiles (the tile contais less features). + +Aggregation is available only for point geometries. During aggregation the points are grouped using a grid; all the points laying in the same cell of the grid are summarized in a single aggregated result point. + - The position of the aggregated point is controlled by the `placement` parameter. + - The aggregated rows always contain at least a column, named `_cdb_feature_count`, which contains the number of the original points that the aggregated point represents. + +### Special default aggregation + +When no placement or columns are specified a special default aggregation is performed. + +This special mode performs only spatial aggregation (using a grid defined by the requested tile and the resolution, parameter, as all the other cases), and returns a _random_ record from each group (grid cell) with all its columns and an additional `_cdb_features_count` with the number of features in the group. + +Regarding the randomness of the sample: currently we use the row with the minimum `cartodb_id` value in each group. + +The rationale behind having this special aggregation with all the original columns is to provide a mostly transparent way to handle large datasets without having to provide special map configurations for those cases (i.e. preserving the logic used to produce the maps with smaller datasets). Overviews have been used so far with this intent, but they are inflexible. + +### User defined aggregations + +When either a explicit placement or columns are requested we no longer use the special, query; we use one determined by the placement (which will default to "centroid"), and it will have as columns only the aggregated columns specified, in addition to `_cdb_features_count`, which is always present. + +We might decide in the future to allow sampling column values for any of the different placement modes. + +### Behaviour for raster and vector tiles + +The vector tiles from a vector-only map will be aggregated by default. +However, Raster tiles (or vector tiles from a map which defines CartoCSS styles) will be aggregated only upon request. + +Aggregation that would otherwise occur can be disabled by passing an `aggregation=false` parameter to the map instantiation HTTP call. + +To control how aggregation is performed, an aggregation option can be added to the layer: + +```json +{ + "layers": [ + { + "options": { + "sql": "SELECT * FROM data", + "aggregation": { + "placement": "centroid", + "columns": { + "value": { + "aggregate_function": "sum", + "aggregated_column": "value" + } + } + } + } + } + ] +} +``` + +Even if aggregation is explicitly requested it may not be activated, e.g., if the geometries are not points +or the whole dataset is too small. The map instantiation response contains metadata that informs if any particular +layer will be aggregated when tiles are requested, both for vector (mvt) and raster (png) tiles. + +```json +{ + "layergroupid": "7b97b6e76590fef889b63edd2efb1c79:1513608333045", + "metadata": { + "layers": [ + { + "type": "mapnik", + "id": "layer0", + "meta": { + "stats": { + "estimatedFeatureCount": 6232136 + }, + "aggregation": { + "png": true, + "mvt": true + } + } + } + ] + } +} +``` + +## Aggregation parameters + +The aggregation parameters for a layer are defined inside an `aggregation` option of the layer: + +```json +{ + "layers": [ + { + "options": { + "sql": "SELECT * FROM data", + "aggregation": {"...": "..."} + } + } + ] +} +``` + +### `placement` + +Determines the kind of aggregated geometry generated: + +#### `point-sample` + +This is the default placement. It will place the aggregated point at a random sample of the grouped points, +like the default aggregation does. No other attribute is sampled, though, the point will contain the aggregated attributes determined by the `columns` parameter. + +Example: here the red dots are the original data points; the greenish bigger dots are the aggregated points and the blue lines show the aggregation grid. + +![point-sample](https://user-images.githubusercontent.com/5909/34304018-38dfe972-e738-11e7-80e3-de5016e76f24.png) + +See the `vector-agg-mapbox-gl.html` example to [check how point-sample works using vector tiles in Mapbox GL](https://bl.ocks.org/rochoa/raw/20df8dcab7325d41249d0f9269250970/vector-agg-mapbox-gl.html). + +#### `point-grid` + +Generates points at the center of the aggregation grid cells (squares). + +![grid-point](https://user-images.githubusercontent.com/5909/34304044-52a019fe-e738-11e7-869d-0ba1cb17ff54.png) + +See the `vector-agg-open-layers.html` example to [check how point-grid works using vector tiles in OpenLayers](https://bl.ocks.org/rochoa/raw/20df8dcab7325d41249d0f9269250970/vector-agg-open-layers.html). + +#### `centroid` + +Generates points with the averaged coordinated of the grouped points (i.e. the points inside each grid cell). + +![centroid](https://user-images.githubusercontent.com/5909/34304047-553324e0-e738-11e7-924c-c3778c2d72a7.png) + +See the `raster-agg-leaflet.html` example to [check how centroid works using raster tiles in Leaflet](https://bl.ocks.org/rochoa/raw/20df8dcab7325d41249d0f9269250970/raster-agg-leaflet.html). + + +### `columns` + +The aggregated attributes defined by `columns` are computed by a applying an _aggregate function_ to all the points in each group. +Valid aggregate functions are `sum`, `avg` (average), `min` (minimum), `max` (maximum) and `mode` (the most frequent value in the group). +The values to be aggregated are defined by the _aggregated column_ of the source data. The column keys define the name of the resulting column in the aggregated dataset. + +For example here we define three aggregate attributes named `total`, `max_price` and `price` which are all computed with the same column, `price`, +of the original dataset applying three different aggregate functions. + +```json +{ + "columns": { + "total": { "aggregate_function": "sum", "aggregated_column": "price" }, + "max_price": { "aggregate_function": "max", "aggregated_column": "price" }, + "price": { "aggregate_function": "avg", "aggregated_column": "price" } + } +} +``` + +> Note that you can use the original column names as names of the result, but all the result column names must be unique. In particular, the names `cartodb_id`, `the_geom`, `the_geom_webmercator` and `_cdb_feature_count` cannot be used for aggregated columns, as they correspond to columns always present in the result. + +### `resolution` + +Defines the cell-size of the spatial aggregation grid. This is equivalent to the [CartoCSS `-torque-resolution`](https://carto.com/docs/carto-engine/cartocss/properties-for-torque/#-torque-resolution-float) property of Torque maps. + +The aggregation cells are `resolution`×`resolution` pixels in size, where pixels here are defined to be 1/256 of the (linear) size of a tile. +The default value is 1, so that aggregation coincides with raster pixels. A value of 2 would make each cell to be 4 (2×2) pixels, and a value of +0.5 would yield 4 cells per pixel. In teneral values less than 1 produce sub-pixel precision. + +> Note that is independent of the number of pixels for raster tile or the coordinate resolution (mvt_extent) of vector tiles. + + +### `threshold` + +This is the minimum number of (estimated) rows in the dataset (query results) for aggregation to be applied. If the number of rows estimate is less than the threshold aggregation will be disabled for the layer; the instantiation response will reflect that and tiles will be generated without aggregation. + +### Example + +```json +{ + "version": "1.7.0", + "extent": [-20037508.5, -20037508.5, 20037508.5, 20037508.5], + "srid": 3857, + "maxzoom": 18, + "minzoom": 3, + "layers": [ + { + "type": "mapnik", + "options": { + "sql": "select * from table", + "cartocss": "#table { marker-width: [total]; marker-fill: ramp(value, (red, green, blue), jenks); }", + "cartocss_version": "2.3.0", + "aggregation": { + "placement": "centroid",s + "columns": { + "value": { + "aggregate_function": "avg", + "aggregated_column": "value" + }, + "total": { + "aggregate_function": "sum", + "aggregated_column": "value" + } + }, + "resolution": 2, // Aggregation cell is 2x2 pixels + "threshold": 500000 + } + } + } + ] +} +``` From e34410fd2cfd3ca8f0410c3648c593cf54837109 Mon Sep 17 00:00:00 2001 From: Javier Goizueta Date: Tue, 9 Jan 2018 15:08:46 +0100 Subject: [PATCH 4/8] Add references to general aggregation documentation in MapConfig spec --- docs/MapConfig-Aggregation-extension.md | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/docs/MapConfig-Aggregation-extension.md b/docs/MapConfig-Aggregation-extension.md index ffcd5144..850b179a 100644 --- a/docs/MapConfig-Aggregation-extension.md +++ b/docs/MapConfig-Aggregation-extension.md @@ -22,7 +22,7 @@ The value of this attribute can be `false` to explicitly disable aggregation for // * "point-sample", the default places geometries at a sample point (one of the aggregated geometries) // * "point-grid" places geometries at the center of the aggregation grid cells // * "centroid" places geometriea at the average position of the aggregated points - // See https://github.com/CartoDB/Windshaft-cartodb/blob/master/docs/... for more details + // See https://github.com/CartoDB/Windshaft-cartodb/blob/master/docs/aggregation.md#placement for more details placement: "point-sample", // OPTIONAL @@ -33,6 +33,7 @@ The value of this attribute can be `false` to explicitly disable aggregation for // is always generated in addition to the defined columns. // The column names `cartodb_id`, `the_geom`, `the_geom_webmercator` and `_cdb_feature_count` cannot be used // for aggregated columns, as they correspond to columns always present in the result. + // See https://github.com/CartoDB/Windshaft-cartodb/blob/master/docs/aggregation.md#columns for more details columns: { "aggregated_column_1": { "aggregate_function": "sum", @@ -43,10 +44,12 @@ The value of this attribute can be `false` to explicitly disable aggregation for // OPTIONAL // Number, defines the cell-size of the spatial aggregation grid as a pixel resolution power of two (1/4, 1/2,... 2, 4, 16) // to scale from 256x256 pixels; the default is 1 corresponding to 256x256 cells per tile. + // See https://github.com/CartoDB/Windshaft-cartodb/blob/master/docs/aggregation.md#resolution for more details resolution: 1, // OPTIONAL // Number, the minimum number of (estimated) rows in the dataset (query results) for aggregation to be applied. + // See https://github.com/CartoDB/Windshaft-cartodb/blob/master/docs/aggregation.md#threshold for more details threshold: 500000 } } From 99324b15ef174059e977ae1280bc05c6412a6226 Mon Sep 17 00:00:00 2001 From: Javier Goizueta Date: Tue, 9 Jan 2018 15:58:20 +0100 Subject: [PATCH 5/8] Remove placement examples --- docs/aggregation.md | 15 --------------- 1 file changed, 15 deletions(-) diff --git a/docs/aggregation.md b/docs/aggregation.md index 37151a7d..84c76415 100644 --- a/docs/aggregation.md +++ b/docs/aggregation.md @@ -105,29 +105,14 @@ Determines the kind of aggregated geometry generated: This is the default placement. It will place the aggregated point at a random sample of the grouped points, like the default aggregation does. No other attribute is sampled, though, the point will contain the aggregated attributes determined by the `columns` parameter. -Example: here the red dots are the original data points; the greenish bigger dots are the aggregated points and the blue lines show the aggregation grid. - -![point-sample](https://user-images.githubusercontent.com/5909/34304018-38dfe972-e738-11e7-80e3-de5016e76f24.png) - -See the `vector-agg-mapbox-gl.html` example to [check how point-sample works using vector tiles in Mapbox GL](https://bl.ocks.org/rochoa/raw/20df8dcab7325d41249d0f9269250970/vector-agg-mapbox-gl.html). - #### `point-grid` Generates points at the center of the aggregation grid cells (squares). -![grid-point](https://user-images.githubusercontent.com/5909/34304044-52a019fe-e738-11e7-869d-0ba1cb17ff54.png) - -See the `vector-agg-open-layers.html` example to [check how point-grid works using vector tiles in OpenLayers](https://bl.ocks.org/rochoa/raw/20df8dcab7325d41249d0f9269250970/vector-agg-open-layers.html). - #### `centroid` Generates points with the averaged coordinated of the grouped points (i.e. the points inside each grid cell). -![centroid](https://user-images.githubusercontent.com/5909/34304047-553324e0-e738-11e7-924c-c3778c2d72a7.png) - -See the `raster-agg-leaflet.html` example to [check how centroid works using raster tiles in Leaflet](https://bl.ocks.org/rochoa/raw/20df8dcab7325d41249d0f9269250970/raster-agg-leaflet.html). - - ### `columns` The aggregated attributes defined by `columns` are computed by a applying an _aggregate function_ to all the points in each group. From e57c4c824b4c4738af46ddc8aff52fcf975c7d4b Mon Sep 17 00:00:00 2001 From: Raul Ochoa Date: Thu, 11 Jan 2018 11:45:44 +0000 Subject: [PATCH 6/8] fix invalid json --- docs/aggregation.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/aggregation.md b/docs/aggregation.md index 84c76415..48e453de 100644 --- a/docs/aggregation.md +++ b/docs/aggregation.md @@ -166,7 +166,7 @@ This is the minimum number of (estimated) rows in the dataset (query results) fo "cartocss": "#table { marker-width: [total]; marker-fill: ramp(value, (red, green, blue), jenks); }", "cartocss_version": "2.3.0", "aggregation": { - "placement": "centroid",s + "placement": "centroid", "columns": { "value": { "aggregate_function": "avg", @@ -177,7 +177,7 @@ This is the minimum number of (estimated) rows in the dataset (query results) fo "aggregated_column": "value" } }, - "resolution": 2, // Aggregation cell is 2x2 pixels + "resolution": 2, "threshold": 500000 } } From 72bebf19607fbbae5bf6ad172561462fe8092d88 Mon Sep 17 00:00:00 2001 From: Raul Ochoa Date: Thu, 11 Jan 2018 15:15:25 +0000 Subject: [PATCH 7/8] Fix typo --- docs/aggregation.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/aggregation.md b/docs/aggregation.md index 48e453de..69318259 100644 --- a/docs/aggregation.md +++ b/docs/aggregation.md @@ -1,6 +1,6 @@ # Tile Aggregation -To be able to represent a large amount of data (say, hundred of thousands to millions of points) in a til. This can be useful both for raster tiles (where the aggregation reduces the number of features to be rendered) and vector tiles (the tile contais less features). +To be able to represent a large amount of data (say, hundred of thousands to millions of points) in a tile. This can be useful both for raster tiles (where the aggregation reduces the number of features to be rendered) and vector tiles (the tile contais less features). Aggregation is available only for point geometries. During aggregation the points are grouped using a grid; all the points laying in the same cell of the grid are summarized in a single aggregated result point. - The position of the aggregated point is controlled by the `placement` parameter. From d9e66c596466d3a49e81dba5c3a66441f6d6e171 Mon Sep 17 00:00:00 2001 From: Raul Ochoa Date: Thu, 11 Jan 2018 15:15:33 +0000 Subject: [PATCH 8/8] Link to overviews doc --- docs/aggregation.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/aggregation.md b/docs/aggregation.md index 69318259..2fb361c3 100644 --- a/docs/aggregation.md +++ b/docs/aggregation.md @@ -14,7 +14,7 @@ This special mode performs only spatial aggregation (using a grid defined by the Regarding the randomness of the sample: currently we use the row with the minimum `cartodb_id` value in each group. -The rationale behind having this special aggregation with all the original columns is to provide a mostly transparent way to handle large datasets without having to provide special map configurations for those cases (i.e. preserving the logic used to produce the maps with smaller datasets). Overviews have been used so far with this intent, but they are inflexible. +The rationale behind having this special aggregation with all the original columns is to provide a mostly transparent way to handle large datasets without having to provide special map configurations for those cases (i.e. preserving the logic used to produce the maps with smaller datasets). [Overviews have been used so far with this intent](https://carto.com/docs/tips-and-tricks/back-end-data-performance/), but they are inflexible. ### User defined aggregations