Introduced in: v26.8.0
Selects up to k time series at each time step of a time grid, in a deterministic pseudo-random fashion.
Each input row is one time series: key identifies the series, sampling_key is a per-series hash used as the
selection order, and values contains the values of the series aligned to a common time grid (so the values arrays
of all rows must have the same size). At each time step, the series with the k smallest sampling keys among the series
with non-NULL values at that step are selected. Sampling key ties are broken by preferring the series with the smaller key.
This function implements the limitk() aggregation operator of PromQL and keeps only one bounded heap of size k per
time step, so its state size does not depend on the number of aggregated series.
Syntax
timeSeriesLimitKMasks(k, key, sampling_key, values)Arguments
k— How many series to select at each time step, either one value for all time steps or an array with one value per time step. Must be the same for all rows.UInt*orArray(UInt*)key— Identifier of the time series.UInt64sampling_key— A per-series hash defining the selection order, e.g.timeSeriesGroupToSamplingKey(key).UInt64values— Values of the time series aligned to the time grid, one element per time step.Array(Nullable(Float32))orArray(Nullable(Float64))orArray(Float32)orArray(Float64)
Returned value
Returns the selected series in the order of ascending key, each together with its per-step mask: steps_mask[t] = 1 if the series is selected at time step t. Series which are selected at no time step are not returned. Array(Tuple(key UInt64, steps_mask Array(UInt8)))
Examples
Selecting 2 series per time step by the smallest sampling keys
SET enable_time_series_aggregate_functions = 1;
WITH [(1, [10., 1., NULL], 300), (2, [20., 2., 2.], 100), (3, [30., NULL, 1.], 200)]::Array(Tuple(UInt64, Array(Nullable(Float64)), UInt64)) AS series
SELECT timeSeriesLimitKMasks(2, s.1, s.3, s.2)
FROM (SELECT arrayJoin(series) AS s);[(1,[0,1,0]),(2,[1,1,1]),(3,[1,0,1])]