Skip to content
ClickHouse Docs

Machine Learning Functions

Autogenerated from ClickHouse system tables

assignCentroid

Introduced in: v26.8.0

Returns the id of the nearest (L2) centroid to a vector. The centroids are given as a constant array of float arrays, where the id is the 0-based position in that array, or as the name of a Dictionary holding the attributes cid and vec, where the id is cid.

Syntax

assignCentroid(vec, centroids | dict_name)

Arguments

  • vec — Vector to assign. Its dimension must match the dimension of the centroids. Widths other than Float32 are converted to Float32, which is what the scoring kernel uses. Array(Float32) or Array(Float64) or Array(BFloat16)
  • centroids — The centroids to score against, which must be constant. Given as an array of equally sized, non-empty float arrays, the id then being the 0-based position in that array; or as the name of a Dictionary with an attribute cid of an unsigned integer type that fits UInt32 and an attribute vec of type Array(Float32), the id then being cid. The dictionary is read once and cached until it reloads. Array(Array(Float32)) or Array(Array(Float64)) or Array(Array(BFloat16)) or String

Returned value

The nearest centroid id. UInt32

Examples

Inline centroids

SELECT assignCentroid([1.0, 2.0]::Array(Float32), [[0.0, 0.0], [1.0, 2.0]]::Array(Array(Float32)))
1

evalMLMethod

Introduced in: v20.1.0

Applies a trained machine learning model to input features to generate predictions.

Syntax

evalMLMethod(model, x1[, x2, ...])

Arguments

Returned value

Returns the predicted value based on the trained model. Float64

Examples

Example usage

CREATE TABLE trips (pickup_datetime DateTime('UTC'), trip_distance Float64, total_amount Float64) ENGINE = Memory;

-- A fare of 3, plus 2.5 for every unit of distance.
INSERT INTO trips
SELECT toDateTime('2020-01-01 00:00:00', 'UTC') + number * 60, number % 10 + 1, 2.5 * (number % 10 + 1) + 3
FROM numbers(1000);

-- One model per year of the data.
CREATE TABLE models ENGINE = Memory AS
SELECT
    toYear(pickup_datetime) AS year,
    stochasticLinearRegressionState(0.01, 0.0, 10, 'SGD')(total_amount, trip_distance) AS model
FROM trips
GROUP BY year;

SELECT
    trip_distance,
    round(evalMLMethod(model, trip_distance), 2) AS predicted,
    total_amount
FROM trips
LEFT JOIN models ON year = toYear(pickup_datetime)
ORDER BY pickup_datetime
LIMIT 5
┌─trip_distance─┬─predicted─┬─total_amount─┐
│             1 │      4.05 │          5.5 │
│             2 │      6.79 │            8 │
│             3 │      9.53 │         10.5 │
│             4 │     12.28 │           13 │
│             5 │     15.02 │         15.5 │
└───────────────┴───────────┴──────────────┘

naiveBayesClassifier

Introduced in: v25.11.0

Classifies input text using a NAIVE_BAYES dictionary. Returns the same predicted class value as dictGet(dictionary_name, class_attribute, input_text), where class_attribute is the name of the class label attribute configured in the dictionary’s layout. Unlike dictGet, the result type is always UInt32 rather than the declared type of the class attribute, and input_text must be a String (no key type conversion is applied).

Syntax

naiveBayesClassifier(dictionary_name, input_text)

Arguments

  • dictionary_name — Name of a dictionary with the NAIVE_BAYES layout. String
  • input_text — Text to classify. String

Returned value

Predicted class ID. UInt32

Examples

Classify text

-- A dictionary built from the token counts of two classes: 0 for a positive review, 1 for a negative one.
CREATE TABLE review_tokens (ngram String, class_id UInt32, count UInt64) ENGINE = Memory;
INSERT INTO review_tokens VALUES ('good', 0, 5), ('great', 0, 4), ('excellent', 0, 3), ('bad', 1, 5), ('awful', 1, 4), ('terrible', 1, 3);

CREATE DICTIONARY sentiment (ngram String, class_id UInt32 DEFAULT 0, count UInt64 DEFAULT 0)
PRIMARY KEY ngram
SOURCE(CLICKHOUSE(TABLE 'review_tokens'))
LAYOUT(NAIVE_BAYES(class_attribute 'class_id' n 1 mode 'token'))
LIFETIME(0);

SELECT naiveBayesClassifier('sentiment', 'a good and great film') AS class_id;
┌─class_id─┐
│        0 │
└──────────┘

naiveBayesClassifierWithAllProbs

Introduced in: v26.7.0

Classifies input text using a NAIVE_BAYES dictionary and returns all classes with their probabilities, ordered from most to least probable.

Syntax

naiveBayesClassifierWithAllProbs(dictionary_name, input_text)

Arguments

  • dictionary_name — Name of a dictionary with the NAIVE_BAYES layout. String
  • input_text — Text to classify. String

Returned value

Array of (class_id, probability) tuples ordered from most to least probable. Array(Tuple(UInt32, Float64))

Examples

All class probabilities

-- A dictionary built from the token counts of two classes: 0 for a positive review, 1 for a negative one.
CREATE TABLE review_tokens (ngram String, class_id UInt32, count UInt64) ENGINE = Memory;
INSERT INTO review_tokens VALUES ('good', 0, 5), ('great', 0, 4), ('excellent', 0, 3), ('bad', 1, 5), ('awful', 1, 4), ('terrible', 1, 3);

CREATE DICTIONARY sentiment (ngram String, class_id UInt32 DEFAULT 0, count UInt64 DEFAULT 0)
PRIMARY KEY ngram
SOURCE(CLICKHOUSE(TABLE 'review_tokens'))
LAYOUT(NAIVE_BAYES(class_attribute 'class_id' n 1 mode 'token'))
LIFETIME(0);

SELECT arrayMap(p -> (p.1, round(p.2, 4)), naiveBayesClassifierWithAllProbs('sentiment', 'a good and great film')) AS predictions;
┌─predictions─────────────┐
│ [(0,0.9677),(1,0.0323)] │
└─────────────────────────┘

naiveBayesClassifierWithProb

Introduced in: v26.7.0

Classifies input text using a NAIVE_BAYES dictionary and returns the predicted class with its probability.

Syntax

naiveBayesClassifierWithProb(dictionary_name, input_text)

Arguments

  • dictionary_name — Name of a dictionary with the NAIVE_BAYES layout. String
  • input_text — Text to classify. String

Returned value

Tuple of (class_id, probability). Tuple(UInt32, Float64)

Examples

Classify with probability

-- A dictionary built from the token counts of two classes: 0 for a positive review, 1 for a negative one.
CREATE TABLE review_tokens (ngram String, class_id UInt32, count UInt64) ENGINE = Memory;
INSERT INTO review_tokens VALUES ('good', 0, 5), ('great', 0, 4), ('excellent', 0, 3), ('bad', 1, 5), ('awful', 1, 4), ('terrible', 1, 3);

CREATE DICTIONARY sentiment (ngram String, class_id UInt32 DEFAULT 0, count UInt64 DEFAULT 0)
PRIMARY KEY ngram
SOURCE(CLICKHOUSE(TABLE 'review_tokens'))
LAYOUT(NAIVE_BAYES(class_attribute 'class_id' n 1 mode 'token'))
LIFETIME(0);

WITH naiveBayesClassifierWithProb('sentiment', 'a good and great film') AS p
SELECT (p.1, round(p.2, 4)) AS prediction;
┌─prediction─┐
│ (0,0.9677) │
└────────────┘