# Functions

> Aggregation, string, numeric, collection, type-conversion, date and geospatial functions in WhisperGraph Cypher, each with a call and its result.

*Source: https://www.whisper.security/docs/cypher/functions*

---
WhisperGraph Cypher ships a function library you use inside `RETURN`, `WITH`, and `WHERE`. The tables below give the call and the value it returns. For the clauses that hold these functions, see [Syntax & Clauses](/docs/cypher/syntax); for the `CALL` procedures (`explain`, `whisper.assess`, `whisper.variants`, `whisper.history`, `whisper.origins`), see [Procedures](/docs/whisper-graph/procedures).

An unknown function name is an error, never a `null`: `RETURN notAFunction(1)` answers `400 query-error` with `Unknown function: notAFunction` and points you at `CALL db.functions()`. `explain` is a procedure with no function form, so call it with `CALL`, not inside an expression.

## Aggregation

Aggregations collapse rows; any non-aggregated column in the same `RETURN` or `WITH` becomes a grouping key. They are planned rather than dispatched, so `CALL db.functions()` does not list them.

| Function | Example | Result |
|----------|---------|--------|
| `count` | `count(*)`, `count(c)` | row count |
| `count(DISTINCT ...)` | `count(DISTINCT ip)` | distinct count |
| `sum` | `sum(x)` over `[1,2,3,4]` | `10.0` |
| `avg` | `avg(x)` over `[1,2,3,4]` | `2.5` |
| `min` / `max` | `min(x)` / `max(x)` over `[5,2,8]` | `2` / `8` |
| `percentileCont` | `percentileCont(x, 0.5)` over `[1,2,3,4]` | `2.5` (interpolated; `0.5` is the median) |
| `percentileDisc` | `percentileDisc(x, 0.5)` over `[1,2,3,4]` | `2.0` (an actual value from the set) |
| `stDev` | `stDev(x)` over `[1,2,3,4]` | `1.29…` (sample standard deviation) |
| `stDevP` | `stDevP(x)` over `[1,2,3,4]` | `1.118…` (population standard deviation) |
| `collect` | `collect(c.name)` | a list |
| `collect(DISTINCT ...)` | `collect(DISTINCT x)` over `[1,1,2]` | `[1,2]` |

Two habits keep aggregations cheap. `collect(DISTINCT x)[0..N]` slices the list after it is built, so the whole fan-out is collected first: bound the input with `WITH x LIMIT n` before you collect. And a grouped result is as large as the number of distinct groups, which a trailing `LIMIT` does not shrink: group on a coarser key (country rather than city, ASN rather than prefix, category rather than feed), or bound the input before you aggregate.

## String

| Function | Example | Result |
|----------|---------|--------|
| `toUpper` / `upper` | `toUpper("abc")` | `ABC` |
| `toLower` / `lower` | `toLower("ABC")` | `abc` |
| `trim` / `ltrim` / `rtrim` | `trim("  hi  ")` | `hi` |
| `replace` | `replace("foobar","bar","baz")` | `foobaz` |
| `substring` | `substring("hello",1,3)` / `substring("hello",2)` | `ell` / `llo` |
| `split` | `split("a,b,c",",")` | `["a","b","c"]` |
| `left` / `right` | `left("hello",2)` / `right("hello",2)` | `he` / `lo` |
| `reverse` | `reverse("abc")` | `cba` |
| `size` / `length` | `size("abc")` | `3` (string length) |
| `isEmpty` | `isEmpty("")` | `true` |
| `toString` | `toString(123)` | `123` |

String concatenation uses `+`. Lowercase an anchor value in your own code, not with `toLower()` in the query: names are stored lowercase, and wrapping the anchor in a function turns an indexed lookup into a scan.

## Numeric

| Function | Example | Result |
|----------|---------|--------|
| `abs` | `abs(-5)` | `5` |
| `ceil` / `ceiling` / `floor` | `ceil(4.2)` / `floor(4.8)` | `5.0` / `4.0` |
| `round` | `round(4.5)` | `5` (an integer) |
| `sign` | `sign(-3)` | `-1` |
| `sqrt` | `sqrt(16)` | `4.0` |
| `log` / `ln` / `log10` / `exp` | `log10(1000)` / `ln(e())` | `3.0` / `1.0` |
| `rand` | `rand()` | a value in [0,1) |
| `e` / `pi` | `pi()` | π |

Arithmetic operators: `+`, `-`, `*`, `/`, `%`, `^` (exponent; `2 ^ 3` is `8.0`).

## Trigonometric

| Function | Example | Result |
|----------|---------|--------|
| `sin` / `cos` / `tan` | `cos(0)` | `1.0` |
| `asin` / `acos` / `atan` / `atan2` | `atan2(0,1)` | `0.0` |
| `degrees` | `degrees(pi())` | `180.0` |
| `radians` | `radians(180)` | π |

## Collection

| Function | Example | Result |
|----------|---------|--------|
| `size` | `size([1,2,3])` | `3` |
| `head` / `last` | `head([10,20,30])` | `10` |
| `tail` | `tail([10,20,30])` | `[20,30]` |
| `range` | `range(1,5)` / `range(0,10,5)` | `[1,2,3,4,5]` / `[0,5,10]` |
| `reverse` | `reverse([1,2,3])` | `[3,2,1]` |
| `keys` | `keys(node)`, `keys(rel)` | property keys (`keys(r)` on a `RESOLVES_TO` edge gives `source`, `inferred`) |
| `isEmpty` | `isEmpty([])` | `true` |

List comprehensions and pattern comprehensions build lists inline; both are covered in [Syntax & Clauses](/docs/cypher/syntax#list-and-pattern-comprehensions).

## Node and relationship

| Function | Example | Result |
|----------|---------|--------|
| `id` | `id(n)` | the node id, as a string (`"906258972"`) |
| `elementId` | `elementId(n)` | the qualified form (`"4:whisper:906258972"`) |
| `label` / `labels` | `label(n)` / `labels(n)` | `"HOSTNAME"` / `["HOSTNAME"]` |
| `type` | `type(r)` | `RESOLVES_TO` |
| `properties` | `properties(n)` | a property map |
| `startNode` / `endNode` | `startNode(r).name` | a node |
| `nodes` / `relationships` | `size(nodes(p))` | node count |
| `length` | `length(p)` | path length (hop count) |

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})
RETURN h.name, id(h) AS id, elementId(h) AS elementId
LIMIT 1
```

Ids are strings, so compare them with `=`: `WHERE id(a) = id(b)`. An ordering comparison such as `WHERE id(n) > 0` compares a string with a number and matches nothing. Never look a node up by id across the whole graph (`MATCH (n) WHERE id(n) = "…"`), which is an unanchored scan. Anchor on `name` instead; on `URL` the name is the kit path. Adding the label does not help: `MATCH (u:URL {id: "…"})` still scans every node of the label, and on `URL` the same id can name a different kit on the next call.

## Type conversion

| Function | Example | Result |
|----------|---------|--------|
| `toInteger` / `toInt` | `toInteger("42")` | `42` |
| `toFloat` | `toFloat("3.14")` | `3.14` |
| `toBoolean` | `toBoolean("true")` | `true` |
| `toIntegerList` / `toFloatList` / `toStringList` / `toBooleanList` | `toIntegerList(["1","2"])` | `[1,2]` |

Input that cannot be parsed yields `null` rather than an error: `toInteger("abc")` and `toFloat("abc")` both return `null`.

## Date and time

| Function | Example | Result |
|----------|---------|--------|
| `timestamp` | `timestamp()` | epoch millis |
| `date` | `date()` | today's date, `YYYY-MM-DD` |
| `datetime` / `localdatetime` | `datetime()` | an ISO-8601 timestamp |
| `time` / `localtime` | `time()` | a time of day |
| `duration` | `duration("P1D")` | `{period: "P1D", duration: "PT0S"}` |
| `duration.between` | `duration.between(date("2020-01-01"), date("2020-03-01"))` | `{period: "P2M", duration: "PT0S"}` |
| `duration.inDays` / `duration.inMonths` / `duration.inSeconds` | `duration.inDays(date("2020-01-01"), date("2020-03-01"))` | `{period: "P60D", duration: "PT0S"}` |

## Geospatial and misc

| Function | Example | Result |
|----------|---------|--------|
| `point` | `point({x: 1.0, y: 2.0})` | `{latitude: 2.0, longitude: 1.0}` |
| `distance` / `point.distance` | `distance(point({x: 0, y: 0}), point({x: 3, y: 4}))` | `555811.94…` |
| `coalesce` | `coalesce(a.missing, "default")` | `default` |
| `randomUUID` | `randomUUID()` | a UUID string |

Points are geodesic, not Cartesian: `x` is longitude and `y` is latitude, and `distance()` returns metres along the globe, which is why the example above is not `5`. `distance` and `point.distance` behave identically.

## Functions that introspection leaves out

`CALL db.functions()` is the right way to check a name, but a few functions run even though the listing omits them: `ln`, `localtime`, `duration.between`, `duration.inDays`, `duration.inMonths`, `duration.inSeconds`, `point.distance`, and `whisper.variants`, which as a function returns the variant list directly (`RETURN whisper.variants("google.com")[0..3]`). An omission from the listing is not a rejection; only the `Unknown function` error is. The aggregation functions above are absent from the listing for the same reason: they are planned, not dispatched.

> **Read `coverage` before `band`.** Only `known-clean` licenses the word "clean"; `no-data` means
> *unknown*, which is a different thing again; `malicious-evidenced` and `ambiguous` mean there is
> evidence, whatever the band says.
> Full contract: [Coverage — what we looked at](/docs/whisper-graph/procedures/coverage).

## Reading threat properties off a node

You do not need a function to read a verdict — the threat posture lives directly on the node. An `IPV4` node carries `threatScore`, `threatLevel`, `isThreat`, `isTor`, and `isAnonymizer`, so a single anchored read gives you the whole posture with no extra hops.

```cypher expect=rows>0 seed=185.220.101.1 verified=2026-09-02
MATCH (ip:IPV4 {name: "185.220.101.1"})
RETURN ip.name AS ip, ip.threatScore AS score, ip.threatLevel AS level,
       ip.isThreat AS isThreat, ip.isTor AS isTor, ip.isAnonymizer AS isAnonymizer
LIMIT 1
```

Name the properties you read. A whole-node projection (`RETURN ip`, `keys(ip)`, `properties(ip)`) may leave the reconciled verdict fields (`verdictLevel`, `verdictScore`, `verdictCoverage`, and their siblings) out for speed, and the response then carries a `projection-verdict-omitted` advisory in its top-level `advisories[]` array. Either project the property you need (`RETURN ip.verdictLevel`) or send `projectionFull: true` in the request body to get the full verdict surface back.

For the scored reasoning behind a verdict — the feeds, weights, and factors — call `explain()`. See [explain() — Threat Verdicts](/docs/whisper-graph/procedures/explain).
