# Syntax & Clauses

> Every Cypher clause WhisperGraph supports, each with an example: MATCH, WHERE, WITH, UNWIND, UNION, CALL subqueries, parameters, EXPLAIN and paths.

*Source: https://www.whisper.security/docs/cypher/syntax*

---
WhisperGraph implements a read-only Cypher dialect. You send Cypher over HTTP and get back columns and rows. Write clauses (`CREATE`, `MERGE`, `SET`, `DELETE`, `REMOVE`, `FOREACH`) are not supported — the parser recognizes them and rejects them with a `readonly_engine` suggestion before anything runs.

Two rules make every query fast: anchor the starting node by its `name` (an indexed lookup), and add a `LIMIT`. This page walks each clause with a runnable example, then covers parameters, batching, and the plan you get from `EXPLAIN`. For the labels and edge directions you traverse, see the [Graph Schema](/docs/whisper-graph/schema); for the procedures you call with `CALL`, the [Procedures](/docs/whisper-graph/procedures) reference. For the golden rules and pitfalls, see [Best Practices](/docs/cypher/best-practices).

## MATCH

`MATCH` finds patterns in the graph. The fastest form anchors a node by its `name`, which is an indexed lookup.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"}) RETURN h.name
```

Chain a relationship to reach the node on the other end:

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(p:ANNOUNCED_PREFIX)
RETURN p.name LIMIT 5
```

A label-only match with no `{name: ...}` scans every node of that label. That is fine on small labels like `CATEGORY`, but it never finishes on billion-node labels like `HOSTNAME` or `IPV4` — always anchor those. State the label as well as the name: a bare `MATCH (h {name: "..."})` is not planned the same way and can miss a sparsely connected name.

Names are stored lowercase, without a trailing dot, and they are matched exactly. Normalize in your own code before you anchor: `Google.com` does not reach the `google.com` node, and `gmail.com.` is not the same node as `gmail.com`. Names with special characters anchor as plain strings, so `*spf.google.com` and punycode names such as `xn--80ak6aa92e.com` need no escaping.

An unknown label or edge name in an anchored pattern is not an error: it matches nothing. When a correct-looking query returns zero rows, check `CALL db.labels()` and `CALL db.relationshipTypes()` first. Labels from other graph products are the one exception: `Domain`, `IpAddress`, and `Certificate` are rejected with a `schema-drift` error that names the label to use instead (`HOSTNAME`, `IPV4`, and `CT_OBSERVATION`).

## OPTIONAL MATCH

`OPTIONAL MATCH` keeps the driving row even when the optional pattern has no match, filling the missing columns with `null`. Use it for sparse fields like WHOIS contacts or geolocation, where a plain `MATCH` would drop the whole row.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})
OPTIONAL MATCH (h)-[:HAS_EMAIL]->(e:EMAIL)
OPTIONAL MATCH (h)-[:HAS_REGISTRAR]->(r:REGISTRAR)
RETURN h.name, collect(DISTINCT e.name) AS emails, collect(DISTINCT r.name) AS registrars
```

## WHERE

`WHERE` filters bound rows. The supported operators:

- **Comparison:** `=`, `<>`, `<`, `>`, `<=`, `>=`
- **Logical:** `AND`, `OR`, `NOT`, `XOR`
- **Null checks:** `IS NULL`, `IS NOT NULL`
- **List membership:** `IN`
- **String predicates:** `STARTS WITH`, `ENDS WITH`, `CONTAINS`, `=~` (regex)

```cypher expect=rows>0 seed=cloudflare. verified=2026-09-02
MATCH (h:HOSTNAME)
WHERE h.name STARTS WITH "cloudflare."
RETURN h.name LIMIT 5
```

```cypher expect=rows>0 seed=.cloudflare.com verified=2026-09-02
MATCH (h:HOSTNAME)
WHERE h.name ENDS WITH ".cloudflare.com"
RETURN h.name LIMIT 5
```

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH (a:ASN)
WHERE a.name IN ["AS13335", "AS15169"]
RETURN a.name
```

`STARTS WITH` and `ENDS WITH` on `.name` are index-backed. Keep the suffix narrow and leading-dot (`ENDS WITH ".cloudflare.com"`, never `ENDS WITH "com"`), and remember that only hostnames carry the suffix index: on `PREFIX` or `ASN`, anchor instead. `CONTAINS` is fine once the query is anchored or paired with `STARTS WITH`; never run it across an unanchored label. For a token you cannot classify, use `CALL whisper.search("token")`, which routes to an indexed lookup and never scans. On `ASN`, `.name` is the AS number (`AS13335`), so match it exactly or with `STARTS WITH "AS"`, not with `CONTAINS`.

Regex is a full match against the whole value and only gets an index when it is a plain prefix or a `.*literal.*` shape, so prefer the string predicates and use `=~` on rows you have already anchored:

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH (a:ASN {name: "AS13335"}) WHERE a.name =~ "AS[0-9]+" RETURN a.name
```

`"AS133"` would not match `AS13335`; the pattern has to cover the whole name. An `IN` list is rewritten into indexed lookups; for a long list, switch to `UNWIND` (below).

## RETURN

`RETURN` selects what comes back. Use `AS` for aliases and `DISTINCT` to deduplicate. `RETURN *` returns every bound variable, and literals of any type (numbers, strings, booleans, lists, maps) can be returned directly.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(i)
RETURN DISTINCT labels(i)[0] AS family
```

One case to plan around: across a chain that runs through an announced prefix (`ANNOUNCED_BY`, then `ROUTES`), aggregate with `count(DISTINCT ...)` or de-duplicate in your client rather than writing `RETURN DISTINCT` over the projection. [Best Practices](/docs/cypher/best-practices) has the pattern.

## WITH

`WITH` pipes results from one part of a query to the next. It is how you aggregate or narrow a set before traversing further, and you can filter after it with `WHERE`.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})<-[:NAMESERVER_FOR]-(ns:HOSTNAME)
WITH ns LIMIT 3
MATCH (ns)-[:NAMESERVER_FOR]->(sibling:HOSTNAME)
RETURN ns.name AS nameserver, collect(DISTINCT sibling.name)[0..8] AS domains
LIMIT 3
```

This anchor-then-narrow-then-expand shape is the single most useful pattern in the language. Bounding the intermediate set with `WITH ... LIMIT` keeps a two-stage query from exploding, and it puts the bound where the fan-out happens: a `LIMIT` at the end of the query does not bound the traversal that feeds it.

## ORDER BY, LIMIT, SKIP

`ORDER BY` sorts, `LIMIT` caps the row count, and `SKIP` offsets for pagination. Always include a `LIMIT`, and use literal numbers in `SKIP` / `LIMIT`.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (sub:HOSTNAME)-[:CHILD_OF]->(:HOSTNAME {name: "google.com"})
RETURN sub.name AS subdomain
ORDER BY sub.name SKIP 0 LIMIT 15
```

To page, keep a stable `ORDER BY` and walk `SKIP` forward: `SKIP 0 LIMIT 15`, then `SKIP 15 LIMIT 15`, and so on. `LIMIT 0` returns no rows, and a `SKIP` past the end returns an empty page. If you bind `SKIP` or `LIMIT` to a parameter and the value resolves to `null`, the bound is dropped and the response carries a `null-pagination-param` advisory; pass a number.

> **Read `coverage` before `band`.** Only `known-clean` licenses the word "clean"; `no-data` means
> *unknown*, which is a different thing again; `malicious-evidenced` and `ambiguous` mean there is
> evidence, whatever the band says.
> Full contract: [Coverage — what we looked at](/docs/whisper-graph/procedures/coverage).

## UNWIND

`UNWIND` turns a list into rows, one per element. It is the right pattern for batch lookups: each element becomes its own anchored query.

```cypher expect=rows>0 seed=185.220.101.1 verified=2026-09-02
UNWIND ["185.220.101.1", "104.16.123.96", "8.8.8.8"] AS addr
MATCH (ip:IPV4 {name: addr})
RETURN ip.name AS ip, ip.threatLevel AS level, ip.isThreat AS isThreat
```

The `MATCH` after `UNWIND` is still anchored — each row binds `ip` on its indexed `name`. `UNWIND` followed by `CALL` runs a procedure once per element (see below), and `UNWIND $names AS n MATCH (h:HOSTNAME {name: n})` is the form to use when an `IN` list grows long.

## UNION

`UNION` combines results from multiple queries and deduplicates; `UNION ALL` keeps duplicates. Every branch must return the same column names, and each branch can carry its own `LIMIT`.

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(p) RETURN p.name AS n LIMIT 5
UNION
MATCH (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR]-(peer:ASN) RETURN peer.name AS n LIMIT 5
```

## CALL procedures

`CALL` runs a procedure. Standalone, or with `YIELD` to name the columns you want and feed the rest of the query. See [Procedures](/docs/whisper-graph/procedures) for the full set.

```cypher expect=rows>0 seed=1.1.1.1 verified=2026-09-02
CALL explain("1.1.1.1") YIELD indicator, score, level
RETURN indicator, score, level LIMIT 1
```

A bare `CALL explain("1.1.1.1")` returns every column the procedure defines, including transport fields such as `available` and `cached` and an `advisory` slot that stays empty when there is nothing to advise. Name the columns you read with `YIELD`, as above, and the result stays stable as the procedure grows.

```cypher expect=rows>0 seed=paypal.com verified=2026-09-02
CALL whisper.variants("paypal.com")
YIELD variant, method, exists
WHERE exists
RETURN variant, method LIMIT 10
```

A `CALL` placed after `UNWIND`, `WITH`, or `MATCH` runs once per incoming row, so you can score a whole list in one query:

```cypher expect=rows>0 seed=1.1.1.1 verified=2026-09-02
UNWIND ["1.1.1.1", "8.8.8.8"] AS ip
CALL explain(ip) YIELD indicator, score, level
RETURN indicator, score, level LIMIT 2
```

Three rules keep procedure calls out of trouble:

- **Quote every argument.** `CALL whisper.identify(ubuntu.com)` is a bad-argument error, and an unquoted IPv6 literal is parsed as something else entirely. Always `CALL whisper.identify("ubuntu.com")`.
- **`YIELD` columns are exact contracts.** A column the procedure does not emit is rejected, not ignored, and the error lists the columns it does emit. `db.relationshipTypes()` emits `type`, not `relationshipType`.
- **`YIELD *` is rejected on a procedure whose columns depend on what you passed in.** `explain` and `whisper.history` are both multi-shape. Name columns from one shape, or call the single-shape variant: `whisper.history.whois(domain)` for WHOIS columns, `whisper.history.bgp(ip|asn|prefix)` for routing columns.

Schema-introspection procedures are cheap and answer immediately, so they are the fastest way to confirm a label or edge exists before you anchor on it:

```cypher expect=rows>0 verified=2026-09-02
CALL db.labels() YIELD label RETURN label ORDER BY label LIMIT 20
```

```cypher expect=rows>0 verified=2026-09-02
CALL db.relationshipTypes() YIELD type, count RETURN type, count ORDER BY type LIMIT 5
```

## CALL subqueries

`CALL { ... }` scopes a subquery, importing outer variables with `WITH`. A standalone `CALL { ... }` with no preceding clause is not allowed — give it an importing clause. Bounding each branch inside its own `CALL {}` is the reliable way to run several per-branch aggregations from one anchor, and the subquery's own `LIMIT` bounds each branch on its own. It is also where a multi-hop routing leg belongs after a `WITH`: keep `ANNOUNCED_BY` and `ROUTES` together inside the subquery, or give each `WITH` stage one computed hop.

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH (a:ASN {name: "AS13335"})
CALL {
  WITH a
  MATCH (a)-[:ROUTES]->(p:ANNOUNCED_PREFIX)
  RETURN count(p) AS pc
}
RETURN a.name, pc
```

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})
CALL {
  WITH h
  MATCH (h)-[:RESOLVES_TO]->(ip:IPV4)
  RETURN ip LIMIT 2
}
RETURN h.name, ip.name
```

## EXISTS and COUNT subqueries

`EXISTS { ... }` tests whether a pattern has at least one match without binding it, and `NOT EXISTS { ... }` negates it. `COUNT { ... }` returns how many matches the pattern has, as an expression you can project or filter on.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})
WHERE EXISTS { MATCH (h)-[:RESOLVES_TO]->(:IPV4) }
RETURN h.name
```

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})
WHERE NOT EXISTS { MATCH (h)-[:SPF_EXISTS]->(:HOSTNAME) }
RETURN h.name
```

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})
RETURN h.name, COUNT { MATCH (h)-[:RESOLVES_TO]->(:IPV4) } AS ipv4_count
```

## List and pattern comprehensions

A list comprehension filters and maps a list inline: `[x IN list WHERE predicate | expression]`. A pattern comprehension does the same over a graph pattern, collecting one value per match, and it can be sliced like any list.

```cypher expect=rows>0 seed=low verified=2026-09-02
RETURN [x IN [1, 2, 3] WHERE x > 1 | x * 10] AS scaled
```

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH (a:ASN {name: "AS13335"})
RETURN a.name, [(a)-[:ROUTES]->(p) | p.name][0..3] AS prefixes
```

The slice runs after the list is built, so a pattern comprehension over a wide fan-out collects everything first. Bound the anchor set with `WITH ... LIMIT` before you comprehend over it.

## CASE

`CASE` is a conditional expression, in both the simple form (match a value) and the searched form (evaluate conditions).

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH (a:ASN {name: "AS13335"})
RETURN CASE a.name WHEN "AS13335" THEN "cloudflare" ELSE "other" END AS label
```

```cypher expect=rows>0 seed=low verified=2026-09-02
UNWIND [1, 5, 10] AS x
RETURN CASE WHEN x < 3 THEN "low" WHEN x < 8 THEN "mid" ELSE "high" END AS bucket
```

## Parameters

Write `$name` placeholders in the query and send the values in the request body's `parameters` object (the field is `parameters`, not `params`). Parameters keep the query plan cacheable and save you from escaping quotes inside JSON. They are available on `POST /api/query` only.

```cypher expect=static verified=2026-09-02
MATCH (h:HOSTNAME {name: $n}) RETURN h.name AS n
```

```json
{"query": "MATCH (h:HOSTNAME {name: $n}) RETURN h.name AS n", "parameters": {"n": "google.com"}}
```

A query that references `$n` without a matching entry in `parameters` is rejected with `400 query-error` and `Missing parameter: $n`; the placeholder is never treated as a literal. See [Parameter binding](/docs/cypher-api/reference/query-post#parameter-binding) for the request in five languages.

## Batching statements

Several statements separated by a top-level `;` in one `query` string run as a batch, and the response changes shape: instead of one `columns`/`rows` envelope you get a `results` array with one entry per statement.

```cypher expect=static verified=2026-09-02
RETURN 1 AS a; RETURN 2 AS b
```

Each entry carries the statement's own `result` (`columns`, `rows`, `rowCount`, `executionTimeMs`), an `outcome` (`OK`, `PARSE_ERROR`, `EXECUTION_ERROR`, or `DEADLINE_EXCEEDED`), and a `success` flag. A failed statement adds `errorMessage` and `errorType` and does not stop the others, so read `outcome` per entry rather than the HTTP status. Splitting happens on top-level semicolons only; one inside a string literal is left alone. Details in [Batch statements](/docs/cypher-api/reference/query-post#batch-statements).

## Patterns & paths

A pattern is a sequence of node and relationship descriptions. The forms:

- **Directed** — `(:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(ip)`
- **Reverse** — `(:HOSTNAME {name: "google.com"})<-[:MAIL_FOR]-(mx)` walks the edge backward (mail and nameserver edges point server → domain)
- **Undirected** — `(:ASN {name: "AS13335"})-[:BGP_NEIGHBOR]-(peer)` matches either arrow; use it for peering, which is symmetric in practice
- **Multi-type** — `(:HOSTNAME {name: "google.com"})-[:RESOLVES_TO|HAS_EMAIL]->(x)` matches either edge
- **Variable-length** — `(:HOSTNAME {name: "www.mail.google.com"})-[:CHILD_OF*1..3]->(parent)` walks one to three hops
- **Named paths** — bind the path with `p = (...)` to use `nodes()`, `relationships()`, and `length()`
- **Anonymous nodes** — `(:ASN {name: "AS13335"})-[:ROUTES]->()` when you don't need the far node bound

```cypher expect=rows>0 seed=github.com verified=2026-09-02
MATCH p = (h:HOSTNAME {name: "github.com"})-[:RESOLVES_TO]->(:IPV4)-[:ANNOUNCED_BY]->(:ANNOUNCED_PREFIX)
RETURN length(p) AS hops, [n IN nodes(p) | n.name] AS chain LIMIT 5
```

```cypher expect=rows>0 seed=google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "google.com"})-[r:RESOLVES_TO|HAS_EMAIL]->(x)
RETURN type(r) AS edge, x.name AS target LIMIT 5
```

When you expand outward from an announced prefix (`ANNOUNCED_PREFIX`, `REGISTERED_PREFIX`), give the relationship a type (`<-[:ROUTES]-`, `-[:CONFLICTS_WITH]->`) or start from the `PREFIX`-labelled node with the same name; an untyped, undirected `-[r]-` from a computed prefix is not a supported shape. On stored labels such as `HOSTNAME` the untyped form is fine.

### Variable-length relationships

Repeat an edge between a lower and an upper bound. Always write the upper bound yourself: a bare `[*]` is read as a very deep walk, which is almost never what you meant.

```cypher expect=rows>0 seed=www.mail.google.com verified=2026-09-02
MATCH (h:HOSTNAME {name: "www.mail.google.com"})-[:CHILD_OF*1..3]->(parent)
RETURN parent.name LIMIT 5
```

The chain ends at the TLD: `www.mail.google.com` → `mail.google.com` → `google.com` → `com`. A range wider than the label chain costs nothing extra, but it also finds nothing extra.

Edges computed at query time (`BGP_NEIGHBOR`, `ROUTES`, `ANNOUNCED_BY`, `LISTED_IN`, `BELONGS_TO`, `CONFLICTS_WITH`) work inside `[*1..N]` as long as one endpoint is anchored, and `length(p)` reports correctly. Keep the range tight, and on a peering walk filter `WHERE n <> a`: a mesh routinely walks back to the origin AS.

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
MATCH p = (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR*1..2]-(n:ASN)
WHERE n <> a
RETURN n.name AS asn, length(p) AS distance LIMIT 5
```

Over a high-fan-out edge, prefer explicit single hops joined with `WITH ... LIMIT`: the variable-length form expands every intermediate row before it returns anything. See [Best Practices](/docs/cypher/best-practices).

### shortestPath

`shortestPath` finds the minimum-hop path between two anchored nodes. Write the variable-length range explicitly and keep it tight; both endpoints must be anchored, and when no path exists within the bound the result is empty rather than an error.

```cypher expect=rows>0 seed=www.google.com verified=2026-09-02
MATCH p = shortestPath(
  (a:HOSTNAME {name: "www.google.com"})-[*1..4]-(b:HOSTNAME {name: "google.com"})
)
RETURN length(p) AS hops
```

## EXPLAIN

`EXPLAIN` returns the query plan without executing it, as a single `plan` column holding the operator tree. Use it to confirm an anchored lookup hits the index rather than scanning.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
EXPLAIN MATCH (h:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(ip) RETURN ip.name
```

A `NodeLookup` at the leaf of the plan means the anchor is indexed. A bare label scan there is a warning that the query will be slow — anchor it before you run it.

### PROFILE

`PROFILE` runs the query and returns the same rendered `plan` together with the `rows` it produced and its `executionTimeMs`, so you can see what a query costs before you put it in a loop.

```cypher expect=rows>0 seed=google.com verified=2026-09-02
PROFILE MATCH (h:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(ip) RETURN ip.name LIMIT 1
```
