# iq jq for NoSQL databases. Every page of https://zsltg.github.io/iq/ concatenated in navigation order, from the Markdown sources under docs/docs/. # Get started docs/docs/index.md # Get started `iq` runs `jq`[^2] filters to query, dump, copy, diff and write data across NoSQL databases, and their dump files, from a single static binary. See [Drivers](drivers.md#drivers) for supported databases.

iq registers a MongoDB source, reads one document by key, filters a scan with a pushed-down predicate, explains the plan, and prints the result as gron

`iq` normalizes fetched values to JSON. The filter runs entirely client-side, so one filter means the same thing everywhere. The URI scheme[^4] chooses the backend. The filter is both the *transform* and the *key selector*. The selector walks the parsed `jq` AST[^ast]. Based on the AST, the selector executes a *bounded read*, a *streaming scan* (with *pushdown*[^pushdown]), or a *materialized scan* (see [How it works](how-it-works.md#how-it-works)). Typed dumps carry native types across stores. As a result, a copy, a restore or a migration is one command instead of an export plus a conversion script. `iq` is inspired by [`sq`](https://sq.io "Command-line tool giving jq-style access to SQL databases and files like CSV or Excel"), whose command set it deliberately follows. !!! note `iq` is built with AI assistance, and every change passes the full test suite, container-backed integration tests for every backend, and a mutation gate before it lands (see [CONTRIBUTING.md](https://github.com/zsltg/iq/blob/main/CONTRIBUTING.md)). Queries are read-only. `--insert`, `--replace`, `iq data clear`, `iq data drop` and `iq data delete` write to the target. `iq exec` forwards a native command to the database, so it can write too. Use `--explain` to see the [query plan](query-plan.md#query-plan) or `--dry-run` to report the effect of a write, without changing anything. `iq exec` has no dry run. Feedback and bug reports are very welcome. Report security problems privately, see the [security policy](https://github.com/zsltg/iq/security/policy). ## Installation `iq` ships as a single static binary (no runtime dependencies, no CGO[^5]). === ":fontawesome-brands-linux: Linux" ```sh curl -fsSL https://raw.githubusercontent.com/zsltg/iq/main/install.sh | sh ``` !!! note "Version & Location" The script downloads the release for your OS/arch, verifies its SHA-256[^6] against the release checksums, and installs the binary. `IQ_VERSION` pins a version and `IQ_INSTALL_DIR` picks the target directory. You can also download a `.deb`, `.rpm`, `.apk`, or Arch `.pkg.tar.zst` from the [releases](https://github.com/zsltg/iq/releases). === ":fontawesome-brands-apple: macOS" ```sh brew install zsltg/tap/iq ``` The Linux install script above works on macOS too. === ":fontawesome-brands-windows: Windows" ```powershell scoop bucket add zsltg https://github.com/zsltg/scoop-bucket scoop install iq ``` === ":fontawesome-brands-golang: Go" ```sh go install github.com/zsltg/iq@latest ``` ### Verify a release Every release signs `checksums.txt` with a keyless [cosign](https://docs.sigstore.dev/) signature from the release workflow, and carries SLSA build provenance for every artifact (`multiple.intoto.jsonl`). The install script checks the SHA-256 of the archive it downloads. To check a download yourself, first make sure that `checksums.txt` comes from the iq release workflow: ```sh cosign verify-blob checksums.txt \ --bundle checksums.txt.sigstore.json \ --certificate-identity-regexp '^https://github.com/zsltg/iq/\.github/workflows/release\.yml@refs/tags/v' \ --certificate-oidc-issuer https://token.actions.githubusercontent.com sha256sum --ignore-missing -c checksums.txt ``` To check the build provenance of an archive, use [slsa-verifier](https://github.com/slsa-framework/slsa-verifier): ```sh slsa-verifier verify-artifact iq_0.37.1_linux_amd64.tar.gz \ --provenance-path multiple.intoto.jsonl \ --source-uri github.com/zsltg/iq --source-tag v0.37.1 ``` Releases before v0.37.1 carry `checksums.txt` only, with no signature or provenance. ## Building from source ```sh git clone https://github.com/zsltg/iq cd iq && make build ``` ## The basics Register a source for each store you work with. Then make one of them active: ```sh iq add -n orders 'mongodb://localhost:27017/shop?collection=orders' iq add -n staging 'mongodb://staging:27017/shop?collection=orders' iq add -n cache redis://localhost:6379/0 iq add -n snap file:///backups/prod.rdb iq src orders ``` ```sh { title='Check the list of sources you added' } iq ls ``` ```sh { title='Inspect the active source' } iq inspect ``` ```sh { title='Explain the query plan for a bounded read, a dry run' } iq '.["o-42"]' --explain -v ``` ```sh { title='Run the query to get the document with id "o-42"' } iq '.["o-42"]' ``` ```sh { title='Run a query to get all documents in batches, a streaming scan' } iq '.[]' ``` ```sh { title='Stream a filtered sample, pushed to the server where it can be' } iq --src orders '.[] | select(.status == "new") | {id, total}' ``` ```sh { title='Query a backup without restoring it' } iq --src snap '.[] | select(.active)' ``` ```sh { title='Copy one store into another, native types intact' } iq --src orders --insert cache ``` ```sh { title='Diff two environments, data or inferred schema' } iq diff orders staging --schema ``` You can find detailed examples in [Sources](sources.md#sources), [Query data](query-data.md#query-data) and [Write data](write-data.md#write-data). For more advanced usage check [Output](output.md#output), [Query plan](query-plan.md#query-plan), [Cookbook](cookbook.md#cookbook) and [Loading exports](loading-exports.md#loading-exports). For debugging, see [Diagnostics & Logging](diagnostics-and-logging.md#diagnostics-logging). Supported data sources are listed in [Drivers](drivers.md#drivers). ## Shell completions The `.deb`, `.rpm`, `.apk` and `.pkg.tar.zst` packages install [Bash](https://tiswww.case.edu/php/chet/bash/bashtop.html "GNU Bourne-Again SHell, the default shell on most Linux distributions"), [Zsh](https://www.zsh.org/ "Extended Bourne shell, the default on macOS since Catalina") and [fish](https://fishshell.com/ "Friendly Interactive SHell, deliberately non-POSIX, with autosuggestions built in") completions for you. For a [brew](https://brew.sh/ "Homebrew, the third-party package manager for macOS and Linux"), [scoop](https://scoop.sh/ "Command-line installer for Windows, installing per-user without admin rights"), [go-install](https://go.dev/ref/mod#go-install "Builds and installs a Go command from its module path into GOPATH/bin") or source build, `iq completion ` prints a script to install by hand. === ":simple-gnubash: Bash" ```sh title="load in the current session" eval "$(iq completion bash)" ``` ```sh title="copy it to the completion path" iq completion bash | sudo tee /usr/share/bash-completion/completions/iq >/dev/null ``` === ":simple-zsh: Zsh" ```sh title="write to a directory on your $fpath, then restart the shell" iq completion zsh > ~/.zsh/completions/_iq ``` === ":simple-fishshell: fish" ```sh title="write to a directory on your $fish_complete_path" iq completion fish > ~/.config/fish/completions/iq.fish ``` === ":material-powershell: PowerShell" ```sh title="append to your profile" iq completion powershell >> $PROFILE ``` Completions cover the commands, their sub-subcommands and flags. They also cover the saved source handles, groups, and config-option keys, which they read live from your config. As a result, `iq --src ` offers the sources `iq ls` lists. A flag that takes a closed set offers that option's own values. `iq inspect --only ` and `iq diff --section ` offer the introspection subcommands of the selected source's backend, worked out from its saved URI. The `jq` filter itself is a program, not a completable value. As a result, `iq` offers no candidates there (and never falls back to filenames). `iq` also offers no candidates for the backend verb of `iq exec` and its operands. Every completion is offline. It reads your config file and nothing else. As a result, a `` never opens a connection, never reads the OS keyring[^keyring], and cannot hang. That is why a collection suffix does not complete. `iq --src shop.` offers nothing, because listing collections needs a connection. ## Man page The `.deb`, `.rpm`, `.apk`, and `.pkg.tar.zst` packages also install an `iq(1)` manual page, so `man iq` works after a package install. For any other install, pipe it into your man path: ```sh iq man | sudo tee /usr/share/man/man1/iq.1 >/dev/null ``` [^2]: `jq` is a widely-used command-line utility and very high-level, functional, domain-specific programming language designed for processing JSON data. https://jqlang.org [^4]: RFC3986 proposes a generic URI syntax and a process for resolving URI references that might be in relative form, along with guidelines and security considerations for the use of URIs on the Internet. https://datatracker.ietf.org/doc/html/rfc3986 [^5]: Cgo enables the creation of Go packages that call C code. https://pkg.go.dev/cmd/cgo [^6]: SHA-256 is a Secure Hash Algorithm with a message digest size of 256. https://nvlpubs.nist.gov/nistpubs/fips/nist.fips.180-4.pdf [^ast]: An abstract syntax tree is the tree that a parser builds from the source of a program. Here, it is the parsed `jq` filter that the key selector inspects to decide how to read (see [How it works](how-it-works.md#read-strategies)). https://en.wikipedia.org/wiki/Abstract_syntax_tree [^pushdown]: Predicate pushdown hands part of the filter to the database, so that the database returns only matching items instead of everything for client-side filtering. Each driver page lists what the driver can push (see [Drivers](drivers.md#drivers)). https://en.wikipedia.org/wiki/Predicate_pushdown [^keyring]: The operating system's credential store (macOS Keychain, Windows Credential Manager, the Secret Service on Linux), where `--store keyring` sources keep their secrets (see [Configuration](configuration.md#keyring-keyring)). https://pkg.go.dev/github.com/zalando/go-keyring # How it works docs/docs/how-it-works.md # How it works A Go command-line tool that runs [jq](https://jqlang.github.io/jq/) filters against NoSQL databases. The URI scheme chooses the backend. The query core is driver-agnostic, so further backends slot in behind the same port. The filter is both the transform and the key selector. Its top-level paths name the keys to fetch, so a normal query reads a bounded set of keys. A `.[]`-rooted filter streams the keyspace in pages. A filter that collapses the keyspace into one value materializes only behind `--unbounded`. `iq` normalizes fetched values to JSON. The filter then runs entirely client-side, so its semantics are identical for every backend. The CLI and `iq mcp` are two thin delivery mechanisms over that one core. The MCP server exposes the CLI's own operations as tools, resolves the same saved sources, and runs the same engine. As a result, it adds no port and changes no classification. It adds only its own bounds, a tool set fixed at startup by `--allow`, per-result item and byte caps, and the CLI's redacted error shape. `iq` is inspired by [sq](https://github.com/neilotoole/sq). Much of its command set (the `.` addressing along with many subcommands and flags) deliberately follows sq's to make the tool feel familiar. The `jq` semantics are identical for any future backend. [gojq](https://github.com/itchyny/gojq) (pure Go, no CGO) provides `jq`, keeps `iq` a single static binary, and exposes the AST the key selector walks. ## Read strategies The shape of the filter decides how much `iq` reads. Every filter takes one of three routes: 1. **Bounded reads** (with explicit keys) only read the specified subset of items from the source. The keys you asked for bound the cost, never the size of the database. 2. **Streaming scans** (a filter rooted at `.[]`, for example `.[]`, `.[] | select()`, `.[].title`) process each value independently. They walk the keyspace in pages, run the filter page by page, and emit as they go. Memory stays constant and results appear progressively. 3. **Materialized scans** (a filter that collapses the collection into one value, `.`, `keys`, `length`, `map()`, `group_by`, `sort_by`, aggregates) read all items into memory in batches before applying the filter. ```mermaid graph LR Q1[".[#quot;1#quot;]"] -->|names a key| T1["bounded read"] --> R1["{ #quot;title#quot;: #quot;The Go…#quot; }
one value"] Q2[".[]"] -->|iterates values| T2["streaming scan"] --> R2["{ … } then { … } then …
each value, streamed"] Q3["."] -->|whole root| T3["materialized scan
(needs --unbounded)"] --> R3["{ #quot;1#quot;: {…}, #quot;2#quot;: {…} }
one object, every key"] classDef bounded fill:#e6f4ea,stroke:#137333,color:#0b3d1f; classDef streaming fill:#fef7e0,stroke:#8a5a00,color:#5c3d00; classDef materialized fill:#fce8e6,stroke:#c5221f,color:#5c0f0a; class T1 bounded; class T2 streaming; class T3 materialized; ``` On a streaming scan, **pushdown** compiles what it can of the filter's `select()` into a backend-neutral predicate. Pushdown then hands the predicate to the driver. This shrinks how much data is transferred or decoded. The predicate is deliberately weaker than the filter, so the engine re-runs the full filter per page to drop the extra matches it admits. !!! tip "Check the strategy of a query" Before you execute a query, use `--explain` to see which strategy `iq` will use for it. !!! note "Cost" A *bounded filter* runs client-side over only the named keys, so its cost is `O(keys requested)`. A *streamable scan* runs in `O(page)` memory. !!! note "Unbounded queries" `--unbounded` means "permit loading the whole dataset into memory". The flag names the cost property (loading everything), not any one store's mechanism, so it will mean the same thing for all backends. You can also pass it on a streaming filter. It then switches that filter from batched streaming to a single materialized pass. The result is key-sorted output and a consistent snapshot instead of scan order. !!! note "Streaming data" Streamed output is *best-effort*. Values arrive in scan order (not key-sorted). If the keyspace is resized mid-scan, an element can repeat. This is the price of never holding more than one page. If you need sorted, exactly-once output, use `--unbounded`. !!! note "Progress" A scan has no reliable upfront total (for example Redis `SCAN`, Mongo cursor). For this reason, `iq` shows an animated spinner with a running `N scanned` count on *stderr*. As a result, a sparse `.[] | select()` over a large keyspace is never silent. When a backend can supply a cheap approximate total (for example MongoDB's `estimatedDocumentCount` for an unfiltered whole-collection scan, Redis's `DBSIZE` for its whole-keyspace `MATCH *` scan), `iq` shows the count against that total as `N scanned (~M est)`. The tilde marks the estimate as a hint. The estimate comes from cached metadata and drifts under concurrent writes. As a result, the scan can exceed the estimate, and the estimate never becomes a percentage bar. `iq` shows no total for a pushed-down filtered scan (it walks a subset) or for a cross-source scan (a per-source estimate misleads the aggregate). ## Query routes The **selector** classifies a jq filter. `iq` optionally **decomposes** a scan into a native predicate. Each backend then maps that predicate its own way. Regardless of pushdown, the full jq re-runs client-side. As a result, the pushed predicate is only ever a conservative pre-filter, and results are identical with or without it. ```mermaid graph TD F["jq filter (CLI)"] --> SEL["selector.Keys
(AST analysis)"] SEL -->|bounded| GET["KVStore.Get(keys)"] SEL -->|"streamable scan"| CMP{"FilteredScanner?
(pushdown on)"} SEL -->|"holistic scan"| MAT["materialize
(--unbounded)"] CMP -->|yes| PD["pushdown.Compile → predicate.Node
(adapter pre-filters server-side)"] CMP -->|no| RS["KVStore.ScanBatches
(full scan, no pushdown)"] GET --> JQ["run full jq
client-side, per batch"] MAT --> JQ PD --> JQ RS --> JQ JQ --> OUT["format renderer
→ output"] ``` A scan emits per-page progress to a stderr spinner (CLI only, off unless attached to a terminal). An unfiltered scan can also fetch a cheap up-front total estimate (where the backend metadata makes it possible). ## Write routes Writes ride the query command. There is no separate copy tool. Items arrive from `--src` (a live source or a `file://` dump) or piped stdin. `iq` reads each item as a typed record, so the native type survives the trip. The `jq` filter transforms each item with its key preserved. This transform is the one place where a filter runs per item rather than over the whole keyspace. Iteration is implicit, and you do not write `.[]`. ```mermaid graph TD MV["iq --insert / --typed (CLI)"] --> MSRC["source: --src (live or file:// dump) / stdin → TypedScan"] MSRC --> TX["per-item jq transform + re-key"] TX --> DST{"--insert or --typed?"} DST -->|--insert| PUT["Putter.Put (upsert / insert-only)"] DST -->|--typed| ENC["emit {key,type,value} → jsonl / json / jsona / yaml"] PUT --> BW["backend adapter:
type-aware native writes"] LF["iq data clear / drop / delete (CLI)"] --> CAP["Clearer.Clear / Dropper.Drop / Deleter.Delete (capability-gated)"] ``` The transformed stream then takes one of two exits. - **`--insert `** hands each record to the destination's `Putter`. Existing keys are overwritten (upsert) unless `--no-overwrite` makes the run insert-only. `--replace` empties the destination first (with confirmation, or `--force`). The backend adapter translates each typed value into its native write, so a Redis hash lands as a hash again, not as a JSON blob. - **`--typed`** skips the store and emits the same records as `{"key", "type":…, "value":…}` envelopes in the chosen format (`--jsonl` by default). The dump re-imports through a `file://` source or a piped `--insert`, which closes the loop back into the diagram's source node. Destructive operations are a separate entry point, not a filter outcome. The `iq data clear / drop / delete` subcommands call their own capability-gated ports. As a result, a backend that has no native drop refuses rather than emulating one. For the same reason, no query ever deletes as a side effect. See the full flag tables and key-mapping rules in [Write data](write-data.md). ## Null vs. Missing Modern query standards treat an *absent* field and an explicit *null* as distinct values. [PartiQL](https://partiql.org/) has both `NULL` and `MISSING`, and `MISSING` drops out of a projection where `NULL` is carried through. [SQL++](https://arxiv.org/abs/1405.3631) makes missing a value of its own and documents real divergence. The same path returns `null` in AsterixDB, `missing` in Couchbase, and an error in SQL. [RFC 9535](https://www.rfc-editor.org/rfc/rfc9535) (JSONPath) models absence as `Nothing`, again distinct from `null`. `iq` takes a deliberate, layered position in that vocabulary rather than one blanket rule. - **The jq layer reads missing as `null`.** gojq evaluates entirely client-side. As a result, `.a` on a document without `a` yields `null`, the same on every backend. This is the uniform semantics that the whole tool promises. A filter behaves identically whether the field is absent, stored as `null`, or the source has no such field at all. - **The schema layer preserves the distinction.** `iq schema` tracks *parent-relative presence*. A field observed on some documents but not others is optional, separate from a field that is present-and-nullable. As a result, the inferred shape measures logical structure, not the jq layer's collapse. - **Drivers decline pushes whose backend semantics diverge.** A conjunct is pushed only when the backend reproduces `jq`'s answer for every input including missing and null. Elasticsearch `== null` (which no single term matches as absent-or-null) is declined and re-filtered client-side rather than pushed with the wrong meaning. The per-driver push/not-push tables record each call. So the collapse is a query-layer convenience, not a loss. `iq` keeps the distinction where it carries information (schema inference, pushdown safety). It hides the distinction where uniformity matters more (the query layer). # Sources docs/docs/sources.md # Sources `iq` connects only through **saved sources**, a named connection you register once, then select by name or as the default. !!! note "Configuration" Sources live in a TOML file at `/iq/iq.toml` (for example `~/.config/iq/iq.toml`), written `0600` because a URI can carry a password. Override the path with `IQ_CONFIG`, or per run with the global `--config ` flag (which wins over `IQ_CONFIG`). A source added with `--store keyring` keeps no password in this file. The password lives in the OS keyring (Secret Service on Linux, Keychain on macOS, Credential Manager on Windows). `iq` splices it back into the URI only when connecting. ## Add `add` `iq add [flags]` Register a source from a connection URI. The URI is the only positional argument. Each driver supports a different set of URI parameters (see [Drivers](drivers.md)). | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | `-a` | `--active` | ✗ | make the new source the active source | | `-d ` | `--driver ` | auto-detect | expected backend driver. It must match the URI scheme | | `-n ` | `--handle ` | the keyspace the URI names | handle for the source, derived from the URI when omitted | | `-p` | `--password` | ✗ | prompt for the URI password or read it from stdin | | | `--skip-verify` | ✗ | skip the post-add reachability check | | | `--store ` | `inline` | where the URI's password is kept, `inline` (in the config file) or `keyring` (OS keyring) | ```sh { title='Add an inactive MongoDB source, defaults to handle "books"' } iq add mongodb://localhost:27017/books ``` ```sh { title='Add an inactive Redis source, sets the handle to "cache"' } iq add -n cache redis://localhost:6379/0 ``` ```sh { title='Add a Cassandra source and make it active, defaults to handle "orders"' } iq add -a 'cassandra://localhost:9042/shop?table=orders' ``` ```sh { title='Add an inactive MongoDB source, prompts for the password for a user named "iq" in the "iq" database, defaults to handle "books"' } iq add -p 'mongodb://iq@localhost:27018/iq?collection=books' ``` ```sh { title='Add an inactive MongoDB source, prompts for the password for a user named "root" in the "admin" database, defaults to handle "books"' } iq add -p 'mongodb://root@localhost:27018/iq?authSource=admin&collection=books' ``` !!! tip "Escaping the source URI" If the URI contains `?` (and `&`), your shell can interpret it as a wildcard, operator or separator. As a result, you must escape it (for example `iq add 'mongodb://localhost:27017/iq?collection=books'`). !!! tip ""add" shadowing" `iq add` shadows `jq`'s built-in `add` filter at the top level. To sum with `jq`, write it inside a larger expression, for example `iq '[ .a, .b ] | add'`. ## Group `group` `iq group [name] [flags]` Show, set, or clear the active group. A `/` in a name groups sources (`prod/books`, `dev/books`). Set an active group with `iq group prod`. An unqualified name then resolves inside it. For example, `iq src books` selects `prod/books`. If the group has none, `iq src books` falls back to a top-level `books`. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--clear` | ✗ | clear the active group | ```sh { title='Show the active group' } iq group ``` ```sh { title='Set the active group to "dev"' } iq group dev ``` ```sh { title='Clear the active group' } iq group --clear ``` !!! tip "Listing groups" You can list the already created groups with `iq ls -g`. ## List `ls` `iq ls [group] [flags]` List saved sources, the active one marked with `*`. An optional `[group]` limits the listing to sources in that group. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | `-v` | `--verbose` | ✗ | show a header and a `FORMAT` (a file source's detected dump format) and an `OPTIONS` (the source's stored option defaults) column :material-earth:{ title="Global flag" } | | `-g` | `--group` | ✗ | lists groups instead of sources | | `-j` | `--json` | ✗ | emit machine-readable JSON output | | `-y` | `--yaml` | ✗ | emit machine-readable YAML output | | | `--reveal` | ✗ | prints a password stored inline in the config verbatim | | | `--expand` | ✗ | resolves a keyring-backed source's stored password and inlines it (combine with `--reveal` to print verbatim) | ```sh { title='List saved sources' } iq ls ``` ```sh { title='List groups' } iq ls -g ``` ```sh { title='List saved sources, the saved options (if there is any) and the format for file sources' } iq ls -v ``` ```sh { title='List saved sources, prints passwords verbatim' } iq ls --reveal ``` ## Move `mv` `iq mv [flags]` Rename a source to a new full handle, or move it into a group by giving a group-qualified target (`iq mv books prod/books`). When `` names a group, every source under it is re-prefixed (`iq mv prod staging` renames `prod/*` to `staging/*`). The active source and group follow the move. A keyring-backed source's stored credential moves with it. ```sh { title='Rename source named "shop" to "catalog"' } iq mv shop catalog ``` ```sh { title='Move source named "books" into the group "prod"' } iq mv books prod/books ``` ```sh { title='Move every source in the group named "prod" to the group named "staging"' } iq mv prod staging ``` ## Ping `ping` `iq ping [name...] [flags]` Open each source and round-trip a cheap command (for example Redis `PING` or MongoDB `{ping:1}`), reporting its driver and the round-trip time, or the error. With no arguments, `iq ping` pings the active source. Otherwise, each argument is a source handle or a group (pinging every member). `--all` pings every saved source and takes no arguments. `--timeout` bounds each check. Exits with a non-zero exit code if any source is unreachable. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--all` | ✗ | ping every saved source, rejects source arguments | | | `--timeout` | `5s` | per-query timeout :material-earth:{ title="Global flag" } | ```sh { title='Ping the active source' } iq ping ``` ```sh { title='Ping the sources named "cache" and "shop"' } iq ping cache shop ``` ```sh { title='Ping the sources in the group named "dev"' } iq ping dev ``` ```sh { title='Ping every saved source' } iq ping --all ``` ## Remove `rm` `iq rm ... [flags]` Remove saved sources or whole groups. Each argument is a source handle or a group name (which removes every source under it). The removal is atomic. If any argument names neither a source nor a group, nothing is removed. A keyring-backed source's stored credential is deleted too. ```sh { title='Remove a single source named "cache"' } iq rm cache ``` ```sh { title='Remove multiple sources named "cache", "shop" and "dev/books"' } iq rm cache shop dev/books ``` ```sh { title='Remove all sources in the group named "dev"' } iq rm dev ``` ## Show/Set `src` `iq src [name] [flags]` Show or set the active source. `--src` :material-earth:{ title="Global flag" } is global. Every command accepts it (see [Global flags](global-flags.md#global-flags)). Once a source is active, every query runs against it. Select a different source for a single command with `--src`/`-s`, without changing the active one. Address a MongoDB collection or a Cassandra table with a dotted `handle.collection` / `handle.table` suffix. With no active source and no `--src`, the command errors. There is no ambient URI or environment fallback. ```sh { title='Show the active source' } iq src ``` ```sh { title='Set the active source to "shop"' } iq src shop ``` !!! tip "Per command source selection" You can run commands against a specific source which can be different from the active one, for example `iq --src books '.["2"]'`. ## Inspect `inspect` `iq inspect [source] [flags]` Show a source's native server/database introspection. The positional argument names the source, like `iq inspect books`. With none, it uses --src or the active source. Certain sources accept `sq`-style addressing (for example MongoDB `.` or Cassandra `.`) to pick the keyspace, overriding the default in the source's URI (for example `?collection=` or `?table=`). | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--list` | ✗ | list the subcommands/sections available for the source | | | `--only ` | ✗ | narrow to these sections or subcommands | | | `--reveal` | ✗ | print an inline-stored password verbatim in the location header instead of redacting it | | | `--expand` | ✗ | resolves a keyring-backed source's stored password and inlines it (combine with `--reveal` to print verbatim) | | `-j` | `--json` | ✗ | emit machine-readable JSON | | `-y` | `--yaml` | ✗ | emit machine-readable YAML | ```sh { title='Inspect the active source on the database level' } iq inspect ``` ```sh { title='Inspect a named source on the database level' } iq inspect shop ``` ```sh { title='Inspect a named source on the collection level, output in JSON' } iq inspect shop.orders -j ``` ```sh { title='Inspect a named source on the database level, narrow the sections to "dbStats" and "serverStatus"' } iq inspect shop --only dbStats,serverStatus ``` ## Diff `diff` `iq diff [=] [=] [flags]` Compare two saved sources across the layers a schemaless store can meaningfully compare. Selecting no layer defaults to `--data`. Layers combine. Each side can carry a `jq` filter, so a diff can be scoped to part of a keyspace. Give it per side as `=`, or for both sides at once with `--filter`. A side's own filter takes precedence. Exits non-zero when the sources differ and zero when they match (diff(1)-style[^2]), so scripts can branch on the exit status. Bounded by `--timeout`. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--data` | ✓ | diff items key by key (cross-driver allowed, for example MongoDB `_id` and Redis key) | | | `--filter ` | none | jq filter you root at `.[]`, scoping both sides. A spec's own `source=` overrides it for that side | | | `--patch` | ✗ | emit an RFC 6902 JSON Patch[^1] that transforms the left source into the right (single layer only). For `--data`, the pointers read `//` over the whole keyspace map. `--patch` excludes `--json`/`--yaml`/`--set-arrays` | | | `--schema` | ✗ | diff an inferred field/type shape (cross-driver allowed) based on a sample (use `--sample` to change the sample size) | | | `--sample ` | `1000` | max items sampled per side for --schema (0 = all) | | | `--stats` | ✗ | diff native introspection trees (same driver only). Use with `--section` to narrow it down | | | `--section ` | full set | introspection section(s) for `--stats` (comma-separated or repeatable) | | | `--set-arrays` | ✗ | compare arrays order-insensitively as multisets (duplicates counted). Membership deltas are reported at the array's own path with no index segment. A pure reorder becomes no difference | | `-j` | `--json` | ✗ | emit machine-readable delta in JSON | | `-y` | `--yaml` | ✗ | emit machine-readable delta in YAML | ```sh { title='Diff the "prod" and "staging" source with a filter that applies to both sides' } iq diff prod staging --filter '.[] | select(.status == "new")' ``` ```sh { title='Diff the "prod" and "staging" source with a per-side filter' } iq diff 'prod=.[] | select(.type == "order")' 'staging=.[] | select(.kind == "ORDER")' ``` ```sh { title='Diff a single key from the "prod" and "staging" source' } iq diff 'prod=.["orders:42"]' 'staging=.["orders:42"]' ``` !!! tip "Filtering" Write the filter `.[]`-rooted, as on a plain query (iteration is not implicit here, unlike `--insert`). There are two reasons. A keyed diff needs each item to keep its key. A pushable `select(...)` narrows the read at the backend rather than merely narrowing the report. `iq diff prod staging --filter '.[] | select(.status == "new")'` A filter can instead name a single key (`prod=.["orders:42"]`) to compare one document. An absent key then reports as removed rather than as a change to null. A filter that collapses the keyspace (`keys`, `map(...)`) or fans one item out into several values (`.[] | .tags[]`) is refused. Neither leaves a key to match on. !!! warning "Data diff" `--data` without filtering reads both keyspaces fully into memory. As a result, it costs memory proportional to the two sources. This is a deliberate tradeoff, because an added/removed diff needs both key sets at once. A filter narrows that read, and a pushable one narrows it at the backend. Cross driver can be useful for verifying a migration. But the identity match is only as meaningful as the keys lining up. It is a power-user tool, not a schema comparison. !!! warning "Schema diff" `--schema` used across drivers gives each field path (with `[]` array-element and `{}` map-value wildcards) a canonical type (`integer`/`number`/`string(fmt)`/`map`/`array`) and a `required`/`optional` presence. As a result, it measures logical shape rather than sampling luck and stays quiet under resampling. The shape is sampled (with size set with `--sample`) and inferred, never declared, so a wider sample yields a truer shape. Two backends that genuinely normalize a native type differently (a timestamp as an RFC3339[^3] string vs an epoch number) still diff. That is the JSON each serves back. The format tags make the row legible rather than mysterious. Arrays are aligned by a longest common subsequence, so a single insertion reports one addition rather than a cascade at every later index. ## Schema `schema` `iq schema [source[=]] [flags]` Sample a source and project a schema inferred from its values (the same inference [`diff --schema`](#diff-diff) uses). Unlike [`inspect`](#inspect-inspect), which shows a backend's native introspection, `schema` infers a driver-agnostic shape, never declared. The field/type structure is sampled (with a sample size defined with `--sample`) from the values themselves. As a result, a wider sample yields a truer shape. It describes values, not keys. A non-object keyspace is legal (a string keyspace emits `{"type":"string"}`). A filter that fans one item out into several values is fine here, even though a keyed [`diff`](#diff-diff) must refuse it. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--filter ` | none | jq filter scoping which items the shape is inferred from. The spec form `source=` sets it per source | | | `--format ` | `jsonschema` | picks the contract dialect the shape projects into: JSON Schema draft 2020-12[^4] (`jsonschema`) or Open Data Contract Standard v3.1.0[^5] (`odcs`) | | | `--sample ` | `1000` | max items sampled (0 = all) | | `-y` | `--yaml` | ✗ | emit YAML instead of JSON (`jsonschema` only, `odcs` is always YAML) | ```sh { title='Schema of the active source' } iq schema ``` ```sh { title='Schema of the keyspace "orders" in the source "shop"' } iq schema shop.orders ``` ```sh { title='Schema of the source "dev" using a larger sample size' } iq schema dev --sample 5000 ``` ```sh { title='Schema of a filtered subset in the keyspace "orders" in the source "dev"' } iq schema 'dev.orders=.[] | select(.active)' ``` ```sh { title='Save schema into a file from keyspace "orders" from the source "dev"' } iq schema dev.orders > orders.schema.json ``` ```sh { title='Save schema into a file from keyspace "orders" from the source "dev" in ODCS format' } iq schema dev.orders --format odcs > orders.odcs.yaml # emit an ODCS v3.1.0 contract ``` !!! tip "Scoped schema" A `source=` spec (or `--filter`) scopes which items the shape is inferred from. For example, `iq schema 'prod=.[] | select(.active)'` describes only the active ones. The sample cap then applies to the survivors, so a selective filter walks further into the keyspace to fill it. !!! note "JSON Schema" JSON Schema draft 2020-12 (`jsonschema`, the default) can be used for interop with code generators, like [`quicktype`](https://github.com/glideapps/quicktype). `iq schema prod.orders > orders.schema.json && quicktype -s schema orders.schema.json -l go` !!! note "Open Data Contract Standard" Open Data Contract Standard v3.1.0 (`odcs`) is a YAML data contract for tools like [`datacontract-cli`](https://github.com/datacontract/datacontract-cli), [Soda](https://github.com/sodadata/soda-core) and [Great Expectations](https://greatexpectations.io/). It carries nine logical types with no binary or decimal member. As a result, `iq`'s date-time and date formats demote to `logicalType: date` with a JDK format pattern. A UUID stays `string` with the `uuid` format. Base64-binary and exact-decimal values stay plain `string`. The contract's identifiers (`id`, `name`, schema-object name) derive deterministically from the source handle and keyspace. They use no timestamps or random ids. [^1]: JSON Patch defines a JSON document structure for expressing a sequence of operations to apply to a JavaScript Object Notation (JSON) document. It is suitable for use with the HTTP PATCH method. The "application/json-patch+json" media type is used to identify such patch documents. https://datatracker.ietf.org/doc/html/rfc6902 [^2]: `diff(1)` is a classic Unix command-line utility that compares two files (or directories) line by line and outputs the differences between them. https://pubs.opengroup.org/onlinepubs/9699919799/utilities/diff.html [^3]: Date and time format for use in Internet protocols that is a profile of the ISO 8601 standard for representation of dates and times using the Gregorian calendar https://datatracker.ietf.org/doc/html/rfc3339 [^4]: JSON Schema is a declarative language for defining structure and constraints for JSON data. https://json-schema.org/draft/2020-12 [^5]: The Open Data Contract Standard (ODCS) is an open-source, vendor-neutral specification (maintained under the Linux Foundation's LF AI & Data) for defining data contracts in machine-readable format (YAML or JSON). https://bitol-io.github.io/open-data-contract-standard # Query data docs/docs/query-data.md # Query data The default action, a `jq` filter run against the active source (see [Sources](sources.md)). The top-level paths name the keys to fetch. The result is printed as pretty JSON by default (see [Output formats](output.md)) ```sh { title='Fetch the key "greeting"' } iq '.greeting' ``` ```sh { title='Fetch the key "book:1" (a key containing a colon needs bracket-quoting)' } iq '.["book:1"]' ``` ```sh { title='Fetch the key "book:1" and extract one field' } iq '.["book:1"].title' ``` ```sh { title='Fetch keys "book:1", "book:2" and project a field from each' } iq '[ .["book:1"].title, .["book:2"].title ]' ``` ```sh { title='Values are strings, convert before arithmetic calculations' } iq '.["book:2"].price | tonumber + 5' ``` !!! warning "Escaping" Always wrap the filter in single quotes. `jq` syntax is full of characters that the shell otherwise expands or splits: brackets (`[ ]`), whitespace, `|`, `*`, `$`. Bracket-quoting a colon key like `.["book:1"]` reads as a glob to `zsh` (`no matches found`) or `bash` unless quoted. ## Cross-source queries Both [Compose](#compose-source) and [Combine](#combine-combine) reduce per source, then combine. Pick what the query needs. | | Compose `source()` | Combine `iq combine` | | --- | --- | --- | | Shape | a single `jq` filter | a positional spec per source, then one `--with` program | | Correlated reads
(B keyed by A's rows) | ✓ nest `source()` | ✗ specs are independent | | Quoting | sub-filter is a quoted string inside the filter | each spec is its own argument | | Memory | each reduced result held until combine, plus a correlated `source()` re-runs per row | each reduced result held until combined | Use `source()` when a read depends on another source's values, or to keep everything in one composable filter. Use `iq combine` for a straightforward join, union, or aggregate across a few sources. ### Compose `source()` `source("name"; "")` runs `` against source `name` (reduced, streamed and pushed down like any query) and yields its results as a stream. A one-argument `source("name")` yields the whole source. Both arguments of `source()` are strings, so they must be quoted. It also yields a stream. Collect it before indexing with `INDEX(source(…); .id)` or `[source(…)]`, not `source(…) | INDEX(.id)`. A filter that calls `source()` runs over a null input. Every read is an explicit `source()` call, and there is no implicit primary source. As a result, it needs no active source. Names resolve through the registry like `--src`, active-group namespacing included. ```sh { title='Join users and orders in a single filter (no active source needed)' } iq 'INDEX(source("users"; ".[]"); .id) as $u | source("orders"; ".[] | select(.total > 99)") | {name: $u[.userId].name, total}' ``` !!! tip "Correlated lookups re-run" A `source()` opened inside a stream runs its sub-filter once per element. The connection is reused, but the sub-filter re-executes. Hoist a constant lookup into a binding `INDEX(source("users"; ".[]"); .id) as $u | …` and index `$u` per element instead. !!! tip "Optimize memory usage" Binding `source()` to a jq variable materializes that call's whole result set in memory (`jq` indexing needs a concrete array), even though the read itself streams. Push the reduction into the sub-filter, `source("orders"; ".[] | select(.total > 99)")`, not `source("orders"; ".[]")` filtered outside, so only the rows you need are held. A bare `source("big")` over a large source buys no streaming benefit. When each side is large and independent, prefer `iq combine`. ### Combine `combine` `iq combine [=]... --with [flags]` Query several sources and combine their results with one `jq` program. Each positional is a source spec (`[=]`) reduced at the source (bounded reads, streaming scans, and predicate pushdown all still apply). The spec's results bind to a `jq` variable named after the source, with `/`, `.` and `-` becoming `_` (so `prod/books=.[]` binds `$prod_books`). `--with` is the final program, so it can join, union (`$a + $b`), aggregate, or fan across any number of sources. Each source reduces at the source, and a pushable filter pushes down. As a result, this never copies whole datasets to join them. A spec with no filter binds the whole keyspace, which is a whole-keyspace read. It needs `--unbounded`, exactly as the same expression does on a plain query. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--no-compile` | ✗ | disable server-side predicate pushdown. Run each spec's filter client-side | | | `--with ` | required | final jq over the bound source results (each spec's results bound to $name), run over a null input | ```sh { title='Join users with orders on a shared id, across two sources' } iq combine 'users=.[] | {id, name}' \ 'orders=.[] | select(.total > 99)' \ --with '($users | INDEX(.id)) as $u | $orders[] | . + {name: $u[.userId].name}' ``` !!! tip "Reduce, then combine" Each spec is evaluated independently. Its (already reduced) result is held in memory before `--with` runs. For this reason, keep a stage's output small with `select`/projection/aggregation. A spec that must materialize its whole source (`keys`, `.`, `map(...)`) still needs `--unbounded`, exactly like a single-source query. A `.[]`-rooted spec streams without it. ## Unbounded `--unbounded` Permit a filter that loads the whole dataset into memory. It also materializes a `.[]`-rooted filter instead of streaming it. As a result, a filter that collapses the keyspace into one value (`keys`, `.`, `map(...)`) runs only with it (see [Read strategies](how-it-works.md#read-strategies)). `iq combine` and `iq data` carry their own copies of the flag where they apply. ## Client-side `--no-compile` Disable pushdown. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--no-compile` | ✗ | disable predicate pushdown. Run the full `.[] | select()` filter client-side | # Write data docs/docs/write-data.md # Write data Data movement lives on the query command, like in `sq`. The `jq` filter is the transform. `--insert` names a destination. Piped stdin is an implicit source. There is no separate copy command. `iq exec` remains the untyped escape hatch for anything the typed path does not cover. ## Insert `--insert` With `--insert`/`--typed`, the filter transforms each item (its key is preserved). Iteration over the source is implicit, so you do not write `.[]`. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--insert ` | ✗ | write each item into this destination source (copy/restore/import) instead of rendering | | | `--typed` | ✗ | emit typed `{"key":…, "type":…, "value":…}` records, a re-importable dump (needed for Redis, a document store's plain output already restores) | ```sh { title='Source → Source, key/id preserving' } iq --src books --insert books2 ``` ```sh { title='Cross-driver (Redis → Mongo), object values only' } iq --src cache --insert docs ``` ```sh { title='Back up Redis losslessly (typed dump)' } iq --src cache --typed -o dump.jsonl ``` ```sh { title='Back up Mongo with plain output (self-describing)' } iq --src books --jsonl -o dump.jsonl ``` ```sh { title='Add a dump file, then restore it into a live source' } iq add file:///dump.jsonl -n snap iq --src snap --insert cache ``` ```sh { title='Move plan, no connection' } iq --src books --insert books2 --explain ``` !!! warning "Implicit iteration" This is the one place where `iq` reads a filter per item. Everywhere else (a plain query and the `=` specs `iq combine`, `iq diff` and `iq schema` take), the filter is rooted at the whole keyspace, and you write `.[]` yourself. The split is deliberate. A keyspace-rooted filter can aggregate across items (`[.[] | .total] | add`) and can name a single key (`.["orders:42"]`). A per-item filter can express neither. But when items arrive from piped stdin, the write path has no keyspace at all. Nothing can tell the two apart automatically (`.name` means the key named "name" keyspace-rooted and the field `name` per item). For this reason, they stay separate rather than guessing. Existing keys are overwritten (upsert) unless `--no-overwrite` makes the run insert-only. `--replace` empties the destination first (with confirmation, or `--force`). A document store (Mongo, CouchDB, Couchbase, Elasticsearch) stores each value exactly as given and so requires it be a JSON object. A bare scalar (for example a Redis string value) is rejected with a hint rather than silently wrapped as `{"value": …}`. As a result, a successful copy round-trips exactly. Shape it explicitly first, for example `iq 'if type == "object" then . else {value: .} end' --insert `. !!! note "Typed format" `--typed` serializes the records in the chosen format: `--jsonl` (default), `--json`, `--jsona`, or `--yaml`. All of these re-import through a `file://` source or a piped `--insert`. `iq` auto-detects them from content by their typed `{"key":…, "type":…, "value":…}` envelope. A huge first record can defeat the content sniff. For this reason, a `.yaml`/`.yml` name or an explicit `?format=` / `--from-format` remains available as an override (`jsonl`, `yaml`, `mongoexport`, `bson`, `rdb`, `dynamodb-json`, `cassandra-csv`, or `neo4j-json`). Aliases like `json` are accepted. The renderings that cannot carry a record back are rejected rather than written. `--raw` (a scalar cannot hold the envelope), `--format parquet` (columnar) and `--gron`/`--grona` (flattened assignments no source decodes) drop `--typed` to grep or export the value stream instead. ## Key mapping `--key*` On a plain copy, each item carries its source key, so no key flag is needed. Foreign JSON (piped stdin, a reshaped stream) has no natural key. `--key-field` takes it from an object field. `--key` computes it with a jq expression. `--key-prefix` prepends a namespace either way. `iq combine --insert` is the one write that always requires `--key` or `--key-field`. A combine's results come out of one program over a null input, so no value has a key to inherit. As a result, the run is refused up front rather than failing partway through a copy. For the same reason, `combine` has no `--typed`. A typed dump is a stream of `{key,type,value}` records and needs the same key. A write flag used without `--insert` is an error, never silently ignored. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--key ` | none | `jq` expression yielding each written item's key (--insert/--typed) | | | `--key-field ` | none | object field to take each written item's key from (--insert) | | | `--key-prefix ` | none | string prepended to every written key (--insert/--typed) | ```sh { title='Persist the joined rows into a third source' } iq combine 'users=.[] | {id, name}' 'orders=.[] | select(.total > 99)' \ --with '($users | INDEX(.id)) as $u | $orders[] | . + {name: $u[.userId].name}' \ --insert joined --key '.userId | tostring' --key-prefix 'j:' ``` ```sh { title='import foreign JSON from stdin, keyed by id' } cat foreign.json | iq --insert books --key-field id ``` ```sh { title='reshape + re-key while copying' } iq '{t: .title}' --src books --insert kv --key '.t' ``` ## Type mapping `--type` A plain copy carries each item's native type along (a Redis hash lands as a hash), so no type flag is needed. A filter that reshapes the value drops that type. The output is plain JSON, and `--type` names the native type that the destination stores it as. Left unset, a typed destination infers it from the shape (Redis writes a scalar as a `string` and an object or array as `json`). As a result, `--type` is only required when you want something else: - A `hash` built from an object - A `list` or `set` from an array - A `zset` from `[{member, score}]` pairs. A document store ignores it. Every value is a document there. With `--typed`, the same flag stamps the `type` field of each dump record instead. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--type ` | none | native type stamped on each written value, for example hash, list, json (--insert/--typed) | ```sh { title='Reshape Mongo documents into Redis hashes' } iq '{title, year: (.year | tostring)}' --src books --insert cache --type hash ``` ```sh { title='Project one field per item into a Redis list' } iq '[.tags[]]' --src books --insert cache --type list --key-prefix 'tags:' ``` ```sh { title='Same reshape, no --type: an object lands as RedisJSON' } iq '{title, year}' --src books --insert cache ``` ```sh { title='Dump reshaped items with an explicit type tag' } iq '{title}' --src books --typed --type json -o titles.jsonl ``` ## Replace `--replace` By default, a write upserts. Existing keys are overwritten, and everything else in the destination stays. `--replace` turns the copy into a restore. It empties the destination first (the same operation as `iq data clear`, a Redis `FLUSHDB`, a Mongo `deleteMany({})`) and then writes. As a result, the destination ends up holding exactly the source. Because it destroys data, it prompts (`clear before writing`). `--force` answers yes. `--dry-run` reports the copy without clearing anything. A destination that cannot be cleared (the read-only file dump) is refused up front. `--no-overwrite` is the opposite choice, insert-only, so the two are mutually exclusive. Neither applies to a `--typed` dump. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--replace` | ✗ | empty the destination before writing, with confirmation (--insert) | | | `--force` | ✗ | skip the confirmation prompt for `--replace` | | | `--no-overwrite` | ✗ | skip keys that already exist (--insert) | ```sh { title='Restore a dump so the destination matches it exactly' } iq --src snap --insert cache --replace ``` ```sh { title='Same, unattended (no prompt)' } iq --src snap --insert cache --replace --force ``` ```sh { title='Preview the restore: reports the copy, clears nothing' } iq --src snap --insert cache --replace --dry-run ``` ```sh { title='Fill gaps only, never touch an existing key' } iq --src books --insert books2 --no-overwrite ``` ## Lifecycle previews Every `iq data` subcommand (`delete`, `clear`, `drop`) shares two previews. `clear` and `drop` also prompt before destroying data. `--force` skips the prompt. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--explain` | ✗ | print the access plan and exit without connecting or changing anything | | | `--dry-run` | ✗ | connect and report the real effect without changing anything | ```sh { title='Plan only: which operation, and whether the driver supports it' } iq data drop cache --explain ``` ```sh { title='Connect and count what a clear would remove, then stop' } iq data clear shop.orders --dry-run ``` ```sh { title='Check which of the named keys exist before deleting' } iq data delete cache book:1 book:2 --dry-run ``` ## Delete `data delete` `iq data delete ` removes a named set of keys, keeping the container. It is the typed, capability-gated, explainable counterpart of a raw per-key delete (HBase's `exec delete` verb, a Redis `DEL`). Each key uses the same spelling as a *Get*: a bare string (`book:1`) or a JSON array for a composite key (`["shop",42]`). A key already absent is not an error (delete is idempotent). The report states it (`deleted N key(s), M already absent`). Unlike `data clear`/`data drop`, it does not prompt. The explicit key list you typed is the confirmation (use `--dry-run` to preview). A backend with no per-key identity (the read-only file dump) rejects it (like Redis rejects `data drop`). ```sh { title='"deleted 2 key(s), 0 already absent"' } iq data delete cache book:1 book:2 ``` ```sh { title='a composite-key row, by its JSON-array spelling' } iq data delete shop.orders '["eu",42]' ``` ## Clear `data clear` Empties a container but keeps it (for example MongoDB `deleteMany({})` or Redis `FLUSHDB`). Both this and `iq data drop` are distinct from `iq rm`, which only unregisters a saved source. These destroy stored data, and prompt for confirmation unless `--force` is used. ```sh { title='Empty a Redis source called "cache" (FLUSHDB)' } iq data clear cache ``` ```sh { title='Empty two MongoDB collections "shop.orders" and "shop.users"' } iq data clear shop.orders shop.users ``` ## Drop `data drop` Removes the container entirely and prompts for confirmation unless `--force` is used. Sources without a droppable container are rejected (for example Redis). ```sh { title='Reports drop as unsupported for Redis source called "cache"' } iq data drop cache --explain ``` ```sh { title='Drop one Mongo collection called "shop.orders"' } iq data drop shop.orders ``` ```sh { title='Drop containers "shop.orders" and "shop.users"' } iq data drop shop.orders shop.users ``` ## Dry run `--dry-run` Reports the effect without writing. It does everything except the final mutation. It connects, opens source and destination, and runs the real scan. It applies the real transform and probes capabilities. Then it suppresses the write. | short :material-flag-outline: | long :material-flag-outline: | default | description | | --- | --- | --- | --- | | | `--dry-run` | ✗ | report the effect of `--insert` without writing anything | ```sh { title='Dry run for a copy into "books2"' } iq --src books --insert books2 --dry-run ``` ```sh { title='Dry run for clearing source "books"' } iq data clear books --dry-run ``` # Output docs/docs/output.md # Output Results print as pretty JSON by default. A single format flag selects another rendering. The flags are mutually exclusive. They apply to the jq read path and to `iq combine`, but not to `exec` (which prints the backend's native reply). ## Format flags Shorthand flags are available for most formats. `-f`, `--format ` selects the same renderings by name (`json`, `jsonl`, `jsona`, `yaml`, `values` with the alias `raw`, `gron`, `grona`, `parquet`). It is mutually exclusive with them, so `-f json --jsonl` is rejected. `parquet` has no shorthand flag. It is a binary columnar format, selected by name only. ### JSON `--json` Pretty JSON, one value per result (default). Shorthand `-j`. ``` iq '.[]' -j ``` ``` { "_id": "1", "author": "Donovan and Kernighan", "price": 39, "tags": [ "go", "programming" ], "title": "The Go Programming Language", "year": 2015 } { "_id": "2", "author": "Martin Kleppmann", "price": 45, "tags": [ "data", "architecture" ], "title": "Designing Data-Intensive Applications", "year": 2017 } ``` ### JSON Lines `--jsonl` Compact JSON, one value per line. Shorthand `-J`. ``` iq '.[]' -J ``` ``` {"_id":"1","author":"Donovan and Kernighan","price":39,"tags":["go","programming"],"title":"The Go Programming Language","year":2015} {"_id":"2","author":"Martin Kleppmann","price":45,"tags":["data","architecture"],"title":"Designing Data-Intensive Applications","year":2017} ``` ### JSON Array `--jsona` Every result wrapped in one array `[ ... ]`. Shorthand `-A`. ``` iq '.[]' -A ``` ``` [ { "_id": "1", "author": "Donovan and Kernighan", "price": 39, "tags": [ "go", "programming" ], "title": "The Go Programming Language", "year": 2015 }, { "_id": "2", "author": "Martin Kleppmann", "price": 45, "tags": [ "data", "architecture" ], "title": "Designing Data-Intensive Applications", "year": 2017 } ] ``` !!! info "`iq --jsona` differs from `sq --jsona`" `iq --jsona` wraps the whole result stream in one array (like `jq -s`). It is the analogue of `sq --json`. `sq --jsona` instead emits one JSON array *per row* with the keys dropped. This is a columnar projection that a heterogeneous `jq` value stream has no exact analogue for. As a result, `iq` keeps the `sq` flag name but its own behaviour. ### Raw `--raw` Unquoted scalars, one per line. Objects and arrays fall back to compact JSON. Shorthand `-r`. ``` iq '.[].title' -r ``` ``` The Go Programming Language Designing Data-Intensive Applications ``` ### YAML `--yaml` YAML documents, separated by `---`. Shorthand `-y`. ``` iq '.[]' -y ``` ``` _id: "1" author: Donovan and Kernighan price: 39 tags: - go - programming title: The Go Programming Language year: 2015 --- _id: "2" author: Martin Kleppmann price: 45 tags: - data - architecture title: Designing Data-Intensive Applications year: 2017 ``` ### gron `--gron` Flattened `json.path = value;` assignment statements, one per line. Shorthand `-g`. It is greppable and reversible with `gron --ungron`[^1]. Each result is rooted at a repeated `json`. ``` iq '.[]' -g ``` ``` json = {}; json._id = "1"; json.author = "Donovan and Kernighan"; json.price = 39; json.tags = []; json.tags[0] = "go"; json.tags[1] = "programming"; json.title = "The Go Programming Language"; json.year = 2015; json = {}; json._id = "2"; json.author = "Martin Kleppmann"; json.price = 45; json.tags = []; json.tags[0] = "data"; json.tags[1] = "architecture"; json.title = "Designing Data-Intensive Applications"; json.year = 2017; ``` !!! note "Paths" `--gron` (and `--grona`) emit one `path = ;` statement per line, object keys sorted. A key that is an ASCII identifier(`^[A-Za-z_$][A-Za-z0-9_$]*$`) follows a bare dot (`json.name`). Any other key is bracketed and JSON-quoted (`json["odd key"]`). This is a deliberate ASCII subset of gron's rule, because over-quoting stays ungron-safe. `--gron` repeats the `json` root for every result, so ungron[^1] is last-write-wins across results. !!! warning "Typed dumps are not supported" `--typed` rejects `--gron` and `--grona`. A flattened assignment stream is a rendering to grep, not a dump. No source re-imports it. To gron the value stream, drop `--typed`. To dump instead, use `--jsonl` (default), `--json`, `--jsona`, or `--yaml`. ### gron Array `--grona` Like `--gron` but result N roots at `json[N]`, so the whole stream ungrons[^1] back to one JSON array (gron's `--stream` style). Shorthand `-G`. ``` iq '.[]' -G ``` ``` json = []; json[0] = {}; json[0]._id = "1"; json[0].author = "Donovan and Kernighan"; json[0].price = 39; json[0].tags = []; json[0].tags[0] = "go"; json[0].tags[1] = "programming"; json[0].title = "The Go Programming Language"; json[0].year = 2015; json[1] = {}; json[1]._id = "2"; json[1].author = "Martin Kleppmann"; json[1].price = 45; json[1].tags = []; json[1].tags[0] = "data"; json[1].tags[1] = "architecture"; json[1].title = "Designing Data-Intensive Applications"; json[1].year = 2017; ``` !!! note "Paths" `--grona` roots result **N** at `json[N]` under a leading `json = [];`, so ungron[^1] rebuilds the full array (an empty stream ungrons[^1] to `[]`, like `--jsona`). ### Parquet `--format parquet` Streams the result values to an [Apache Parquet](https://parquet.apache.org/) file ([Apache Arrow](https://arrow.apache.org/) columnar format). Parquet is the bridge to [pandas](https://pandas.pydata.org/), [Polars](https://pola.rs/), [DuckDB](https://duckdb.org/), and the wider data-science ecosystem. Because it is binary, `iq` refuses to write it to a terminal. Redirect it or use `-o out.parquet`. A pipe or file is required. ``` iq '.[]' --format parquet -o out.parquet ``` ``` iq '.[]' --format parquet > out.parquet ``` ```bash iq '.[]' --format parquet | python3 -c " import sys, pyarrow.parquet as pq, io table = pq.read_table(io.BytesIO(sys.stdin.buffer.read())) print(table) " ``` !!! note "Schema" The schema is inferred from the first 1000 result values (schema-inference sample). It is projected onto Arrow types: - `integer→int64` - `number→float64` - `boolean→bool` - `string→utf8` - A `date-time` string→`timestamp[ns, UTC]` - A `date` string→`date32` - `object→struct` - `array→list` - An id-keyed map→`map`. A column whose sampled shape is heterogeneous or null-only falls back to the `arrow.json` canonical extension (utf8 storage holding byte-lossless canonical JSON), marked in field metadata. The Arrow schema is embedded in the file (`ARROW:schema`), so exact types survive a read-back. A value that does not fit its inferred column type past the sample fails the export, naming the column. For fully heterogeneous data, switch to `--format jsonl` rather than coercing. !!! info "Presence caveat" Arrow's validity bitmaps cannot distinguish a *missing* field from a field present as `null`. Both collapse to a null in the column. `iq` preserves the distinction inferred from the sample in field metadata (`iq:presence` = `required` | `optional`), so it survives in the schema even though the values collapse. !!! warning "Typed dumps are not supported" `--typed` dumps cannot use `parquet` (they carry a `{key,type,value}` envelope). Run the query without `--typed` to export a columnar file. ## Compact `--compact` Collapses the pretty renderings to single-line: `--json` becomes one compact value per line (equivalent to `--jsonl`) and `--jsona` becomes a single-line `[ ... ]`. It is a no-op for `--jsonl`, `--raw`, `--yaml`, `--gron`, `--grona`, and `--format parquet`, which are already condensed or binary (`--gron` and `--grona` are inherently line-based). ## File `--output ` `--output` :material-earth:{ title="Global flag" } is global. Every command honours it (for example `iq inspect -o report.json`). See [Global flags](global-flags.md#global-flags). Writes results to `` instead of stdout, truncating an existing file, with the shorthand `-o`. It is orthogonal to the format flags. Color is off for a file unless you force it with `-C`. Progress and errors still go to stderr. ## Numbers `--format.decimal` `--format.decimal` :material-earth:{ title="Global flag" } is global. `--format.decimal ` chooses how a **non-integer decimal** from the backend is presented to the filter. | value | behavior | | --- | --- | | `auto` (default) | each backend keeps its faithful form. MongoDB `Decimal128` is an exact string. A Redis fractional number is a `float64` | | `number` | decimals become bare numbers (`float64`), convenient for arithmetic but lossy beyond `float64` | | `string` | decimals become their exact literal as a string, precision-safe. Use `tonumber` to compute | ```bash { title='"19.99" — exact, precision-safe' } ./iq '.book.price' --format.decimal=string ``` ```bash title="Compute on the exact decimal" ./iq '.book.price | tonumber * 1.2' ``` ```bash title="Bare numbers, ready for jq arithmetic" ./iq '.[].price' --format.decimal=number ``` !!! warning Because the `jq` filter runs client-side over the fetched value, this choice is made at normalization time. It changes what the filter computes on, not only how the result prints (unlike `sq`, where `jq` is not involved). !!! note Integers are always exact regardless of the mode. They arrive as an `int`, or as a big integer when they exceed 64 bits. As a result, `.count + 1` stays exact rather than rounding through `float64`. A backend can round before `iq` sees the value. For example, RedisJSON stores an integer larger than 64 bits as a double, so it arrives already in scientific notation. A big integer renders as a bare number in the JSON formats but as a quoted string under `--yaml` (a `yaml.v3` limitation). Exactness is kept in preference to YAML's numeric form. ## Color `--color` `--color` :material-earth:{ title="Global flag" }, shorthand `-C`, is global. Output is syntax-highlighted when `iq` writes to a terminal. It is left plain when it is piped or redirected, so captured output stays free of color codes. TTY detection is where capture safety comes from. ```bash title="Colored on a terminal, plain when piped" ./iq '.[]' ``` ```bash title="Never colored" ./iq '.[]' -M ``` ```bash title="Keep color through a pager" ./iq '.[]' -C | less -R ``` !!! note "Rendering" Every rendering syntax-highlights on a terminal (`--json`, `--jsonl`, `--jsona`, `--yaml`, `--raw`, `--gron`, `--grona`). Under `--raw`, strings and nulls still print bare and uncolored, keeping shell substitution exact. The human commands color their signal too: - `ping` shows `ok`/`error` in green/red. - `diff` shows additions green, removals red, and changes yellow. - `ls`/`inspect` highlight the active source and section headers. The raw reply bodies from `exec` and `inspect` are colored in their native form. For example, MongoDB gets JSON syntax highlighting, and Redis gets redis-cli-style value tokens. ## No color `--monochrome` `--monochrome` :material-earth:{ title="Global flag" }, shorthand `-M`, is global. Disable colored output. Color is on by default only when writing to a terminal. !!! tip "NO_COLOR" Colored output is also disabled if the `NO_COLOR` environment variable is set. `-C` forces colored output and overrides `NO_COLOR`. [^1]: `gron` flattens JSON into one `json.path = value;` assignment per line, so it can be grepped. `gron --ungron` reverses that and rebuilds the JSON from the assignments. This is what makes `--gron` output round-trippable. https://github.com/tomnomnom/gron#ungronning # Global flags docs/docs/global-flags.md # Global flags Flags every `iq` command accepts, a query, `iq inspect`, `iq diff`, `iq data drop` alike. Each one is documented on the page that owns its subject, and carries a :material-earth:{ title="Global flag" } there to mark that it works everywhere. | short :material-flag-outline: | long :material-flag-outline: | documented in | | --- | --- | --- | | `-s` | `--src` | [Sources](sources.md#showset-src) | | | `--config` | [Configuration](configuration.md#configuration-config) | | `-o` | `--output` | [Output](output.md#file-output-file) | | `-C` | `--color` | [Output](output.md#color-color) | | `-M` | `--monochrome` | [Output](output.md#no-color-monochrome) | | | `--format.decimal` | [Output](output.md#numbers-formatdecimal) | | | `--no-cache` | [Drivers](drivers.md#file-dumps) | | | `--no-cache-index` | [Drivers](drivers.md#file-dumps) | | | `--no-progress` | [Diagnostics & Logging](diagnostics-and-logging.md#flags) | | | `--log*` | [Diagnostics & Logging](diagnostics-and-logging.md#flags) | | | `--error*` | [Diagnostics & Logging](diagnostics-and-logging.md#flags) | | | `--debug.pprof` | [Diagnostics & Logging](diagnostics-and-logging.md#flags) | | `-v` | `--verbose` | [Diagnostics & Logging](diagnostics-and-logging.md#flags), plus [Query plan](query-plan.md#query-plan) for the formatted plan and [Sources](sources.md#list-ls) for the extra listing columns | The three below have no topic page of their own. ## Timeout `--timeout` Timeout for the whole operation, a query, a diff, a combine, or an `--insert` copy (default is `5s`). ## Help `--help` Display help on the command line for `iq`. ## Version `--version` Prints the bare version and exits, with no name or prefix. As a result, a script can use it directly (`v=$(iq --version)`). The version has one of these forms: - `v1.2.3` for a release build - The git-describe form (`v1.2.3-14-gabc1234`) for an untagged build - `dev+` only for a plain `go build` with no version metadata at all. It always writes to stdout, unaffected by `--output`. For the human form (version, commit, build date, Go version), run `iq version`. ## Not global The query command carries flags of its own, `--unbounded` and `--no-compile` on [Query data](query-data.md#query-data). `--explain` on [Query plan](query-plan.md#query-plan). `--from-format` on [Write data](write-data.md#write-data) and the output-format and write families on their own pages. `iq combine` and `iq data` carry their own copies of `--unbounded`/`--explain` where they apply. # Query plan docs/docs/query-plan.md # Query Plan `--explain` prints a formatted **query plan** and exits without connecting or executing. `-v`/`--verbose` prints the same plan to stderr, then runs, tracing each backend command. ## Examples ```bash title="Plan only, no connection" ./iq --src orders '.[] | select(.total > 99) | {id, total}' --explain ``` ```bash title="Annotated plan, a note per pipe stage" ./iq --src orders '.[] | select(.total > 99) | {id, total}' --explain -v ``` ```bash title="Redis SCAN + typed reads" ./iq --src cache '.[] | select(.active)' --explain ``` ```bash title="One pushed, one client-side" ./iq --src orders '.[] | select(.total > 99 and (.active | not))' --explain ``` ```bash { title='Plan + live "mongo> find(...)" trace' } ./iq --src orders '.[] | select(.total > 99)' -v ``` ```bash title="Trace on stderr, stdout stays pure data" ./iq '.[]' -v 2>/dev/null ``` ## Breakdown The plan shows four things, syntax-highlighted when the destination is a terminal: - **Filter** - The `jq` filter pretty-printed with real line breaks. Nested `source("name"; "")` sub-filters and every `iq combine` fragment are formatted too. - Under `-v`/`--verbose`, each top-level pipe stage also carries a short right-aligned note describing it (`— keep inputs where …`). The stage that reads from the store is marked with its route, colored by cost. - A green `bounded read` (keyed lookup) - A yellow `streaming scan` (batched over `.[]`) - A red `materialized scan` (an aggregate, a non-`.[]` root, or any scan under `--unbounded`) - **Access Plan** - The concrete backend calls each source will make, derived from the filter's route (bounded keys, streaming scan, or materialize) - **Pushdown** - The breakdown: one line per top-level `select(...)` conjunct. Each line says whether the backend evaluates it (`pushed`) or it re-runs client-side (`client-side`). It also says why a client-side conjunct did not push. - A conjunct is `client-side` when the compiler cannot express it as a provable superset (an inexact negation, a non-portable regex, an unsafe field name, or any other unpushable construct). It is also `client-side` when the backend's translator declines the compiled predicate (a range on Elasticsearch, say), or when the source does no server-side filtering at all (the read-only file driver). - This is the observable split of what the backend evaluated versus what ran client-side. It is absent under `--no-compile`. - **Compiled Filter** - As JSON, the merged fragment of the pushed conjuncts. !!! danger "Redaction" Credentials are never traced (for example Redis `AUTH` and the MongoDB auth handshake are redacted or skipped). !!! info "Explain is a dry run" `--explain` never opens a connection, so it works offline against any saved source. !!! note "Verbosity" Under `-v`, the live trace shows the actual commands (`redis> TYPE …`, `mongo> find …`) as they run. # Configuration docs/docs/configuration.md # Configuration `config` The config file lives at `/iq/iq.toml`. `IQ_CONFIG` points it elsewhere, and `--config` :material-earth:{ title="Global flag" } overrides both for one run (precedence: `--config` > `IQ_CONFIG` > default). It holds the saved sources and the stored option defaults below. `iq config location` prints the resolved path. iq writes the config file with mode `0600`, so only you can read it. A source saved with `--store inline` (the default) keeps its password in this file. If the file holds an inline password and other users can access it, for example after an edit or a copy, every command prints a warning to stderr with the fix (`chmod 600 `) and then runs as usual. Windows has no such mode, so there is no warning there. `iq config keyring migrate` moves inline passwords into the OS keyring. Inspect the config file and manage stored option defaults. Persist a flag's value once so you need not retype it. Set an option globally, or scope it to one source with `--src`. At query time, the precedence is **explicit flag > per-source option > base option > built-in default**. As a result, a saved default fills any flag you leave unset. An explicit flag on the command line always takes precedence. ```sh { title='Every query defaults to YAML output' } iq config set format yaml ``` ```sh { title='30s timeout only when querying "prod"' } iq config set --src prod timeout 30s ``` ```sh { title='Effective value for "prod" (source > base > default)' } iq config get --src prod format ``` ```sh { title='Every persistable option: value, default, and help' } iq config ls -v ``` ```sh { title='Renders YAML (the stored default)' } iq '.[]' ``` ```sh { title='Explicit flag overrides the stored default' } iq -f json '.[]' ``` !!! note "Persistable options" Persistable options are the flags whose default it is reasonable to persist: - `--format`, `--format.decimal`, `--compact` - `--timeout` - `--monochrome`, `--color`, `--no-progress` - `--no-cache`, `--no-cache-index` - `--verbose`, `--log*`, `--error*` Per-invocation flags are not storable: - `--src` - `--explain` - `--unbounded` - `--no-compile` - `--reveal`, `--expand` - `--debug.pprof` An `iq combine` query has no single source, so it uses the base options only, never a per-source override. !!! warning "Logging options" The `log*` options are the one place where a stored default and the environment overlap. A stored `log*` value fills an unset flag. But an `$IQ_LOG*` environment variable still wins over it (the `flag > env > default` chain for logging applies before a stored default is treated as "set"). An explicit `--log*` flag beats both. No other option reads the environment, so this interaction is unique to the logging family. ## Edit `edit` Open the config file in `$IQ_EDITOR` (then `$VISUAL`, `$EDITOR`, else `vi`). ```sh iq config edit ``` ## Get `get` Print an option's effective value at that scope. ```sh iq config get [--src ]
…` is the typed, capability-gated per-key delete that formalizes the raw `delete` verb above. The raw `put`/`delete` verbs remain the lower-level raw path (a single cell, a column), mirroring the other drivers' raw paths. `iq inspect` lists the source namespace's tables (`tables`). ## MongoDB [MongoDB](https://www.mongodb.com) is a document database that stores JSON-like documents in collections. After you register a `mongodb://` source (`mongodb+srv://` for SRV discovery), the same jq interface works against a collection. **The collection is the keyspace, a document's `_id` is the key and the document is the value**. The database comes from the URI path. The collection comes from the URI's `?collection=` (the driver's own connection option, overridable per run with a dotted `handle.collection`). ```sh { title='Register a MongoDB source' } iq add -n books 'mongodb://localhost:27017/iq?collection=books' ``` ```sh { title='Fetch the document whose _id is "2"' } iq --src books '.["2"]' ``` ```sh { title='Stream the collection, filtered' } iq --src books '.[] | select(.year > 2015) | .title' ``` ```sh { title='Materialize every _id' } iq --src books --unbounded 'keys' ``` Because Mongo values are natively typed, numeric comparisons like `.year > 2015` need no `tonumber`. This is unlike Redis, where everything is a string. Documents normalize to JSON with the same rules everywhere: - An `ObjectID` becomes its hex string. - A date becomes an RFC 3339 string. - Numbers stay numbers. - Nested documents and arrays are preserved. A missing `_id` reads as `null`. The `--unbounded` / streaming rules are identical to every backend (`.[]`-rooted filters stream a cursor in constant memory, `keys`/`.`/`map` materialize and require the flag). ### Pushdown By default, the driver translates the **equality**, **range**, **regex**, **existence**, **length** and **array** clauses of a `.[] | select(...)` filter into a native Mongo query. As a result, the server does the filtering (and can use an index) before the documents ever reach iq: ```sh { title='Pushed down: the server filters by author' } iq --src books '.[] | select(.author == "Robert C. Martin") | .title' ``` ```sh { title='Forced client-side with --no-compile' } iq --src books --no-compile '.[] | select(.author == "Robert C. Martin") | .title' ``` Pushdown never changes results, only speed. The full jq always re-runs client-side over whatever comes back, so a pushed filter is only ever a conservative pre-filter. Pass `--no-compile` to skip it and stream the whole collection, filtering entirely client-side. What it can push: | `select(...)` clause | Pushed | MongoDB translation | Notes | | --- | :---: | --- | --- | | `.a == x` | ✓ | `{a: x}` | number, string, bool, or null literal | | `.a == 1 or .a == 2` | ✓ | `{a: {$in: [1, 2]}}` | an `or` of equalities on one field | | `.a > n`, `>=`, `<`, `<=` | ✓ | native op + `$type` guards (an `$or`) | number/string literal. The translation reproduces jq's cross-type order, so the match is never a subset | | `.a \| test("re")` | ✓ | `{a: {$regex: "re", $options: "is"}}` | portable patterns only (below), and jq's `i` and `m` flags. jq's `m` (dot-matches-newline) maps to PCRE's `s` | | `has("a")`, `.a \| has("k")` | ✓ | `{a: {$exists: true}}` | exact, key presence, like jq's `has()` | | `.a \| length == n` | ✓ | `{$size: n}` + `$type` guards (an `$or`) | jq `length` is polymorphic (array/string/object/number), so guards keep it a superset | | `.a \| any(cond)` | ✓ | `{a: {$elemMatch: cond}}` (an `$or` with an object guard) | an array element satisfying a pushable element predicate. `cond` can combine the rows above | | `.a != x` | ✓ | `{$or: [{a: {$ne: x}}, {a: {$type: "array"}}]}` | exact negation of equality (the guard keeps arrays, which jq never equates to a scalar) | | `has("a") \| not` | ✓ | `{a: {$exists: false}}` | exact negation of existence | | `.a \| any(.f == v) \| not` | ✓ | `{a: {$not: {$elemMatch: …}}}` | no array element matches. The element condition must be exact equality | | `E1 and E2`, `E1 or E2` | ✓ | `$and` / `$or` of the above | an `and` can push only its pushable parts and drop the rest | | negated range/regex/`size` | — | — | their filters are supersets and a negated superset is a subset (unrecoverable) | | `.a > true`, `.a < null` | — | — | a range against bool/null has no clean superset | | non-portable regex | — | — | engine-specific construct (below) | | anything else | — | — | runs client-side, as under `--no-compile` | **Portable regex.** iq's jq is [gojq](https://github.com/itchyny/gojq), which compiles a `test()` pattern with Go's RE2. MongoDB uses PCRE. A pattern is pushed only when every construct it uses means the same, or a superset, in both. These constructs qualify: - Literals - Anchors (`^` `$`) - `.` - Quantifiers (`* + ? {n,m}`) - Alternation (`|`) - Groups - Character classes - The ASCII `\d` `\w` `\s` `\D` `\W` shorthands - Word boundaries (`\b`, `\B`). `\S` is the one shorthand held back. RE2's `\s` omits the vertical tab that PCRE's `\s` matches. As a result, RE2's `\S` matches a vertical tab that PCRE's does not. If pushed, it drops a document jq keeps (`\s` diverges the other way, a superset the client-side re-run corrects). Flags follow the same rule. gojq accepts only `i`, `m`, `g`, and iq pushes `i` (case-insensitive) and `m`. In jq, `m` means "`.` matches newline" (dotall), so it maps to PCRE's `s`, not PCRE's `m`. A pattern that uses one of these constructs is not portable and stays client-side: - Lookaround (`(?=…)`) - Backreferences (`\1`) - Unicode properties (`\p{…}`) - POSIX classes (`[[:…:]]`) - Possessive quantifiers. As a result, the pushed set always equals jq's. ### Raw commands `iq exec` runs a single JSON command document with `runCommand` and prints the reply as JSON. It is the raw path for server-side queries, aggregation and administration: ```sh { title='Run a native find command' } iq --src books exec '{"find":"books","filter":{"year":{"$gt":2015}}}' ``` ```sh { title='Run an aggregation pipeline' } iq --src books exec '{"aggregate":"books","pipeline":[{"$group":{"_id":null,"avg":{"$avg":"$price"}}}],"cursor":{}}' ``` `iq inspect` runs these diagnostic database commands: - `dbStats` - `serverStatus` - `listCollections` - `collStats` (needs a collection, address it as `source.collection` or set `?collection=` on the source URI) - `buildInfo` - `hostInfo`. `--only` narrows to those subcommands. ## Neo4j [Neo4j](https://neo4j.com) is a graph database of nodes and relationships, queried with Cypher. After you register a `neo4j://` source, the same jq interface works against a node label. **The node label is the keyspace, a node's key is the key and the node is the value**. Neo4j has no single keyspace, so a label is the addressable collection (like a Mongo collection or a Cassandra table). The host is the bolt server. The label rides in the URI's `?label=` (overridable per run with a dotted `handle.label`). Neo4j is multi-database, and the database is `?database=` (default `neo4j`). Use `neo4j+s://` (or `bolt://` for a single instance, `+s`/`+ssc` for TLS). **Credentials travel in the URI userinfo** (bolt basic auth), so `--store keyring` moves the password to the OS keyring exactly as for the other backends. **The key is the elementId by default, or a property you name with `?key=`.** `elementId(n)` is always present and unique, but opaque and not stable across database reloads. As a result, a `?key=` property (a stable, human-meaningful id) reads better. The value carries `_id` (the elementId) and `_labels` alongside the node's properties, so identity survives whichever key you choose. ```sh { title='Register a Neo4j source keyed by the "id" property' } iq add -n graph 'neo4j://neo4j:password@localhost:7687/?label=Person&key=id' ``` ```sh { title='Fetch the Person whose id is 1' } iq --src graph '.["1"]' ``` ```sh { title='Stream the label, filtered' } iq --src graph '.[] | select(.age > 40) | .name' ``` ```sh { title='Materialize every key in the label' } iq --src graph --unbounded 'keys' ``` ```sh { title='Query a different label on the same database' } iq --src graph.Book '.[]' ``` Neo4j values map to JSON directly. Integers keep exact precision. Bytes become base64. Temporal and spatial values become their canonical ISO strings and `{x,y,srid}` objects. A scan pages the label with keyset pagination ordered by `elementId(n)`. A `?key=` property is not guaranteed unique (unlike a primary key). For this reason, a scan falls back to a node's elementId whenever the key collides within a page, so no node is ever silently dropped. A bounded `.["v"]` lookup that matches more than one node is an error rather than an arbitrary pick. ### Relationship collections A **relationship type** is an addressable collection too, so you can query a graph's edges the same way. Name it with `?rel=KNOWS` on the source, or address one per run with the `:` marker (`handle.:KNOWS`). A leading colon can never be a valid label, so it unambiguously selects a relationship type. A source names either a label or a relationship type, not both. ```sh { title='Register a relationship-type source' } iq add -n edges 'neo4j://neo4j:password@localhost:7687/?rel=WROTE' ``` ```sh { title='Stream WROTE edges, filtered (pushdown on the edge)' } iq --src edges '.[] | select(.year > 2015)' ``` ```sh { title='Address the type per run from any source' } iq --src graph.:WROTE '.[]' ``` Each relationship's value is its properties plus a self-describing envelope: `_type` (the type), `_id` (its elementId) and `_start` / `_end` (the endpoint node elementIds). The scan, key, count, and `select(...)` pushdown rules are identical to nodes (the predicate is pushed onto the edge variable). **Relationship collections are read-only for now.** Creating an edge needs endpoint resolution (which nodes to connect and by which key), and that is a further follow-up. As a result, `iq` refuses a copy or `iq data` write into a relationship source with a clear message. Write nodes with `?label=`. ### Pushdown By default, the driver translates the **equality** and **existence** clauses of a `.[] | select(...)` filter into a Cypher `WHERE` clause. As a result, the server filters before nodes reach iq. The clause uses dynamic `n[$prop]` access, so the property name is a parameter, never string-built: | `select(...)` clause | Pushed | Cypher | Notes | | --- | :---: | --- | --- | | `.a == x` | ✓ | `n[$p] = $v` | equality. A `null` literal becomes `n[$p] IS NULL` (a missing property) | | `.a \| has` / `has("a")` | ✓ | `n[$p] IS NOT NULL` | key presence, exact | | `has("a") \| not` | ✓ | `n[$p] IS NULL` | key absence, exact | | `E1 and E2` | ✓ | `(… AND …)` | drops any conjunct it cannot push (widening) | | `E1 or E2` | ✓ | `(… OR …)` | pushed only when **every** branch is pushable | | `.a > x` / `.a <= x`, `!=`, `length`, regex, `any`, nested paths | — | — | run client-side. Cypher compares mismatched types as null rather than by jq's cross-type ordering. A nested path has no flat Neo4j property. As a result, pushing these can wrongly exclude a node that jq keeps | Pushdown never changes results, only speed. The full jq always re-runs client-side, so a pushed filter is a conservative pre-filter. `--explain` shows the `WHERE` clause. `--no-compile` streams the whole label and filters entirely client-side. ### Writing A copy into a Neo4j label upserts each node with `MERGE (n:Label {key}) SET n += props`, so a re-run converges. **Writing needs a `?key=` property** (a MERGE key must be stable and the elementId is server-assigned) **and a uniqueness constraint on it** (`CREATE CONSTRAINT ... REQUIRE n. IS UNIQUE`). Without the constraint, a MERGE can match and overwrite several nodes at once. For this reason, the write is refused up front rather than fanning out. `iq data clear` detach-deletes every node in the label (and the relationships they hold). A label is not a droppable container, so `iq data drop` is unsupported. Writes set node properties only. Relationships are a follow-up. ### Raw commands `iq exec` runs raw, parameterized [Cypher](https://neo4j.com/docs/cypher-manual/current/). The first argument is the statement. An optional second argument is a JSON object of parameters (passed as parameters, never string-built into the statement). It prints the result rows as JSON: ```sh { title='Run parameterized Cypher' } iq --src graph exec 'MATCH (n:Person) WHERE n.age > $min RETURN n.name, n.age' '{"min": 40}' ``` ```sh { title='Count the nodes' } iq --src graph exec 'MATCH (n) RETURN count(n) AS nodes' ``` `iq inspect` reads deployment and schema metadata through these subcommands: - `server` (components and version) - `databases` (the deployment's databases) - `labels` (the addressable node labels) - `reltypes` (relationship types) - `constraints` (which shows the uniqueness constraint a `?key=` write needs). `--only` narrows to those subcommands. ## Redis [Redis](https://redis.io) is an in-memory key-value store used as a cache, database and message broker. After you register a `redis://` source (`rediss://` for TLS), the same jq interface works against the Redis keyspace. **A key maps directly to a Redis key and the value is whatever that key holds**. The database index comes from the URI path (`/0`). Every value is a string, so numeric comparisons need `tonumber`. ```sh { title='Register a Redis source' } iq add -n cache redis://localhost:6379/0 ``` ```sh { title='Fetch the key "greeting"' } iq --src cache '.greeting' ``` ```sh { title='Stream the keyspace, filtered (string values need tonumber)' } iq --src cache '.[] | select((.year|tonumber) > 2015) | .title' ``` ```sh { title='Materialize every key' } iq --src cache --unbounded 'keys' ``` The `--unbounded` / streaming rules match every backend (`.[]`-rooted filters stream in constant memory, `keys`/`.`/`map` materialize and require the flag). ### Value encoding Each fetched Redis value is normalized to JSON by type: | Redis type | JSON shape | | --- | --- | | string | the string verbatim (numeric strings stay strings, use `tonumber`) | | hash | object `{field: value}` | | list | array, in list order | | set | array, sorted lexically (sets have no native order) | | sorted set | array of `{"member": ..., "score": ...}`, in ascending score order | | stream | array of `{"id": ..., "fields": {field: value}}`, in entry order | | RedisJSON | the stored document, parsed as JSON | | missing key | `null` | Other module types (time series, bloom, …) have no frozen encoding yet. `iq` refuses a named read of one with a clear message. ### Pushdown Redis has no server-side filtering. As a result, a compiled predicate drives a **client-side raw-byte prefilter** instead. On a streaming scan, the prefilter tests each RedisJSON value against the predicate on its raw JSON.GET bytes. When the value provably cannot match, the prefilter drops it before the (dominant) decode. Every other type is decoded and included unchanged. The full jq always re-runs client-side, so output is identical with or without the prefilter. The prefilter only skips decoding documents that the filter rejects. `--no-compile` turns it off. ### Raw commands `iq exec` forwards a command to the database verbatim and prints the reply in redis-cli style. It is the raw path for writes, administration and seeding that the jq read path does not cover: ```sh { title='Set a key, replies "OK"' } iq --src cache exec SET greeting hello ``` ```sh { title='Read it back, replies "hello"' } iq --src cache exec GET greeting ``` ```sh { title='Increment a counter, replies (integer) 1' } iq --src cache exec INCR counter ``` ```sh { title='Read a missing key, replies (nil)' } iq --src cache exec GET missing ``` Its output mirrors redis-cli's cooked style: - Bulk strings are quoted. - Integers appear as `(integer) N`. - A missing value appears as `(nil)`. - Arrays appear as a numbered, indented list. The client uses RESP2, so aggregate replies match redis-cli's classic flat output. Status replies such as `OK` and `PONG` appear quoted. This is a limitation of the underlying client, which does not distinguish them from bulk strings. `iq inspect` runs `INFO`. `--only` narrows it to sections (`server`, `clients`, `memory`, `persistence`, `stats`, `replication`, `cpu`, `keyspace`). With no section, it runs the full `INFO`. ## File dumps A `file://` source reads a database dump straight from disk. As a result, you can do these operations on a snapshot with the same jq interface, with **no running server**: - Query it - Inspect it for shape - Diff it - Restore it. A `file://` source is read-only. A `file://` endpoint is never a copy *destination*. `iq exec`/`iq inspect` (which need a live server) do not apply. ```sh { title='Register a dump like any source' } iq add file:///backups/prod.rdb -n snap ``` ```sh { title='Bounded read of one key' } iq --src snap '.["session:42"]' ``` ```sh { title='Streamed scan' } iq --src snap '.[] | select(.active)' ``` ```sh { title='Whole-dataset filters obey --unbounded' } iq --src snap 'keys' --unbounded ``` ```sh { title='Restore the dump into a live source' } iq --src snap --insert prod ``` ```sh { title='Diff a dump against a live source' } iq diff snap prod --data ``` The format is detected from the file's content (or forced with a `?format=` query, for example `file:///d.bin?format=bson`). A gzipped dump is unwrapped automatically. A gzipped dump must pass `?format=`, because its content is not sniffable through the compression. On Windows, a drive path takes the `file:///C:/path/to/dump.json` form (forward slashes, three slashes before the drive letter). | Format | Produced by | Notes | | --- | --- | --- | | Typed JSONL | `iq --src --typed -o ` | iq's own dump, lossless round-trip | | Typed YAML | `iq --src --typed -y -o ` | the same records as YAML documents, auto-detected by a `.yaml`/`.yml` name, else `?format=yaml` | | Redis RDB | `redis-cli --rdb`, `SAVE` | values match a live scan, RDB ≤ v12 (Redis ≤ 7.2) | | Mongo BSON | `mongodump` | single `.bson` file | | Mongo Extended JSON | `mongoexport` | one document per line, or a `--jsonArray` array | | DynamoDB JSON | S3 `export-table-to-point-in-time`, `aws dynamodb scan` | needs `?format=dynamodb-json` and a `?keys=pk[:S][,sk[:N]]` key schema (a dump carries items but not the table's key schema). Export files are gzipped NDJSON | | Cassandra CSV | `cqlsh COPY … TO 'f.csv'` | needs `?format=cassandra-csv`, `?keys=col1[,col2]` naming the primary-key columns, and `?types=col=cqltype,…` for the non-text columns (COPY writes every value as text). Column names come from a `WITH HEADER=TRUE` row, else `?columns=`. Scalar columns only | | Neo4j APOC JSON | `CALL apoc.export.json.all('g.json',{})` | needs `?format=neo4j-json` and a keyspace selector: either `?label=