Skip to main content
Kinetica can query tables managed in an external data lake, reading them through the catalog that owns them rather than by pointing at the underlying data files directly. This allows Kinetica to honor each table’s current state, schema, and partition layout as the data lake evolves, without the table having to be redefined in the database. Data lake access is read-only. A data lake table is queried through a logical external table; it cannot be written to, created, or dropped from Kinetica. Querying a data lake table takes three objects:
  1. A credential - holds the authentication details for the catalog &, where applicable, the object store
  2. A catalog - holds the location of, and connection information for, the data lake catalog
  3. A logical external table - maps one data lake table into Kinetica for querying
This page covers the behavior common to every supported table format. See the format-specific pages for catalog configuration, catalog path rules, and supported features.

Supported Formats

Apache Hudi is not supported. For Apache Iceberg, Hive Metastore, JDBC, and filesystem/Hadoop catalogs are not supported.

Catalogs

A catalog holds the location of, and connection information for, a data lake catalog that is external to the database. Each catalog is created with a TABLE FORMAT identifying the table format it serves and a TYPE identifying the catalog implementation; only the combinations listed in Supported Formats are accepted. Catalogs are created via the /create/catalog native API call. The credential shape and options each catalog type requires differ considerably by format: The skip_validation option, which bypasses validation of the connection to the remote source, is common to all catalog types and defaults to false.

External Tables

A data lake table is mapped into Kinetica as a logical external table. The CATALOG PATH clause identifies the table within the catalog, and the datalake_catalog option names the catalog to resolve it against.
LOGICAL must be given explicitly. The default external table type is materialized, which is not available for data lake tables, so omitting LOGICAL fails with an error rather than defaulting to a logical table.
The accepted form of the catalog path differs by table format; see Iceberg catalog paths & Delta Lake catalog paths.
In the native API, the catalog path is given as the datalake_path option of /create/table/external.
The following external table features are not available for data lake tables, in either format:
  • SUBSCRIBE
  • transformations
  • text search columns
Column projection is applied when the underlying data files are read, rather than during scan planning.

Required Privileges

A catalog is resolved on every query against a data lake external table, not only when the table is created, so the following are required on an ongoing basis: Creating or dropping a catalog requires catalog administration permission.
A missing privilege is not reported as a permission error. Without read permission on the catalog, the catalog is reported as not found. Without connect permission on the data source, the catalog is built with no object store connection details at all, which surfaces later as an obscure object store failure rather than as a permission message. Grant these permissions explicitly rather than diagnosing that symptom as a storage misconfiguration.

Snapshot & Version Selection

Queries always read the table’s current state—the latest Iceberg snapshot, or the latest Delta Lake version.
The datalake_snapshot option of /create/table/external is not implemented for either table format and has no effect. There is no way to pin an external table to a specific snapshot or version.

Data Type Mapping

Data lake column types are mapped to Kinetica types as follows. The mapping itself is shared across table formats; what differs is which source type names each format can produce, noted per row. A column whose type is not listed above will cause creation of the external table to fail, rather than being silently dropped. Three mappings are lossy and worth noting:
  • timestamptz & timestamp_ntz both map to the same Kinetica timestamp as a plain timestamp; the distinction is not preserved.
  • A decimal whose precision exceeds that of the largest Kinetica decimal type falls back to double. This preserves magnitude rather than truncating high-order digits, but loses precision, and is not reported as a warning.
  • struct & map columns—and list columns whose elements are not simple scalars—are readable as json, not as structured columns.

Caching & Tuning

Kinetica caches data lake table metadata, to avoid re-reading it from the catalog on every query. The cache is governed by the following configuration parameters: Because table metadata is cached for the duration of datalake.table_metadata_cache_ttl, a query can read a state that is up to that stale. Where eventual consistency is not acceptable, enable datalake.table_metadata_cache_snapshot_check, which validates the cached entry against the catalog on every use. Note that iceberg.manifest_cache_enabled applies per cached table and is therefore subordinate to datalake.table_metadata_cache_enabled; disabling the latter makes the former moot. All of these parameters can be changed at runtime via /alter/system/properties and are reported by SHOW SYSTEM PROPERTIES.
The property names differ by surface. The configuration file and SHOW SYSTEM PROPERTIES use a dotted form (iceberg.manifest_cache_enabled), while the /alter/system/properties property map uses an underscored form (iceberg_manifest_cache_enabled).