diff --git a/modules/ROOT/nav.adoc b/modules/ROOT/nav.adoc index a413d36df3..882d606c7d 100644 --- a/modules/ROOT/nav.adoc +++ b/modules/ROOT/nav.adoc @@ -209,6 +209,7 @@ *** xref:manage:iceberg/query-iceberg-topics.adoc[Query Iceberg Topics] *** xref:manage:iceberg/migrate-iceberg-catalog.adoc[Migrate Iceberg Catalogs] *** xref:manage:iceberg/iceberg-performance-tuning.adoc[Tune Iceberg Performance] +*** xref:manage:iceberg/monitor-iceberg-health.adoc[] *** xref:manage:iceberg/iceberg-troubleshooting.adoc[Troubleshoot Iceberg Topics] ** xref:manage:schema-reg/index.adoc[Schema Registry] *** xref:manage:schema-reg/schema-reg-overview.adoc[Overview] diff --git a/modules/get-started/pages/release-notes/redpanda.adoc b/modules/get-started/pages/release-notes/redpanda.adoc index 468dccf544..1eac7e182b 100644 --- a/modules/get-started/pages/release-notes/redpanda.adoc +++ b/modules/get-started/pages/release-notes/redpanda.adoc @@ -7,240 +7,8 @@ This topic includes new content added in version {page-component-version}. For a * xref:cloud-data-platform:get-started:whats-new-cloud.adoc[] * xref:cloud-data-platform:get-started:cloud-overview.adoc#redpanda-cloud-vs-self-managed-feature-compatibility[Redpanda Cloud vs Self-Managed feature compatibility] +== Iceberg: Health monitoring -== Cloud Topics +Redpanda now provides an Admin API endpoint to monitor the health of Iceberg topics, including external catalog reachability and per-partition commit lag. Use it to detect an unreachable catalog and identify the partitions with the highest commit lag. -xref:develop:manage-topics/cloud-topics.adoc[Cloud Topics] are now available, making it possible to use durable cloud storage (S3, ADLS, GCS) as the primary backing store instead of local disk, eliminating over 90% of cross-AZ replication costs. This makes them ideal for latency-tolerant, high-throughput workloads such as observability streams, analytics pipelines, and AI/ML training data feeds, where cross-AZ networking charges are the dominant cost driver. - -You can use Cloud Topics exclusively in Redpanda Streaming clusters, or in combination with traditional Tiered Storage and local storage topics on a shared cluster supporting low latency workloads. - -Cloud Topics require Tiered Storage and an Enterprise license. For setup instructions and limitations, see xref:develop:manage-topics/cloud-topics.adoc[]. - -== Group-based access control (GBAC) - -Redpanda {page-component-version} introduces xref:manage:security/authorization/gbac.adoc[group-based access control (GBAC)], which extends OIDC authentication to support group-based permissions. In addition to assigning roles or ACLs to individual users, you can assign them to OIDC groups. Users inherit permissions from all groups reported by their identity provider (IdP) in the OIDC token claims. - -GBAC supports two authorization patterns: - -* Assign a group as a member of an RBAC role so that all users in the group inherit the role's ACLs. -* Create ACLs directly with a `Group:` principal. - -Group membership is managed entirely by your IdP. Redpanda reads group information from the OIDC token at authentication time and works across the Kafka API, Schema Registry, and HTTP Proxy. - -== FIPS 140-3 validation and FIPS Docker image - -Redpanda's cryptographic module has been upgraded from FIPS 140-2 to https://csrc.nist.gov/pubs/fips/140-3/final[FIPS 140-3^] validation. Additionally, Redpanda now provides a FIPS-specific Docker image (`docker.redpanda.com/redpandadata/redpanda:-fips`) for `amd64` and `arm64` architectures, with the required OpenSSL FIPS module pre-configured. - -NOTE: If you are upgrading with FIPS mode enabled, ensure all SASL/SCRAM user passwords are at least 14 characters before upgrading. FIPS 140-3 enforces stricter HMAC key size requirements. - -See xref:manage:security/fips-compliance.adoc[] for configuration details. - -== Iceberg: Expanded JSON Schema support - -Redpanda now supports additional JSON Schema patterns when translating to Iceberg tables: - -* `$ref` support: Internal references using `$ref` (for example, `"$ref": "#/definitions/myType"`) are resolved from schema resources declared in the same document. External references are not yet supported. -* Map type from `additionalProperties`: `additionalProperties` objects that contain subschemas now translate to Iceberg `map`. -* `oneOf` nullable pattern: The `oneOf` keyword is now supported for the standard nullable pattern if exactly one branch is `{"type":"null"}` and the other is a non-null schema. - -See xref:manage:iceberg/specify-iceberg-schema.adoc#how-iceberg-modes-translate-to-table-format[Specify Iceberg Schema] for JSON types mapping and updated requirements. - -== Ordered rack preference for Leader Pinning - -xref:develop:produce-data/leader-pinning.adoc[Leader Pinning] now supports the `ordered_racks` configuration value, which lets you specify preferred racks in priority order. Unlike `racks`, which distributes leaders uniformly across all listed racks, `ordered_racks` places leaders in the highest-priority available rack and fails over to subsequent racks only when higher-priority racks become unavailable. - -== User-based throughput quotas - -Redpanda now supports throughput quotas based on authenticated user principals. Unlike client-based quotas (which rely on self-declared `client-id` values), user-based quotas enforce limits using verified identities from SASL, mTLS, or OIDC authentication. - -You can set quotas for individual users, default users, or fine-grained user/client combinations. See xref:manage:cluster-maintenance/about-throughput-quotas.adoc[] for conceptual details, and xref:manage:cluster-maintenance/manage-throughput.adoc#set-user-based-quotas[Set user-based quotas] to get started. - -== Cross-region Remote Read Replicas - -Remote Read Replica topics on AWS can be deployed in a different region from the origin cluster's S3 bucket. This enables cross-region disaster recovery and data locality scenarios while maintaining the read-only replication model. - -To create cross-region Remote Read Replica topics, configure dynamic upstreams that point to the origin cluster's S3 bucket location. Redpanda manages the number of concurrent dynamic upstreams based on your `cloud_storage_url_style` setting (virtual_host or path style). - -See xref:manage:tiered-storage.adoc#remote-read-replicas[Remote Read Replicas] for setup instructions and configuration details. - -== Automatic broker decommissioning - -When continuous partition balancing is enabled, Redpanda can automatically decommission brokers that remain unavailable for a configured duration. The xref:reference:properties/cluster-properties.adoc#partition_autobalancing_node_autodecommission_timeout_sec[`partition_autobalancing_node_autodecommission_timeout_sec`] property triggers permanent broker removal, unlike xref:reference:properties/cluster-properties.adoc#partition_autobalancing_node_availability_timeout_sec[`partition_autobalancing_node_availability_timeout_sec`] which only moves partitions temporarily. - -Key characteristics: - -* Disabled by default -* Requires `partition_autobalancing_mode` set to `continuous` -* Permanently removes the node from the cluster (the node cannot rejoin automatically) -* Processes one decommission at a time to maintain cluster stability -* Manual intervention required if decommission stalls - -See xref:manage:cluster-maintenance/continuous-data-balancing.adoc[] for configuration details. - -== Schema Registry contexts - -xref:manage:schema-reg/schema-reg-contexts.adoc[Schema Registry contexts] are namespaces that isolate schemas, subjects, and configurations within a single Schema Registry instance. Each context maintains its own schema ID counter, mode settings, and compatibility settings, so teams can share a Schema Registry without risking naming collisions or configuration drift. - -To enable contexts, set xref:reference:properties/cluster-properties.adoc#schema_registry_enable_qualified_subjects[`schema_registry_enable_qualified_subjects`] to `true` and restart your brokers. Once enabled, you reference schemas using qualified subject syntax, for example `:.staging:my-topic-value`. See xref:manage:schema-reg/schema-reg-contexts.adoc[] for configuration examples and limitations. - -== Schema Registry metadata properties - -xref:manage:schema-reg/schema-reg-overview.adoc#metadata-properties[Schema Registry metadata properties] let you store and retrieve arbitrary key-value pairs alongside schemas. Properties such as `owner`, `team`, or `application.version` travel with the schema through its lifecycle, making it easier to track ownership and lineage without modifying the schema itself. - -You can set metadata when registering a schema using the `POST /subjects/\{subject\}/versions` endpoint or with the xref:reference:rpk/rpk-registry/rpk-registry-schema-create.adoc[`--metadata-properties`] flag in `rpk registry schema create`. Metadata is returned in `GET /subjects/\{subject\}/versions/\{version\}` and `GET /schemas/ids/\{id\}` responses, and viewable with `rpk registry schema get --print-metadata`. New schema versions automatically inherit metadata from the most recent version of the subject unless you register with an explicit empty `metadata` object. - -== OIDC authentication with rpk - -Starting in v26.1.7, `rpk` supports the `OAUTHBEARER` SASL mechanism, so you can authenticate `rpk` to the Kafka API with an OIDC access token issued by your identity provider (IdP) instead of a SASL/SCRAM username and password. Pass the token through xref:reference:rpk/rpk-x-options.adoc#sasl-mechanism[`-X sasl.mechanism=OAUTHBEARER`] and xref:reference:rpk/rpk-x-options.adoc#pass[`-X pass="token:"`]. - -For end-to-end steps and prerequisites, see xref:manage:security/authentication.adoc#oidc-rpk[Connect to Redpanda with OIDC using rpk]. - -== New configuration properties - -**Storage mode:** - -* xref:reference:properties/cluster-properties.adoc#default_redpanda_storage_mode[`default_redpanda_storage_mode`]: Set the default storage mode for new topics (`local`, `tiered`, `cloud`, or `unset`) -* xref:reference:properties/topic-properties.adoc#redpanda-storage-mode[`redpanda.storage.mode`]: Set the storage mode for an individual topic, superseding the legacy `redpanda.remote.read` and `redpanda.remote.write` properties - -**Cloud Topics:** - -NOTE: Cloud Topics requires an Enterprise license. For more information, contact https://redpanda.com/try-redpanda?section=enterprise[Redpanda sales^]. - -* xref:reference:properties/cluster-properties.adoc#cloud_topics_allow_materialization_failure[`cloud_topics_allow_materialization_failure`]: Enable recovery from missing L0 extent objects -* xref:reference:properties/cluster-properties.adoc#cloud_topics_compaction_interval_ms[`cloud_topics_compaction_interval_ms`]: Interval for background compaction -* xref:reference:properties/cluster-properties.adoc#cloud_topics_compaction_key_map_memory[`cloud_topics_compaction_key_map_memory`]: Maximum memory per shard for compaction key-offset maps -* xref:reference:properties/cluster-properties.adoc#cloud_topics_compaction_max_object_size[`cloud_topics_compaction_max_object_size`]: Maximum size for L1 objects produced by compaction -* xref:reference:properties/cluster-properties.adoc#cloud_topics_epoch_service_max_same_epoch_duration[`cloud_topics_epoch_service_max_same_epoch_duration`]: Maximum duration a node can use the same epoch -* xref:reference:properties/cluster-properties.adoc#cloud_topics_fetch_debounce_enabled[`cloud_topics_fetch_debounce_enabled`]: Enable fetch debouncing -* xref:reference:properties/cluster-properties.adoc#cloud_topics_gc_health_check_interval[`cloud_topics_gc_health_check_interval`]: L0 garbage collector health check interval -* xref:reference:properties/cluster-properties.adoc#cloud_topics_l1_indexing_interval[`cloud_topics_l1_indexing_interval`]: Byte interval for index entries in long-term storage objects -* xref:reference:properties/cluster-properties.adoc#cloud_topics_long_term_file_deletion_delay[`cloud_topics_long_term_file_deletion_delay`]: Delay before deleting stale long-term files -* xref:reference:properties/cluster-properties.adoc#cloud_topics_long_term_flush_interval[`cloud_topics_long_term_flush_interval`]: Interval for flushing long-term storage metadata to object storage -* xref:reference:properties/cluster-properties.adoc#cloud_topics_metastore_lsm_apply_timeout_ms[`cloud_topics_metastore_lsm_apply_timeout_ms`]: Timeout for applying replicated writes to LSM database -* xref:reference:properties/cluster-properties.adoc#cloud_topics_metastore_replication_timeout_ms[`cloud_topics_metastore_replication_timeout_ms`]: Timeout for L1 metastore Raft replication -* xref:reference:properties/cluster-properties.adoc#cloud_topics_num_metastore_partitions[`cloud_topics_num_metastore_partitions`]: Number of partitions for the metastore topic -* xref:reference:properties/cluster-properties.adoc#cloud_topics_parallel_fetch_enabled[`cloud_topics_parallel_fetch_enabled`]: Enable parallel fetching -* xref:reference:properties/cluster-properties.adoc#cloud_topics_preregistered_object_ttl[`cloud_topics_preregistered_object_ttl`]: Time-to-live for pre-registered L1 objects -* xref:reference:properties/cluster-properties.adoc#cloud_topics_produce_no_pid_concurrency[`cloud_topics_produce_no_pid_concurrency`]: Concurrent Raft requests for producers without a producer ID -* xref:reference:properties/cluster-properties.adoc#cloud_topics_produce_write_inflight_limit[`cloud_topics_produce_write_inflight_limit`]: Maximum in-flight write requests per shard -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_max_interval[`cloud_topics_reconciliation_max_interval`]: Maximum reconciliation interval for adaptive scheduling -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_max_object_size[`cloud_topics_reconciliation_max_object_size`]: Maximum size for L1 objects produced by the reconciler -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_min_interval[`cloud_topics_reconciliation_min_interval`]: Minimum reconciliation interval for adaptive scheduling -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_parallelism[`cloud_topics_reconciliation_parallelism`]: Maximum concurrent objects built by reconciliation per shard -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_slowdown_blend[`cloud_topics_reconciliation_slowdown_blend`]: Blend factor for slowing down reconciliation -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_speedup_blend[`cloud_topics_reconciliation_speedup_blend`]: Blend factor for speeding up reconciliation -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_target_fill_ratio[`cloud_topics_reconciliation_target_fill_ratio`]: Target fill ratio for L1 objects -* xref:reference:properties/cluster-properties.adoc#cloud_topics_upload_part_size[`cloud_topics_upload_part_size`]: Part size for multipart uploads -* xref:reference:properties/cluster-properties.adoc#cloud_topics_epoch_service_epoch_increment_interval[`cloud_topics_epoch_service_epoch_increment_interval`]: Interval for cluster epoch incrementation -* xref:reference:properties/cluster-properties.adoc#cloud_topics_epoch_service_local_epoch_cache_duration[`cloud_topics_epoch_service_local_epoch_cache_duration`]: Cache duration for local epoch data -* xref:reference:properties/cluster-properties.adoc#cloud_topics_long_term_garbage_collection_interval[`cloud_topics_long_term_garbage_collection_interval`]: Interval for long-term storage garbage collection -* xref:reference:properties/cluster-properties.adoc#cloud_topics_produce_batching_size_threshold[`cloud_topics_produce_batching_size_threshold`]: Object size threshold that triggers upload -* xref:reference:properties/cluster-properties.adoc#cloud_topics_produce_cardinality_threshold[`cloud_topics_produce_cardinality_threshold`]: Partition cardinality threshold that triggers upload -* xref:reference:properties/cluster-properties.adoc#cloud_topics_produce_upload_interval[`cloud_topics_produce_upload_interval`]: Time interval that triggers upload -* xref:reference:properties/cluster-properties.adoc#cloud_topics_reconciliation_interval[`cloud_topics_reconciliation_interval`]: Interval for moving data from short-term to long-term storage -* xref:reference:properties/cluster-properties.adoc#cloud_topics_short_term_gc_backoff_interval[`cloud_topics_short_term_gc_backoff_interval`]: Backoff interval for short-term storage garbage collection -* xref:reference:properties/cluster-properties.adoc#cloud_topics_short_term_gc_interval[`cloud_topics_short_term_gc_interval`]: Interval for short-term storage garbage collection -* xref:reference:properties/cluster-properties.adoc#cloud_topics_short_term_gc_minimum_object_age[`cloud_topics_short_term_gc_minimum_object_age`]: Minimum age for objects to be eligible for short-term garbage collection - -**Object storage:** - -* xref:reference:properties/object-storage-properties.adoc#cloud_storage_gc_max_segments_per_run[`cloud_storage_gc_max_segments_per_run`]: Maximum number of log segments to delete from object storage during each xref:manage:tiered-storage.adoc#object-storage-housekeeping[housekeeping run] -* xref:reference:properties/object-storage-properties.adoc#cloud_storage_prefetch_segments_max[`cloud_storage_prefetch_segments_max`]: Maximum number of small segments to prefetch during sequential reads - -**Authentication:** - -* xref:reference:properties/cluster-properties.adoc#nested_group_behavior[`nested_group_behavior`]: Control how Redpanda handles nested groups extracted from authentication tokens -* xref:reference:properties/cluster-properties.adoc#oidc_group_claim_path[`oidc_group_claim_path`]: JSON path to extract groups from the JWT payload -* xref:reference:properties/cluster-properties.adoc#schema_registry_enable_qualified_subjects[`schema_registry_enable_qualified_subjects`]: Enable parsing of qualified subject syntax in Schema Registry - -**Other:** - -* xref:reference:properties/cluster-properties.adoc#delete_topic_enable[`delete_topic_enable`]: Enable or disable topic deletion via the Kafka DeleteTopics API -* xref:reference:properties/cluster-properties.adoc#internal_rpc_request_timeout_ms[`internal_rpc_request_timeout_ms`]: Default timeout for internal RPC requests between nodes -* xref:reference:properties/cluster-properties.adoc#log_compaction_max_priority_wait_ms[`log_compaction_max_priority_wait_ms`]: Maximum time a priority partition (such as `__consumer_offsets`) waits before preempting regular compaction -* xref:reference:properties/cluster-properties.adoc#partition_autobalancing_node_autodecommission_timeout_sec[`partition_autobalancing_node_autodecommission_timeout_sec`]: Duration a node must be unavailable before Redpanda automatically decommissions it - -=== Changes to default values - -* xref:reference:properties/cluster-properties.adoc#log_compaction_tx_batch_removal_enabled[`log_compaction_tx_batch_removal_enabled`]: Changed from `false` to `true`. -* xref:reference:properties/cluster-properties.adoc#tls_v1_2_cipher_suites[`tls_v1_2_cipher_suites`]: Changed from OpenSSL cipher names to IANA cipher names. - -==== v26.1.4 - -* xref:reference:properties/cluster-properties.adoc#max_concurrent_producer_ids[`max_concurrent_producer_ids`]: Changed from unlimited to `100000` to prevent unbounded memory growth with heavy producer usage. - -* xref:reference:properties/cluster-properties.adoc#max_transactions_per_coordinator[`max_transactions_per_coordinator`]: Changed from unlimited to `10000` to prevent resource exhaustion from excessive transaction sessions. - -=== Removed properties - -The following deprecated configuration properties have been removed in v26.1.1. If you have any of these in your configuration files, update them according to the guidance below. - -**RPC timeout properties:** - -Replace with xref:reference:properties/cluster-properties.adoc#internal_rpc_request_timeout_ms[`internal_rpc_request_timeout_ms`]. - -* `alter_topic_cfg_timeout_ms` -* `create_topic_timeout_ms` -* `metadata_status_wait_timeout_ms` -* `node_management_operation_timeout_ms` -* `recovery_append_timeout_ms` -* `rm_sync_timeout_ms` -* `tm_sync_timeout_ms` -* `wait_for_leader_timeout_ms` - -**Client throughput quota properties:** - -Use xref:reference:rpk/rpk-cluster/rpk-cluster-quotas.adoc[`rpk cluster quotas`] to manage xref:manage:cluster-maintenance/manage-throughput.adoc#client-throughput-limits[client throughput limits]. - -* `kafka_admin_topic_api_rate` -* `kafka_client_group_byte_rate_quota` -* `kafka_client_group_fetch_byte_rate_quota` -* `target_fetch_quota_byte_rate` -* `target_quota_byte_rate` - -**Quota balancer properties:** - -Use xref:manage:cluster-maintenance/manage-throughput.adoc#broker-wide-throughput-limit-properties[broker-wide throughput limit properties]. - -* `kafka_quota_balancer_min_shard_throughput_bps` -* `kafka_quota_balancer_min_shard_throughput_ratio` -* `kafka_quota_balancer_node_period` -* `kafka_quota_balancer_window` -* `kafka_throughput_throttling_v2` - -**Timestamp alert properties:** - -* `log_message_timestamp_alert_after_ms`: Use xref:reference:properties/cluster-properties.adoc#log_message_timestamp_after_max_ms[`log_message_timestamp_after_max_ms`] -* `log_message_timestamp_alert_before_ms`: Use xref:reference:properties/cluster-properties.adoc#log_message_timestamp_before_max_ms[`log_message_timestamp_before_max_ms`] - -**Other removed properties:** - -No replacement needed. These properties were deprecated placeholders that have been silently ignored and will continue to be ignored even after removal. - -* `cloud_storage_disable_metadata_consistency_checks` -* `cloud_storage_reconciliation_ms` -* `coproc_max_batch_size` -* `coproc_max_inflight_bytes` -* `coproc_max_ingest_bytes` -* `coproc_offset_flush_interval_ms` -* `datalake_disk_space_monitor_interval` -* `enable_admin_api` -* `enable_coproc` -* `find_coordinator_timeout_ms` -* `full_raft_configuration_recovery_pattern` -* `id_allocator_replication` -* `kafka_memory_batch_size_estimate_for_fetch` -* `log_compaction_adjacent_merge_self_compaction_count` -* `max_version` -* `min_version` -* `raft_max_concurrent_append_requests_per_follower` -* `raft_recovery_default_read_size` -* `rm_violation_recovery_policy` -* `schema_registry_protobuf_renderer_v2` -* `seed_server_meta_topic_partitions` -* `seq_table_min_size` -* `tm_violation_recovery_policy` -* `transaction_coordinator_replication` -* `tx_registry_log_capacity` -* `tx_registry_sync_timeout_ms` -* `use_scheduling_groups` +See xref:manage:iceberg/monitor-iceberg-health.adoc[] for details. diff --git a/modules/manage/pages/iceberg/monitor-iceberg-health.adoc b/modules/manage/pages/iceberg/monitor-iceberg-health.adoc new file mode 100644 index 0000000000..7b7fae2b59 --- /dev/null +++ b/modules/manage/pages/iceberg/monitor-iceberg-health.adoc @@ -0,0 +1,128 @@ += Monitor Iceberg Health +:description: Monitor the health of Iceberg topics using the Admin API to check external catalog reachability and per-partition commit lag. +:page-categories: Iceberg, Monitoring + +// tag::single-source[] +:page-topic-type: how-to +:personas: ops_admin, streaming_developer +:learning-objective-1: Check whether the external Iceberg catalog is reachable +:learning-objective-2: Identify Iceberg topics with the highest commit lag + +ifndef::env-cloud[] +[NOTE] +==== +include::shared:partial$enterprise-license.adoc[] +==== +endif::[] + +{description} + +After reading this page, you will be able to: + +* [ ] {learning-objective-1} +* [ ] {learning-objective-2} + +Use the `GetIcebergStatus` Admin API endpoint to check whether your external Iceberg catalog is reachable and to see how far behind catalog commit is for each Iceberg topic and partition. This helps you detect a misconfigured or unreachable catalog and identify the partitions with the highest commit lag on a busy cluster. + +NOTE: This endpoint is available in Redpanda version 26.2 and later. + +[IMPORTANT] +==== +This endpoint requires superuser privileges, even though it is read-only. See xref:manage:use-admin-api.adoc#authentication[Admin API authentication]. +==== + +== Check Iceberg health + +Send a POST request to the link:/api/doc/admin/v2/operation/operation-redpanda-core-admin-v2-icebergservice-geticebergstatus[`redpanda.core.admin.v2.IcebergService/GetIcebergStatus`] endpoint. With an empty request body, the response covers all Iceberg topics that Redpanda is currently tracking for translation: + +[,bash] +---- +curl \ + -u : \ + --request POST 'http://:/redpanda.core.admin.v2.IcebergService/GetIcebergStatus' \ + --header "Content-Type: application/json" \ + --data '{}' +---- + +To report on specific topics only, set `topicsFilter` to a list of Iceberg topic names: + +[,bash] +---- +curl \ + -u : \ + --request POST 'http://:/redpanda.core.admin.v2.IcebergService/GetIcebergStatus' \ + --header "Content-Type: application/json" \ + --data '{ "topicsFilter": ["orders", "clicks"] }' +---- + +* Request headers `Connect-Protocol-Version` and `Connect-Timeout-Ms` are optional. +* v2 endpoints also accept binary-encoded Protobuf request bodies. Use the `Content-Type: application/proto` header. + +A response looks like the following: + +[,json] +---- +{ + "catalog": { + "reachable": true, + "errorCode": "", + "errorMessage": "" + }, + "topics": [ + { + "topic": "orders", + "lifecycleState": "ICEBERG_LIFECYCLE_STATE_LIVE", + "partitions": [ + { + "partition": 0, + "lastTranslatedOffset": "10432", + "lastCommittedOffset": "10400", + "commitLag": "32" + } + ] + } + ] +} +---- + +== Interpret catalog reachability + +The `catalog` object reports whether Redpanda can reach the external Iceberg catalog: + +* *`reachable`*: `true` if the catalog responded, `false` otherwise. +* *`errorCode`*: When the catalog is unreachable, a short error identifier from the catalog error set, such as `io_error` or `timedout`. Empty when `reachable` is `true`. +* *`errorMessage`*: A human-readable diagnostic for the failure. Empty when `reachable` is `true`. + +If `reachable` is `false`, no commits can succeed until the catalog is available again. Check the catalog endpoint configuration, network connectivity, and credentials. + +== Interpret commit lag + +For each Iceberg topic, `topics[]` reports a `lifecycleState` and a list of `partitions`. The `lifecycleState` is one of `ICEBERG_LIFECYCLE_STATE_LIVE`, `ICEBERG_LIFECYCLE_STATE_CLOSED`, `ICEBERG_LIFECYCLE_STATE_PURGED`, or `ICEBERG_LIFECYCLE_STATE_UNSPECIFIED`. + +Each entry in `partitions[]` reports the following offsets: + +* *`partition`*: The partition ID. +* *`lastTranslatedOffset`*: The highest Kafka offset written to Parquet. Unset if nothing has been translated yet. +* *`lastCommittedOffset`*: The last Kafka offset committed to the Iceberg catalog. Unset if nothing has been committed yet. +* *`commitLag`*: The number of records written to Parquet but not yet committed to the catalog. + +A steadily rising `commitLag` across many partitions can indicate that the cluster is underpowered for the current write rate, or that catalog commits are slow or failing. To triage, sort partitions by `commitLag` and investigate the topics with the highest lag first. + +== Commit lag compared to translation lag + +[IMPORTANT] +==== +This endpoint returns *commit lag* only: the number of records written to Parquet but not yet committed to the catalog. It does not return *translation lag* (the gap between the topic's high watermark and the last translated offset), because translation lag is a client-side calculation against the Kafka API. +==== + +To monitor both dimensions over time, use the corresponding metrics: + +* xref:reference:public-metrics-reference.adoc#redpanda_iceberg_pending_commit_lag[`redpanda_iceberg_pending_commit_lag`]: The metric equivalent of `commitLag` from this endpoint. +* xref:reference:public-metrics-reference.adoc#redpanda_iceberg_pending_translation_lag[`redpanda_iceberg_pending_translation_lag`]: Records that are pending translation to Iceberg. This is the translation lag that the endpoint does not return. + +== Next steps + +* xref:manage:iceberg/iceberg-performance-tuning.adoc[Tune Iceberg performance] if commit lag is consistently high. +* xref:manage:iceberg/iceberg-troubleshooting.adoc[Troubleshoot Iceberg topics] to diagnose translation errors and inspect the dead-letter queue. + +// end::single-source[] diff --git a/modules/manage/pages/use-admin-api.adoc b/modules/manage/pages/use-admin-api.adoc index a3b6717faf..ce844732ef 100644 --- a/modules/manage/pages/use-admin-api.adoc +++ b/modules/manage/pages/use-admin-api.adoc @@ -81,11 +81,11 @@ The new endpoints differ from the legacy endpoints in the following ways: * URL paths use the fully-qualified names of the ConnectRPC services. * ConnectRPC endpoints accept only POST requests. -Use ConnectRPC endpoints with features introduced in v25.3 such as: +Use ConnectRPC endpoints with features such as: -// TODO: Add links to docs when they are merged -* Shadowing -* Connected client monitoring +* xref:manage:disaster-recovery/shadowing/index.adoc[Shadowing] +* xref:manage:cluster-maintenance/manage-throughput.adoc#view-connected-client-details[Connected client monitoring] +* xref:manage:iceberg/monitor-iceberg-health.adoc[Iceberg health monitoring] For a full list of available endpoints, see the link:/api/doc/admin/v2/[Admin API Reference]. Select "v2" in the version selector to view the ConnectRPC endpoints.