# Cloud - Full Markdown Export > This file contains all Cloud documentation pages in markdown format for AI agent consumption. > Generated from 908 pages on 2026-08-28T16:28:36.112Z > Component: cloud-data-platform | Version: > Site: https://docs.redpanda.com ## About This Export This export includes the **latest version** () of the Cloud documentation. ### AI-Friendly Documentation Formats We provide multiple formats optimized for AI consumption: - **https://docs.redpanda.com/llms.txt**: Curated overview of all Redpanda documentation - **https://docs.redpanda.com/llms-full.txt**: Complete documentation export with all components - **https://docs.redpanda.com/cloud-data-platform-full.txt**: This file - Cloud documentation only - **Individual markdown pages**: Each HTML page has a corresponding .md file --- # Page 1: Manage Billing **URL**: https://docs.redpanda.com/cloud-data-platform/billing.md --- # Manage Billing > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Billing latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/index.adoc description: Learn about the metrics Redpanda uses to measure consumption and about subscriptions with committed use. page-git-created-date: "2024-06-06" page-git-modified-date: "2024-08-01" --- - [Billing and Support](billing/) Learn about the metrics Redpanda uses to measure consumption in Redpanda Cloud. - [Manage Payment Methods](manage-payment-methods/) Add a credit card, set a default payment method, and update billing contact information in Redpanda Cloud. - [View Billing Activity](view-billing-activity/) View charges, filter the resources breakdown, and export billing activity to CSV in Redpanda Cloud. - [Manage Billing Notifications](billing-notifications/) Manage billing notifications in Redpanda Cloud: what alerts you receive, who receives them, and how to configure your notification preferences. - AWS - [Use AWS Commitments](aws-commit/) Subscribe to Redpanda in AWS Marketplace with committed use. - [Use AWS Pay As You Go](aws-pay-as-you-go/) Subscribe to Redpanda in AWS Marketplace with pay-as-you-go billing, and cancel anytime. - Azure - [Use Azure Commitments](azure-commit/) Subscribe to Redpanda in Azure Marketplace with committed use. - GCP - [Use GCP Commitments](gcp-commit/) Subscribe to Redpanda in Google Cloud Marketplace with committed use. - [Use GCP Pay As You Go](gcp-pay-as-you-go/) Subscribe to Redpanda in Google Cloud Marketplace with pay-as-you-go billing, and cancel anytime. --- # Page 2: Use AWS Commitments **URL**: https://docs.redpanda.com/cloud-data-platform/billing/aws-commit.md --- # Use AWS Commitments > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use AWS Commitments latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: aws-commit page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: aws-commit.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/aws-commit.adoc description: Subscribe to Redpanda in AWS Marketplace with committed use. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-10-17" --- You can subscribe to Redpanda Cloud through AWS Marketplace and use your existing marketplace billing and credits to quickly provision clusters. View your bills and manage your subscription directly in the marketplace. With a usage-based billing commitment, you sign up for a minimum spend amount. Commitments are minimums: - If you use less than your committed amount, you still pay the minimum. Any unused amount on a monthly commitment rolls over to the next month until the end of your term. - If you use more than your committed amount, you can continue using Redpanda Cloud without interruption. You’re charged for any additional usage until the end of your term. > ❗ **IMPORTANT** > > When you subscribe to Redpanda Cloud through AWS Marketplace, you can only create clusters on AWS. ## [](#sign-up-in-aws-marketplace)Sign up in AWS Marketplace 1. Contact [Redpanda Sales](https://redpanda.com/contact) to request a private offer with possible discounts. 2. You will receive a private offer on AWS Marketplace. Review the policy and required terms, and click **Accept**. > 📝 **NOTE** > > If you don’t have a billing account associated with your project, you’re prompted to enable billing to link the subscription with a billing account. You are taken to the Redpanda sign-up page. 3. On the Redpanda sign-up page: - For **Email**, enter your email address to register with Redpanda. - For **Organization name**, enter a name for your new organization connected through AWS Marketplace. Redpanda organizations contain all resources, including clusters and networks. - Click **Sign up and create organization**. You will receive an email sent to the address you entered. 4. In the email, click **Verify email address**. This completes the registration and associates the email with a Redpanda account. 5. On the **Accept your invitation to sign up** page, click **Sign up** or **Log in**. You can now create resource groups, clusters, and networks in your organization. ## [](#next-steps)Next steps - [Create a Serverless cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#create-a-serverless-cluster) - [Create a BYOC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/) - [Create a Dedicated cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster/#create-a-dedicated-cluster) --- # Page 3: Use AWS Pay As You Go **URL**: https://docs.redpanda.com/cloud-data-platform/billing/aws-pay-as-you-go.md --- # Use AWS Pay As You Go > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use AWS Pay As You Go latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: aws-pay-as-you-go page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: aws-pay-as-you-go.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/aws-pay-as-you-go.adoc description: Subscribe to Redpanda in AWS Marketplace with pay-as-you-go billing, and cancel anytime. page-git-created-date: "2024-09-19" page-git-modified-date: "2026-05-05" --- Subscribe to Redpanda Cloud through AWS Marketplace to quickly provision Serverless and Dedicated clusters. With a usage-based pay-as-you-go subscription, you only pay for what you use and can cancel anytime. > ❗ **IMPORTANT** > > When you sign up for Redpanda Cloud through AWS Marketplace, you can only create clusters on AWS. ## [](#sign-up-in-aws-marketplace)Sign up in AWS Marketplace 1. In the AWS Marketplace, select [**Redpanda Cloud - The proven Apache Kafka alternative (Pay as You Go)**](https://aws.amazon.com/marketplace/pp/prodview-ecbu7wwsfh644?applicationId=AWSMPContessa&ref_=beagle&sr=0-3). 2. On the **Redpanda Cloud - Pay as You Go** overview page, click **View purchase options**, then click **Subscribe**. > 📝 **NOTE** > > If you don’t have a billing account associated with your project, you’re prompted to link the subscription with a billing account. 3. On the **Subscribe to Redpanda Cloud** page, click **Set up your account**. You’re taken to the Redpanda sign-up page. 4. On the Redpanda sign-up page: - For **Email**, enter your email address to register with Redpanda. - For **Organization name**, enter a name for your new organization connected through AWS Marketplace. > 💡 **TIP** > > This process creates a new organization, even for existing Redpanda customers. Organizations contain all resources, including clusters and networks. - Click **Sign up and create organization**. You will receive an email sent to the address you entered. 5. In the email, click **Verify email address**. This associates the email with a Redpanda account. 6. On the **Accept your invitation to sign up** page, enter the credentials you want to use for Redpanda Cloud. You can now create resource groups, networks, and clusters in your organization. ## [](#next-steps)Next steps - [Create a Serverless cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#create-a-serverless-cluster) - [Create a Dedicated cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster/#create-a-dedicated-cluster) --- # Page 4: Use Azure Commitments **URL**: https://docs.redpanda.com/cloud-data-platform/billing/azure-commit.md --- # Use Azure Commitments > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use Azure Commitments latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: azure-commit page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: azure-commit.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/azure-commit.adoc description: Subscribe to Redpanda in Azure Marketplace with committed use. page-git-created-date: "2024-10-30" page-git-modified-date: "2025-10-17" --- You can subscribe to Redpanda Cloud through Azure Marketplace and use your existing marketplace billing and credits to quickly provision clusters. View your bills and manage your subscription directly in the marketplace. With a usage-based billing commitment, you sign up for a monthly or an annual minimum spend amount. Commitments are minimums: - If you use less than your committed amount, you still pay the minimum. Any unused amount on a monthly commitment rolls over to the next month until the end of your term. - If you use more than your committed amount, you can continue using Redpanda Cloud without interruption. You’re charged for any additional usage until the end of your term. > ❗ **IMPORTANT** > > When you subscribe to Redpanda Cloud through Azure Marketplace, you can only create clusters on Azure. ## [](#sign-up-in-azure-marketplace)Sign up in Azure Marketplace 1. Contact [Redpanda sales](https://redpanda.com/contact) to request a private offer with possible discounts. You will receive a private offer on Azure Marketplace. This offer is associated with an Azure user account that has access to the Azure subscription used for billing. 2. In Azure Marketplace, review the policy and required terms, and click **Accept**. You are taken to the Redpanda sign-up page. 3. On the Redpanda sign-up page: - For **Email**, enter your email address to register with Redpanda. - For **Organization name**, enter a name for your new organization connected through Azure Marketplace. Redpanda organizations contain all resources, including clusters and networks. - Click **Sign up and create organization**. You will receive an email sent to the address you entered. 4. In the email, click **Verify email address**. This completes the registration and associates the email with a Redpanda account. 5. On the **Accept your invitation to sign up** page, click **Sign up** or **Log in**. You can now create resource groups, clusters, and networks in your organization. ## [](#next-steps)Next steps - [Create a BYOC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/) - [Create a Dedicated cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster/) --- # Page 5: Manage Billing Notifications **URL**: https://docs.redpanda.com/cloud-data-platform/billing/billing-notifications.md --- # Manage Billing Notifications > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Billing Notifications latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: billing-notifications page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: billing-notifications.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/billing-notifications.adoc description: "Manage billing notifications in Redpanda Cloud: what alerts you receive, who receives them, and how to configure your notification preferences." page-topic-type: how-to personas: platform_admin, evaluator learning-objective-1: Identify the billing notifications Redpanda Cloud sends and their thresholds learning-objective-2: Configure which users in your organization receive billing notifications learning-objective-3: Opt out of billing notification emails for yourself or your organization page-git-created-date: "2026-03-27" page-git-modified-date: "2026-04-07" --- Redpanda Cloud sends email notifications to help you monitor your billing balance. Organization admins receive alerts when credit or commit balances reach spending thresholds. In this guide, you will: - Identify the billing notifications Redpanda Cloud sends and their thresholds - Configure which users in your organization receive billing notifications - Opt out of billing notification emails for yourself or your organization ## [](#what-notifications-you-receive)What notifications you receive Redpanda Cloud monitors your balance and sends a notification when it crosses each threshold. Each threshold triggers one notification. If your balance crosses the same threshold again after adding credits, you may receive another notification at that level. | Notification | Description | Thresholds | | --- | --- | --- | | Low credit balance | Sent when your pre-paid credit balance is running low. Credits are drawn down by usage, similar to a prepaid account. | 50%, 30%, 10%, 0% remaining | | Low commit balance | Sent when your contractual commit balance is running low. Commits represent a minimum spend over a contract period. | 50%, 30%, 10%, 0% remaining | Notifications are sent to email only. The subject line follows this format: `Action Required: Your Redpanda Cloud is % remaining` ## [](#who-receives-notifications)Who receives notifications All users with the **Admin** role in your organization receive billing notifications by default. To change who receives notifications, update role assignments on the **Organization IAM** page. See [Role-Based Access Control](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/) or [Group-Based Access Control](https://docs.redpanda.com/cloud-data-platform/security/authorization/gbac/gbac/). ## [](#opt-out-of-notifications)Opt out of notifications ### [](#individual-opt-out)Individual opt-out To stop receiving billing notification emails: - Open any billing notification email. - Click the **Unsubscribe** or **Manage notification preferences** link at the bottom of the email. No support ticket is needed. The change takes effect within 24-48 hours. ### [](#organization-wide-opt-out)Organization-wide opt-out To disable billing notifications for all admins in your organization, contact [Redpanda support](https://support.redpanda.com/hc/en-us/requests/new). > 📝 **NOTE** > > If billing notifications are enabled for the organization, individual admins who have not unsubscribed will continue to receive notifications. ## [](#common-questions)Common questions - I didn’t sign up for these emails. Why am I receiving them? Billing notifications are sent automatically to all organization admins. If you don’t want to receive them, click the **Unsubscribe** link at the bottom of the email. - I got an alert but I already added credits. Why? Notifications are triggered when your balance crosses a threshold. If you added credits after the threshold was crossed, the notification was already queued. If your balance later crosses the same threshold again (for example, after adding credits and then using them), you may receive another notification. - Who else in my organization is getting these? All users with the Admin role receive billing notifications. To see who has the Admin role, check the **Organization IAM** > **Users** page in Redpanda Cloud. - I unsubscribed but still received a notification. What happened? Unsubscribe requests take 24-48 hours to process. If you receive a notification during that window, it was sent before your request was fully applied. - What should I do when I get an alert? Review your current balance on the **Billing** page. You can add credits or contact your Redpanda account team to discuss your usage and plan options. - Do trial accounts get notifications? Only if the trial has promotional credits. Standard trial accounts without a credit balance do not receive billing notifications. --- # Page 6: Billing and Support **URL**: https://docs.redpanda.com/cloud-data-platform/billing/billing.md --- # Billing and Support > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Billing and Support latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: billing page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: billing.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/billing.adoc description: Learn about the metrics Redpanda uses to measure consumption in Redpanda Cloud. page-topic-type: reference personas: platform_admin, evaluator page-git-created-date: "2024-06-06" page-git-modified-date: "2026-05-22" --- Redpanda Cloud uses various [metrics](#redpanda-streaming-billing-metrics) to measure the consumption of resources. - All pricing is set in US dollars (USD). - All usage-based billing computations are conducted in Coordinated Universal Time (UTC). Billing accrues at hourly intervals. Any usage that is less than an hour is billed for the full hour. - The **Billing** page shows detailed billing activity for your organization and lets you [manage payment methods](https://docs.redpanda.com/cloud-data-platform/billing/manage-payment-methods/). Redpanda charges the credit card marked as the default. > 📝 **NOTE** > > - Redpanda Cloud can notify you when your credit or commit balance is running low. See [Manage Billing Notifications](https://docs.redpanda.com/cloud-data-platform/billing/billing-notifications/). > > - Pricing information is available on [redpanda.com](https://www.redpanda.com/price-estimator). For questions about billing, contact [billing@redpanda.com](mailto:billing@redpanda.com). ## [](#redpanda-streaming-billing-metrics)Redpanda Streaming billing metrics ### Serverless Pricing for Serverless clusters depends on the data in, data out, data stored, partitions (virtual streams), and the time the instance is up. The cost for each Serverless metric varies based on the region you select for your cluster. | Metric | Description | | --- | --- | | Uptime | Tracks the number of hours the instance is running.NOTE: Uptime is not charged if partitions = 0 and storage = 0. This condition is met when all topics are deleted. | | Ingress | Tracks the data written into Redpanda (in GB).All Kafka protocol requests (except message headers) are counted as ingress as soon as they are read by Redpanda’s proxy process. | | Egress | Tracks the data read out of Redpanda (in GB).All Kafka protocol responses generated by the cluster (except message headers) are counted as egress as soon as the cluster processes the request, even if the client drops the connection before they are delivered. | | Partitions | Tracks the number of partitions used per hour. | | Storage | Tracks the data in object storage per hour (in GB). | See also: [Serverless limits](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#serverless-usage-limits) ### Dedicated Pricing for Dedicated clusters depends on the time the instance is up, the data in, data out, and data stored. | Metric | Description | | --- | --- | | Uptime | Tracks the number of hours the instance is running.The cost varies based on the region and tier you select for your cluster. | | Ingress | Tracks the data written into Redpanda (in GB).All Kafka protocol requests (including message headers) are counted as ingress as soon as they are read by Redpanda’s proxy process.The cost varies based on the region you select for your cluster. | | Egress | Tracks the data read out of Redpanda (in GB).All Kafka protocol responses generated by the cluster (including message headers) are counted as egress as soon as the cluster processes the request, even if the client drops the connection before they are delivered.The cost varies based on the number of availability zones (AZ) you select for your cluster. | | Storage | Tracks the usage of object storage on an hourly basis during the billing period (in GB-hours).Replication to object storage is implemented with Tiered Storage. All topics have a fixed replication factor of 3, but Redpanda counts each byte only once. | ### BYOC Pricing for BYOC clusters depends on compute, data in, data out, and data stored. The rate decreases as usage increases. | Metric | Description | | --- | --- | | Compute | Tracks the server resources (vCPU and memory) a cluster uses on an hourly basis in Redpanda units (RPUs). Where:1 RPU = 2 vCPU + 8 GB memory | | Ingress | Tracks the data written into Redpanda (in GB).All Kafka protocol requests (including message headers) are counted as ingress as soon as they are read by Redpanda’s proxy process. | | Ingress to Iceberg topics | Tracks the data written to Iceberg tables per hour (in GB).NOTE: This metric applies only if you write to Iceberg topics. This charge is in addition to the standard ingress charge. | | Egress | Tracks the data read out of Redpanda (in GB).All Kafka protocol responses generated by the cluster (including message headers) are counted as egress as soon as the cluster processes the request, even if the client drops the connection before they are delivered.The cost varies based on the number of availability zones (AZ) you select for your cluster. | | Storage | Tracks the usage of object storage on an hourly basis during the billing period (in GB-hours).Replication to object storage is implemented with Tiered Storage. All topics have a fixed replication factor of 3, but Redpanda counts each byte only once. | ## [](#redpanda-sql-billing-metrics)Redpanda SQL billing metrics Pricing for Redpanda SQL depends on the provisioned compute resources allocated to the SQL workload, measured in RPUs. The cost of an RPU can vary based on the cloud provider and region you select. | Metric | Description | | --- | --- | | Compute | Tracks the server resources (vCPU and memory) Redpanda SQL uses on an hourly basis in Redpanda units (RPUs). Where:1 RPU = 2 vCPU + 8 GB memory | > 📝 **NOTE** > > Redpanda SQL uses the same RPU definition as Redpanda Streaming BYOC clusters, but the price per RPU differs. Contact your Redpanda account team for current pricing. ## [](#redpanda-connect-billing-metrics)Redpanda Connect billing metrics Pricing per pipeline depends on the compute units you allocate. The cost of a compute unit can vary based on the cloud provider and region you select for your cluster. | Metric | Description | | --- | --- | | Compute | Tracks the server resources (vCPU and memory) a pipeline uses in compute units per hour. Where:1 compute unit = 0.1 CPU + 400 MB memory | ## [](#support-plans)Support plans All organizations in Redpanda require one of the following support plans: | Support plan | Features | | --- | --- | | Basic | Designed for non-production environmentsProvides minimal support: priority 3 tickets within 8 business hours response time and priority 4 tickets with no target response timeSupport availability is 8:00 AM to 5:00 PM Pacific Time, Monday through Friday, excluding federal US holidays | | Enterprise | Designed for production environments needing continuous availabilityP1/P2 tickets may be submittedSupport availability is 24/7, including holidays | | Premium | Designed for mission-critical workloads30-minute response times for production outagesIncludes a named Customer Success Manager to support planning and coordination, and 10 hours per month of consulting from a Solutions ArchitectRequired for deployments with BYOVPC/BYOVnet clusters | ## [](#next-steps)Next steps - [Use AWS Commitments](https://docs.redpanda.com/cloud-data-platform/billing/aws-commit/) - [Use Azure Commitments](https://docs.redpanda.com/cloud-data-platform/billing/azure-commit/) - [Use GCP Commitments](https://docs.redpanda.com/cloud-data-platform/billing/gcp-commit/) - [Create a Serverless cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#create-a-serverless-cluster) - [Create a Dedicated cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster/#create-a-dedicated-cluster) - [Create a BYOC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/) --- # Page 7: Use GCP Commitments **URL**: https://docs.redpanda.com/cloud-data-platform/billing/gcp-commit.md --- # Use GCP Commitments > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use GCP Commitments latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: gcp-commit page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: gcp-commit.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/gcp-commit.adoc description: Subscribe to Redpanda in Google Cloud Marketplace with committed use. page-git-created-date: "2024-06-06" page-git-modified-date: "2026-05-05" --- You can subscribe to Redpanda Cloud through Google Cloud Marketplace and use your existing marketplace billing and credits to quickly provision clusters. View your bills and manage your subscription directly in the marketplace. With a usage-based billing commitment, you sign up for a monthly or an annual minimum spend amount. Commitments are minimums: - If you use less than your committed amount, you still pay the minimum. Any unused amount on a monthly commitment rolls over to the next month until the end of your term. - If you use more than your committed amount, you can continue using Redpanda Cloud without interruption. You’re charged for any additional usage until the end of your term. > ❗ **IMPORTANT** > > When you subscribe to Redpanda Cloud through Google Cloud Marketplace, you can only create clusters on GCP. ## [](#sign-up-in-google-cloud-marketplace)Sign up in Google Cloud Marketplace 1. Contact [Redpanda sales](https://redpanda.com/contact) to request a private offer with possible discounts. 2. You will receive a private offer on Google Cloud Marketplace. Review the policy and required terms, and click **Accept**. > 📝 **NOTE** > > If you don’t have a billing account associated with your project, you’re prompted to enable billing to link the subscription with a billing account. You are taken to the Redpanda sign-up page. 3. On the Redpanda sign-up page: - For **Email**, enter your email address to register with Redpanda. - For **Organization name**, enter a name for your new organization connected through Google Cloud Marketplace. Redpanda organizations contain all resources, including clusters and networks. - Click **Sign up and create organization**. You will receive an email sent to the address you entered. 4. In the email, click **Verify email address**. This completes the registration and associates the email with a Redpanda account. 5. On the **Accept your invitation to sign up** page, click **Sign up** or **Log in**. You can now create resource groups, clusters, and networks in your organization. ## [](#next-steps)Next steps - [Create a Serverless cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#create-a-serverless-cluster) - [Create a BYOC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/) - [Create a Dedicated cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster/) --- # Page 8: Use GCP Pay As You Go **URL**: https://docs.redpanda.com/cloud-data-platform/billing/gcp-pay-as-you-go.md --- # Use GCP Pay As You Go > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use GCP Pay As You Go latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: gcp-pay-as-you-go page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: gcp-pay-as-you-go.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/gcp-pay-as-you-go.adoc description: Subscribe to Redpanda in Google Cloud Marketplace with pay-as-you-go billing, and cancel anytime. page-git-created-date: "2026-05-05" page-git-modified-date: "2026-05-05" --- Subscribe to Redpanda Cloud through Google Cloud Marketplace to provision Serverless and Dedicated clusters. With a usage-based pay-as-you-go subscription, you only pay for what you use and can cancel anytime. > ❗ **IMPORTANT** > > When you sign up for Redpanda Cloud through Google Cloud Marketplace, you can only create clusters on GCP. > 📝 **NOTE** > > Serverless on GCP is currently in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release. ## [](#sign-up-in-google-cloud-marketplace)Sign up in Google Cloud Marketplace 1. In the Google Cloud Marketplace, select [**Redpanda Cloud - The proven Apache Kafka alternative (Pay as You Go)**](https://console.cloud.google.com/marketplace/product/redpanda-public/redpanda-cloud-platform?project=redpanda-public). 2. On the **Redpanda Cloud - Pay as You Go** overview page, click **Subscribe**. > 📝 **NOTE** > > If you don’t have a billing account associated with your project, you’re prompted to link the subscription with a billing account. 3. On the **Subscribe to Redpanda Cloud** page, click **Set up your account**. You’re taken to the Redpanda sign-up page. 4. On the Redpanda sign-up page: - For **Email**, enter your email address to register with Redpanda. - For **Organization name**, enter a name for your new organization connected through Google Cloud Marketplace. > 💡 **TIP** > > This process creates a new organization, even for existing Redpanda customers. Organizations contain all resources, including clusters and networks. - Click **Sign up and create organization**. Redpanda sends a verification email to the address you entered. 5. In the email, click **Verify email address**. This associates the email with a Redpanda account. 6. On the **Accept your invitation to sign up** page, enter the credentials you want to use for Redpanda Cloud. You can now create resource groups, networks, and clusters in your organization. ## [](#next-steps)Next steps - [Create a Serverless cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#create-a-serverless-cluster) - [Create a Dedicated cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster/) --- # Page 9: Manage Payment Methods **URL**: https://docs.redpanda.com/cloud-data-platform/billing/manage-payment-methods.md --- # Manage Payment Methods > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Payment Methods latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: manage-payment-methods page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: manage-payment-methods.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/manage-payment-methods.adoc description: Add a credit card, set a default payment method, and update billing contact information in Redpanda Cloud. page-topic-type: how-to personas: platform_admin, evaluator learning-objective-1: Add a credit card as a payment method in Redpanda Cloud learning-objective-2: Set a default payment method that Redpanda charges automatically learning-objective-3: Update the billing contact information for your organization page-git-created-date: "2026-05-11" page-git-modified-date: "2026-05-21" --- To pay for usage in Redpanda Cloud, you must add a credit card on the **Billing** page. The card you add is the payment method for all billable resources in your organization, including Serverless, Dedicated, and BYOC clusters, Redpanda Connect pipelines, and your support plan. The most recently added card becomes the default payment method, but you can change the default at any time. After reading this page, you will be able to: - Add a credit card as a payment method in Redpanda Cloud - Set a default payment method that Redpanda charges automatically - Update the billing contact information for your organization ## [](#prerequisites)Prerequisites - You have the **Admin** role in your Redpanda Cloud organization. See [Role-Based Access Control](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/). - You have a valid credit card. ## [](#add-a-payment-method)Add a payment method 1. Sign in to [Redpanda Cloud](https://cloud.redpanda.com). 2. From the navigation menu, click **Billing**. 3. On the **Billing** page, select the **Payment methods** tab. 4. Click **Add payment method**. 5. Enter your card details and billing address, then click **Save**. The new card appears on the **Payment methods** tab. The most recently added card becomes the default payment method unless you select a different one. > 📝 **NOTE** > > - After you add a credit card, any remaining credit balance is applied first. After that, Redpanda charges the default card on the first of each month. > > - Serverless free trials do not require a credit card to start. After your trial ends, you have a 7-day grace period to add a payment method before your clusters are suspended. See [Serverless Clusters](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/). ## [](#set-the-default-payment-method)Set the default payment method If you have more than one card on file, you can choose which one Redpanda charges: 1. Sign in to [Redpanda Cloud](https://cloud.redpanda.com). 2. From the navigation menu, click **Billing**. 3. On the **Billing** page, select the **Payment methods** tab. 4. Find the card you want to use, and mark it as the default. The card you selected is now labeled **Default payment method** and is the one Redpanda charges for usage. ## [](#remove-a-payment-method)Remove a payment method 1. Sign in to [Redpanda Cloud](https://cloud.redpanda.com). 2. From the navigation menu, click **Billing**. 3. On the **Billing** page, select the **Payment methods** tab. 4. Find the card you want to remove, and delete it. The card is removed from the **Payment methods** tab. If the card you want to remove is the default and the only card on file, add another card and set it as the default first. For help, contact [billing@redpanda.com](mailto:billing@redpanda.com). ## [](#update-billing-contact-information)Update billing contact information The billing contact is the person and address Redpanda uses for invoices and billing-related communication. It is separate from the recipients of low-balance email alerts, which are all users with the **Admin** role in your organization. See [Manage Billing Notifications](https://docs.redpanda.com/cloud-data-platform/billing/billing-notifications/). To update the billing contact: 1. Sign in to [Redpanda Cloud](https://cloud.redpanda.com). 2. From the navigation menu, click **Billing**. 3. On the **Billing** page, select the **Settings** tab. 4. Next to **Billing contact information**, click **Edit**. 5. Update all required fields, then click **Save**. The updated billing contact appears on the **Settings** tab. ## [](#next-steps)Next steps - [Billing and Support](https://docs.redpanda.com/cloud-data-platform/billing/billing/) - [View Billing Activity](https://docs.redpanda.com/cloud-data-platform/billing/view-billing-activity/) - [Manage Billing Notifications](https://docs.redpanda.com/cloud-data-platform/billing/billing-notifications/) --- # Page 10: View Billing Activity **URL**: https://docs.redpanda.com/cloud-data-platform/billing/view-billing-activity.md --- # View Billing Activity > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: View Billing Activity latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: view-billing-activity page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: view-billing-activity.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/billing/pages/view-billing-activity.adoc description: View charges, filter the resources breakdown, and export billing activity to CSV in Redpanda Cloud. page-topic-type: how-to personas: platform_admin, evaluator learning-objective-1: View a summary of charges for your organization learning-objective-2: Filter the per-resource breakdown of billing activity learning-objective-3: Export billing activity to a CSV file page-git-created-date: "2026-05-11" page-git-modified-date: "2026-05-21" --- The **Billing activity** tab on the **Billing** page shows a summary of charges and a per-resource breakdown for your organization. Use it to review usage for the current or a previous month, drill into per-resource costs, and export charges for record keeping or external billing systems. After reading this page, you will be able to: - View a summary of charges for your organization - Filter the per-resource breakdown of billing activity - Export billing activity to a CSV file ## [](#prerequisites)Prerequisites - You have the **Admin** role in your Redpanda Cloud organization. See [Role-Based Access Control](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/). ## [](#view-a-summary-of-charges)View a summary of charges 1. Sign in to [Redpanda Cloud](https://cloud.redpanda.com). 2. From the navigation menu, click **Billing**. 3. On the **Billing** page, select the **Billing activity** tab. 4. From the **Time range** menu, choose the period to display, for example, the current month or a previous month. The tab has three sections: - **Usage totals**: A high-level summary that subtotals usage by resource type (Dedicated, BYOC, Serverless) and shows the total amount owed for the selected time range. Expand any row to see the metrics that contribute to that subtotal. Amounts are shown before any discounts are applied. - **Resources breakdown**: A per-resource list showing the cost of each cluster, Redpanda Connect pipeline, and your support plan. Filter the list by: - Resource type: All, Dedicated, Serverless, Redpanda Connect pipeline, or Support. - Resource group: Limit results to a specific resource group. - Resource name: Search by name. - Show deleted resources: Toggle to include resources that have been deleted in the selected time range. Expand any resource to see a table of activity (for example, Compute, Ingress, Egress, Storage, Uptime, Partitions) with quantity, unit price, and amount. - **Plan details**: A side panel showing your pricing plan (for example, Pay-as-you-go) and the payment method on file. ## [](#download-charges-as-csv)Download charges as CSV To download a CSV summary of your monthly charges per resource, click the download icon next to **Resources breakdown**. You can use the exported file for record keeping or to import into your billing system. ## [](#next-steps)Next steps - [Billing and Support](https://docs.redpanda.com/cloud-data-platform/billing/billing/) - [Manage Payment Methods](https://docs.redpanda.com/cloud-data-platform/billing/manage-payment-methods/) - [Manage Billing Notifications](https://docs.redpanda.com/cloud-data-platform/billing/billing-notifications/) --- # Page 11: Develop **URL**: https://docs.redpanda.com/cloud-data-platform/develop.md --- # Develop > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Develop latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/index.adoc description: Develop doc topics. page-git-created-date: "2024-06-06" page-git-modified-date: "2024-06-07" --- - [Kafka Compatibility](kafka-clients/) Kafka clients, version 0.11 or later, are compatible with Redpanda. Validations and exceptions are listed. - [Topics](topics/) Overview of standard topics in Redpanda Cloud. - [Produce Data](produce-data/) Learn how to configure producers and idempotent producers. - [Consume Data](consume-data/) Learn about consumer offsets and follower fetching. - [Use Redpanda with the HTTP Proxy API](http-proxy/) HTTP Proxy exposes a REST API to list topics, produce events, and subscribe to events from topics using consumer groups. - [Redpanda Cloud Management MCP Server](cloud-mcp/) Manage your Redpanda Cloud clusters, topics, and users through AI agents using natural language commands. - [Data Transforms](data-transforms/) Learn about WebAssembly data transforms within Redpanda Cloud. - [Transactions](transactions/) Learn how to use transactions; for example, you can fetch messages starting from the last consumed offset and transactionally process them one by one, updating the last consumed offset and producing events at the same time. - [Kafka Connect](managed-connectors/) Use Kafka Connect to stream data into and out of Redpanda. --- # Page 12: Redpanda Cloud Management MCP Server **URL**: https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp.md --- # Redpanda Cloud Management MCP Server > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Cloud Management MCP Server page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cloud-mcp/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cloud-mcp/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/cloud-mcp/index.adoc # Beta release status page-beta: "true" description: Manage your Redpanda Cloud clusters, topics, and users through AI agents using natural language commands. page-git-created-date: "2026-06-15" page-git-modified-date: "2026-06-15" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- - [Redpanda Cloud Management MCP Server](overview/) Let AI agents securely operate your Redpanda Cloud clusters, topics, and users through natural language commands. - [Redpanda Cloud Management MCP Server Quickstart](quickstart/) Connect your Claude AI agent to your Redpanda Cloud account and clusters using the Redpanda Cloud Management MCP Server. - [Configure the Redpanda Cloud Management MCP Server](configuration/) Learn how to configure the Redpanda Cloud Management MCP Server, including auto and manual client setup, enabling deletes, and security considerations. --- # Page 13: Configure the Redpanda Cloud Management MCP Server **URL**: https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/configuration.md --- # Configure the Redpanda Cloud Management MCP Server > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure the Redpanda Cloud Management MCP Server page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cloud-mcp/configuration page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cloud-mcp/configuration.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/cloud-mcp/configuration.adoc # Beta release status page-beta: "true" description: Learn how to configure the Redpanda Cloud Management MCP Server, including auto and manual client setup, enabling deletes, and security considerations. page-topic-type: how-to personas: agent_developer, platform_admin learning-objective-1: Configure MCP clients learning-objective-2: Enable delete operations safely learning-objective-3: Troubleshoot common configuration issues page-git-created-date: "2026-06-15" page-git-modified-date: "2026-06-15" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- After installing the Redpanda Cloud Management MCP Server, you can configure it for different AI clients, customize security settings, and troubleshoot common issues. After reading this page, you will be able to: - Configure MCP clients - Enable delete operations safely - Troubleshoot common configuration issues ## [](#prerequisites)Prerequisites - At least version 25.2.3 of [`rpk` installed on your local machine](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) - Access to a Redpanda Cloud account - An MCP-compatible AI client such as Claude, Claude Code, or another tool that supports MCP > 💡 **TIP** > > The MCP server exposes Redpanda Cloud API endpoints for both the [Control Plane](https://docs.redpanda.com/api/doc/cloud-controlplane/) and the [Data Plane](https://docs.redpanda.com/api/doc/cloud-dataplane/). Available endpoints depend on your `rpk` version. Keep `rpk` updated to access new Redpanda Cloud features through the MCP server. New MCP endpoints are documented in Redpanda [release notes](https://github.com/redpanda-data/redpanda/releases). ## [](#install-the-integration-for-claude-or-claude-code)Install the integration for Claude or Claude Code For some supported clients, you can install and configure the MCP integration using the [`rpk cloud mcp install` command](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-mcp-install/). For Claude and Claude Code, run one of these commands: ```bash # Choose one rpk cloud mcp install --client claude rpk cloud mcp install --client claude-code ``` If you need to update the integration, re-run the install command for your client. ## [](#configure-other-mcp-clients-manually)Configure other MCP clients manually If you’re using another MCP-compatible client, manually configure it to use the Redpanda Cloud Management MCP Server. Follow these steps: Add an MCP server entry to your client’s configuration (example shown in JSON). Adjust paths for your system. ```json "mcpServers": { "redpandaCloud": { "command": "rpk", "args": [ "--config", "", (1) "cloud", "mcp", "stdio" ] } } ``` | 1 | Optional: The --config flag lets you target a specific rpk.yaml, which contains the configuration for connecting to your cluster. Always use the same configuration path as you used for rpk cloud login to ensure it has your token. Default paths vary by operating system. See the rpk cloud login reference for the default paths. | | --- | --- | You can also [start the server manually in a terminal to observe logs and troubleshoot](#local). ## [](#enable-delete-operations)Enable delete operations The server disables destructive operations by default. To allow delete operations, add `--allow-delete` to the MCP server invocation. > ⚠️ **CAUTION** > > Enabling delete operations permits actions like **deleting topics or clusters**. Restrict access to your AI client and double-check prompts. ### Auto-configured clients ```bash # Choose one rpk cloud mcp install --client claude --allow-delete rpk cloud mcp install --client claude-code --allow-delete ``` ### Manual configuration example ```json "mcpServers": { "redpandaCloud": { "command": "rpk", "args": [ "cloud", "mcp", "stdio", "--allow-delete" ] } } ``` ## [](#specify-configuration-file-paths)Specify configuration file paths All `rpk` commands accept a `--config` flag, which lets you specify the exact `rpk.yaml` configuration file to use for connecting to your Redpanda cluster. This flag overrides the default search path and ensures that the command uses the credentials and settings from the file you provide. Always use the same configuration path for both `rpk cloud login` and any MCP server setup or install commands to avoid authentication issues. By default, `rpk` searches for config files in standard locations depending on your operating system. See the [reference documentation](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-login/) for details. Use an absolute path and make sure your user has read and write permissions. > ⚠️ **CAUTION** > > The `rpk` configuration file contains your Redpanda Cloud token. Keep the file secure and never share it. For example, if you want to use a custom config path, specify it for both login and the MCP install command: ```bash rpk cloud login --config /Users//my-rpk-config.yaml rpk cloud mcp install --client claude --config /Users//my-rpk-config.yaml ``` Or for Claude Code: ```bash rpk cloud login --config /Users//my-rpk-config.yaml rpk cloud mcp install --client claude-code --config /Users//my-rpk-config.yaml ``` ## [](#remove-the-mcp-server)Remove the MCP server To remove the MCP server, delete or disable the `mcpServers.redpandaCloud` entry in your client’s config (steps vary by client). ## [](#security-considerations)Security considerations - Avoid enabling `--allow-delete` unless required. - For most local use cases, such as with Claude or Claude Code, log in with your personal Redpanda Cloud user account for better security and easier management. - If you are deploying the MCP server as part of an application or shared environment, consider using a [service account](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#authenticate-to-the-cloud-api) with tailored roles. To log in as a service account, use: ```bash rpk cloud login --client-id --client-secret --save ``` - Regularly review and rotate your credentials. ## [](#troubleshooting)Troubleshooting ### [](#verify-your-installation)Verify your installation 1. Make sure you are using at least version 25.2.3 of `rpk`. 2. If you see authentication errors, run `rpk cloud login` again. 3. Ensure you installed for the right client: ```bash rpk cloud mcp install --client claude # or rpk cloud mcp install --client claude-code ``` 4. If using another MCP client, verify your `mcpServers.redpandaCloud` entry (paths, JSON syntax, and args order). 5. Start the server manually using the [`rpk cloud mcp stdio` command](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-mcp-stdio/) (one-time login required) to verify connectivity to Redpanda Cloud endpoints: ```bash rpk cloud login rpk cloud mcp stdio ``` 1. Send the following newline-delimited JSON-RPC messages (each on its own line): ```json {"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{"roots":{},"sampling":{},"elicitation":{}},"clientInfo":{"name":"ManualTest","version":"0.1.0"}}} {"jsonrpc":"2.0","method":"notifications/initialized"} {"jsonrpc":"2.0","id":2,"method":"tools/list"} ``` Expected response shapes (examples): ```json {"jsonrpc":"2.0","id":1,"result":{"capabilities":{...}}} {"jsonrpc":"2.0","id":2,"result":{"tools":[{"name":"...","description":"..."}, ...]}} ``` 2. Stop the server with `Ctrl+C`. ### [](#client-cant-find-the-mcp-server)Client can’t find the MCP server - Re-run the install for your MCP client. - Confirm the path in `--config /path/to/rpk.yaml` exists and is readable. - Double-check your client’s configuration format and syntax. ### [](#unauthorized-errors-or-token-errors)Unauthorized errors or token errors Your capabilities depend on your Redpanda Cloud account permissions. If an operation fails with a permissions error, contact your account admin. - Run `rpk cloud login` to refresh the token. - Ensure your account has the necessary permissions for the requested operation. ### [](#deletes-not-working)Deletes not working - The server disables delete operations by default. Add `--allow-delete` to the server invocation (auto or manual configuration) and restart the client. - For auto-configured clients, you may need to edit the generated config or re-run the install command and adjust the entry. --- # Page 14: Redpanda Cloud Management MCP Server **URL**: https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/overview.md --- # Redpanda Cloud Management MCP Server > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Cloud Management MCP Server page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cloud-mcp/overview page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cloud-mcp/overview.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/cloud-mcp/overview.adoc # Beta release status page-beta: "true" description: Let AI agents securely operate your Redpanda Cloud clusters, topics, and users through natural language commands. page-topic-type: overview personas: evaluator, agent_developer, platform_admin learning-objective-1: Explain what the Redpanda Cloud Management MCP Server does learning-objective-2: Identify what operations are available through MCP learning-objective-3: Identify security considerations for MCP authentication page-git-created-date: "2026-06-15" page-git-modified-date: "2026-08-14" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- The Redpanda Cloud Management MCP Server lets AI agents securely access and operate your Redpanda Cloud account and clusters through natural language commands. After reading this page, you will be able to: - Explain what the Redpanda Cloud Management MCP Server does - Identify what operations are available through MCP - Identify security considerations for MCP authentication ![A terminal window showing Claude Code invoking the Redpanda Cloud Management MCP Server to list topics in a cluster.](https://docs.redpanda.com/cloud-data-platform/shared/_images/cloud-mcp.gif) ## [](#what-you-can-do)What you can do The MCP server provides controlled access to: - [Control Plane](https://docs.redpanda.com/api/doc/cloud-controlplane/) API endpoints, for example to create a Redpanda Cloud cluster or list clusters. - [Data Plane](https://docs.redpanda.com/api/doc/cloud-dataplane/) API endpoints, for example to create topics or list topics. The MCP server runs on your computer and authenticates to Redpanda Cloud using a Redpanda Cloud token. You can do anything that’s available in the Control Plane or Data Plane APIs. Typical requests you can make to your assistant once connected include: - Create a Redpanda Cloud cluster named `dev-mcp`. - List topics in `dev-mcp`. - Create a topic `orders-raw` with 6 partitions. > 📝 **NOTE** > > The MCP server does **not** expose delete endpoints by default. You can enable delete endpoints when you create the server if you intentionally want to allow delete operations. ## [](#use-cases)Use cases - Test automation: Create short-lived clusters, create topics, and validate pipelines quickly. - Operational assistance: Inspect a cluster’s health or list topics during incidents. - Onboarding and demos: Let team members issue high-level requests without memorizing every CLI flag. ## [](#how-it-works)How it works 1. Authenticate to Redpanda Cloud and receive a token using the [`rpk cloud login` command](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-login/). 2. Configure your MCP client using the [`rpk cloud mcp install`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-mcp-install/) command. Your client then starts the server on-demand using [`rpk cloud mcp stdio`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-mcp-stdio/), authenticating with the Redpanda Cloud token from `rpk cloud login`. 3. Prompt your assistant to perform Redpanda operations. The MCP server executes them in your Redpanda Cloud account using your Redpanda Cloud token. ### [](#components)Components The Redpanda Cloud Management MCP Server requires these components: - AI client (Claude, Claude Code, or any other MCP client) that connects to the MCP server. - Redpanda CLI (`rpk`) for obtaining a token and starting the MCP server. - Redpanda Cloud account that the MCP server can connect to and issue API requests. ## [](#security-considerations)Security considerations MCP servers authenticate to Redpanda Cloud using your personal or service account credentials. However, there is **no auditing or access control** that distinguishes between actions performed by MCP servers versus direct API calls: - All API actions appear in Redpanda Cloud’s internal logs as coming from the authenticated user account, not the specific MCP server. - You cannot audit which MCP server performed which operations, as Redpanda Cloud logs are not accessible to users. - You cannot restrict specific MCP servers to only certain API endpoints or resources. ## [](#next-steps)Next steps - [Redpanda Cloud Management MCP Server Quickstart](https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/quickstart/) - [Configure the Redpanda Cloud Management MCP Server](https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/configuration/) > 💡 **TIP** > > The Redpanda documentation site has a read-only MCP server that provides access to Redpanda docs and examples. This server has no access to your Redpanda Cloud account or clusters. See [How to Use These Docs](https://docs.redpanda.com/home/how-to-use-these-docs/). --- # Page 15: Redpanda Cloud Management MCP Server Quickstart **URL**: https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/quickstart.md --- # Redpanda Cloud Management MCP Server Quickstart > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Cloud Management MCP Server Quickstart page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cloud-mcp/quickstart page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cloud-mcp/quickstart.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/cloud-mcp/quickstart.adoc # Beta release status page-beta: "true" description: Connect your Claude AI agent to your Redpanda Cloud account and clusters using the Redpanda Cloud Management MCP Server. page-topic-type: tutorial personas: agent_developer, platform_admin learning-objective-1: Authenticate to Redpanda Cloud with rpk learning-objective-2: Install the MCP integration for Claude learning-objective-3: Issue natural language commands to manage clusters page-git-created-date: "2026-06-15" page-git-modified-date: "2026-06-15" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- In this quickstart, you’ll get your Claude AI agent talking to Redpanda Cloud using the [Redpanda Cloud Management MCP Server](https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/overview/). After completing this quickstart, you will be able to: - Authenticate to Redpanda Cloud with rpk - Install the MCP integration for Claude - Issue natural language commands to manage clusters ## [](#prerequisites)Prerequisites - At least version 25.2.3 of [`rpk` installed on your computer](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) - Access to a Redpanda Cloud account - [Claude](https://support.anthropic.com/en/articles/10065433-installing-claude-desktop) or [Claude Code](https://docs.anthropic.com/en/docs/claude-code/setup) installed > 💡 **TIP** > > For other clients, see [Configure the Redpanda Cloud Management MCP Server](https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/configuration/). ## [](#set-up-the-mcp-server)Set up the MCP server 1. Verify your `rpk` version. ```bash rpk version ``` Ensure the version is at least 25.2.3. 2. Log in to Redpanda Cloud. ```bash rpk cloud login ``` A browser window opens. Sign in to grant access. After you sign in, `rpk` stores a token locally. This token is not shared with your AI agent. It is used by the MCP server to authenticate requests to your Redpanda Cloud account. 3. Install the MCP integration. Choose one client: ```bash # Claude desktop rpk cloud mcp install --client claude # Claude Code (IDE) rpk cloud mcp install --client claude-code ``` This command configures the MCP server for your client. If you need to update the integration, re-run the install command for your client. ## [](#start-prompting)Start prompting Launch Claude or Claude Code and try one of these prompts: - “Create a Redpanda Cloud cluster named `dev-mcp`.” - “List topics in `dev-mcp`.” - “Create a topic `orders-raw` with 6 partitions.” > 📝 **NOTE: Delete operations are opt-in** > > The MCP server does **not** expose API endpoints that result in delete operations by default. Use `--allow-delete` only if you intentionally want to enable delete operations. See [Enable delete operations](https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/configuration/#enable_delete_operations). ## [](#next-steps)Next steps - [Configure the Redpanda Cloud Management MCP Server](https://docs.redpanda.com/cloud-data-platform/develop/cloud-mcp/configuration/) --- # Page 16: Redpanda Connect in Redpanda Cloud **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/about.md --- # Redpanda Connect in Redpanda Cloud > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Connect in Redpanda Cloud latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/about.adoc description: Learn about Redpanda Connect in Redpanda Cloud and its wide range of connectors. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Redpanda Connect in Redpanda Cloud lets you quickly build and deploy streaming data pipelines on your clusters from a fully-integrated UI or using the [Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/group/endpoint-redpanda-connect-pipeline). Choose from a [wide range of connectors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/about/) to suit your use case, including connectors to: - Integrate data sources ([inputs](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/about/)) - Write to data sinks ([outputs](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about/)) - Transform data ([processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/)) Comprehensive data pipeline metrics are also available to help you to [monitor your data pipelines](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/monitor-connect/) and [per pipeline scaling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/resource-management/). Try this [quickstart](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/). > 💡 **TIP** > > If you’re new to Redpanda Connect, try [building and testing data pipelines locally](https://docs.redpanda.com/connect/get-started/quickstarts/rpk/) before deploying to the Cloud. --- # Page 17: Components Catalog **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/about.md --- # Components Catalog > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Components Catalog latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/about.adoc description: A searchable catalog of available Redpanda Connect components. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Use the following table to search for available inputs, outputs, and processors. Type: All Types Selected ▼ Processor Input Output Scanner Metric Cache Tracer Rate limit Buffer | Name | Connector Type | | --- | --- | | a2a_message | Processor | | amqp_0_9RabbitMQ AMQP | Input, Output | | arc | Output | | archiveZIP TAR GZIP | Processor | | avro | Processor, Scanner | | aws_bedrock_chatAmazon AWS Bedrock Chat | Processor | | aws_bedrock_embeddingsAmazon AWS Bedrock Embeddings | Processor | | aws_cloudwatch_logsAWS CloudWatch Logs Amazon CloudWatch Logs | Input | | aws_dynamodbAWS DynamoDB Amazon DynamoDB DynamoDB | Cache, Output | | aws_dynamodb_cdcAmazon DynamoDB CDC | Input | | aws_dynamodb_partiqlAmazon AWS DynamoDB PartiQL | Processor | | aws_kinesisAWS Kinesis Amazon Kinesis Kinesis | Input, Output | | aws_kinesis_firehoseAWS Kinesis Firehose Amazon Kinesis Firehose Kinesis Firehose | Output | | aws_lambdaAWS Lambda Amazon Lambda Lambda | Processor | | aws_s3AWS S3 Amazon S3 S3 Simple Storage Service | Cache, Input, Output | | aws_snsAWS SNS Amazon SNS SNS Simple Notification Service | Output | | aws_sqsAWS SQS Amazon SQS SQS Simple Queue Service | Input, Output | | azure_blob_storageAzure Blob Storage Microsoft Azure Storage | Input, Output | | azure_cosmosdbMicrosoft Azure Azure | Input, Output, Processor | | azure_data_lake_gen2Microsoft Azure Azure | Output | | azure_queue_storageAzure Queue Storage Microsoft Azure Queue | Input, Output | | azure_table_storageAzure Table Storage Microsoft Azure Table | Input, Output | | batched | Input | | benchmark | Processor | | bloblang | Processor | | bounds_check | Processor | | branch | Processor | | broker | Input, Output | | cache | Output, Processor | | cached | Processor | | catch | Processor | | chunker | Scanner | | cohere_chat | Processor | | cohere_embeddings | Processor | | cohere_rerank | Processor | | compress | Processor | | csvComma-Separated Values | Scanner | | cyborgdb | Output | | decompress | Processor, Scanner | | dedupe | Processor | | drop | Output | | drop_on | Output | | elasticsearch_v8 | Output | | fallback | Output | | for_each | Processor | | gateway | Input | | gcp_bigqueryGCP BigQuery Google BigQuery BigQuery | Output | | gcp_bigquery_selectGCP BigQuery Google Cloud GCP | Input, Processor | | gcp_bigquery_write_apiGCP BigQuery | Output | | gcp_cloud_storageGCP Cloud Storage Google Cloud Storage GCS | Cache, Input, Output | | gcp_cloudtraceGCP Cloud Trace | Tracer | | gcp_pubsubGCP PubSub Google Cloud Pub/Sub GCP Pub/Sub Google Pub/Sub | Input, Output | | gcp_spanner_cdcGoogle Cloud GCP | Input | | gcp_vertex_ai_chatGCP Vertex AI Google Cloud GCP | Processor | | gcp_vertex_ai_embeddingsGoogle Cloud GCP | Processor | | generate | Input | | git | Input | | google_drive_download | Processor | | google_drive_list_labels | Processor | | google_drive_search | Processor | | group_by | Processor | | group_by_value | Processor | | http | Processor | | http_clientHTTP REST API REST | Input, Output | | http_serverHTTP REST API REST Gateway | Input | | icebergApache Iceberg Apache Polaris AWS Glue Databricks Unity Catalog | Output | | inproc | Input, Output | | insert_part | Processor | | jiraAtlassian Jira | Input, Processor | | jmespath | Processor | | jq | Processor | | json_array | Scanner | | json_documents | Scanner | | json_schemaJSON Schema | Processor | | kafkaApache Kafka | Input, Output | | kafka_franzApache Kafka Kafka | Input, Output | | lines | Scanner | | local | Rate_limit | | log | Processor | | lru | Cache | | mapping | Processor | | memcached | Cache | | memory | Buffer, Cache | | metric | Processor | | microsoft_sql_server_cdc | Input | | mongodbMongo | Cache, Input, Output, Processor | | mongodb_cdcMongoDB CDC | Input | | mqtt | Input, Output | | multilevel | Cache | | mutation | Processor | | mysql_cdc | Input | | natsNATS.io | Input, Output | | nats_jetstreamNATS JetStream NATS | Input, Output | | nats_kvNATS KV | Cache, Input, Output, Processor | | nats_request_replyNATS Request Reply | Processor | | none | Buffer, Metric, Tracer | | noop | Cache, Processor | | open_telemetry_collectorOpenTelemetry | Metric, Tracer | | openai_chat_completion | Processor | | openai_embeddings | Processor | | openai_image_generation | Processor | | openai_speech | Processor | | openai_transcription | Processor | | openai_translation | Processor | | opensearch | Output | | oracledb_cdcOracle CDC OracleDB CDC Oracle Database CDC | Input | | otlp_grpcOpenTelemetry OTLP OTel gRPC | Input, Output | | otlp_httpOpenTelemetry OTLP OTel | Input, Output | | parallel | Processor | | parquet_decode | Processor | | parquet_encode | Processor | | parse_log | Processor | | pg_stream | | | pinecone | Output | | postgres_cdc | Input | | processors | Processor | | prometheus | Metric | | qdrant | Output, Processor | | questdb | Output | | rate_limit | Processor | | re_match | Scanner | | read_until | Input | | redis | Cache, Processor, Rate_limit | | redis_hashRedis Hash Redis | Output | | redis_listRedis List Redis Lists Redis | Input, Output | | redis_pubsubRedis PubSub Redis Pub/Sub Redis | Input, Output | | redis_scanRedis | Input | | redis_scriptRedis Script | Processor | | redis_streamsRedis Streams Redis | Input, Output | | redpanda | Cache, Input, Output, Tracer | | redpanda_common | Input, Output | | redpanda_migrator | Input, Output | | reject | Output | | reject_errored | Output | | resource | Input, Output, Processor | | retry | Output, Processor | | ristretto | Cache | | salesforce | Input | | salesforce_cdcSalesforce Salesforce CDC | Input | | salesforce_graphqlSalesforce Salesforce GraphQL | Input | | salesforce_sinkSalesforce Salesforce Sink | Output | | schema_registry | Input, Output | | schema_registry_decode | Processor | | schema_registry_encode | Processor | | select_parts | Processor | | sequence | Input | | sftp | Input, Output | | skip_bom | Scanner | | slack | Input | | slack_postSlack Post | Output | | slack_reactionSlack Reaction | Output | | slack_threadSlack Thread | Processor | | slack_usersSlack Users | Input | | sleep | Processor | | snowflake_putSnowflake | Output | | snowflake_streamingSnowflake Streaming | Output | | spicedb_watch | Input | | split | Processor | | splunk | Input | | splunk_hecSplunk | Output | | sql | Cache | | sql_driver_clickhouseClickHouse | | | sql_driver_mysqlMYSQL | | | sql_driver_oracleOracle | | | sql_driver_postgresPostgreSQL | | | sql_driver_sqliteSQLite | | | sql_insertSQL PostgreSQL MySQL Microsoft SQL Server ClickHouse Trino | Output, Processor | | sql_rawSQL PostgreSQL MySQL Microsoft SQL Server ClickHouse Trino | Input, Output, Processor | | sql_selectSQL PostgreSQL MySQL Microsoft SQL Server ClickHouse Trino | Input, Processor | | string_split | Processor | | switch | Output, Processor, Scanner | | sync_response | Output, Processor | | system_window | Buffer | | tar | Scanner | | text_chunker | Processor | | timeplus | Input, Output | | to_the_end | Scanner | | try | Processor | | try_catch | Processor | | ttlru | Cache | | unarchiveZIP TAR GZIP Archive | Processor | | while | Processor | | workflow | Processor | | xml | Processor | ## [](#about-components)About Components Every Redpanda Connect pipeline has at least one [input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/about/), an optional [buffer](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/buffers/about/), an [output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about/) and any number of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/): ```yaml input: kafka: addresses: [ TODO ] topics: [ foo, bar ] consumer_group: foogroup buffer: type: none pipeline: processors: - mapping: | message = this meta.link_count = links.length() output: aws_s3: bucket: TODO path: '${! meta("kafka_topic") }/${! json("message.id") }.json' ``` These are the main components within Redpanda Connect and they provide the majority of useful behavior. ## [](#observability-components)Observability components There are also the observability components: [logger](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/logger/about/), [metrics](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/), and [tracing](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/about/), which allow you to specify how Redpanda Connect exposes observability data. ```yaml http: address: 0.0.0.0:4195 enabled: true debug_endpoints: false logger: format: json level: WARN metrics: statsd: address: localhost:8125 flush_period: 100ms tracer: jaeger: agent_address: localhost:6831 ``` ## [](#resource-components)Resource components Finally, there are [caches](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) and [rate limits](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/). These are components that are referenced by core components and can be shared. ```yaml input: http_client: # This is an input url: TODO rate_limit: foo_ratelimit # This is a reference to a rate limit pipeline: processors: - cache: # This is a processor resource: baz_cache # This is a reference to a cache operator: add key: '${! json("id") }' value: "x" - mapping: root = if errored() { deleted() } rate_limit_resources: - label: foo_ratelimit local: count: 500 interval: 1s cache_resources: - label: baz_cache memcached: addresses: [ localhost:11211 ] ``` It’s also possible to configure inputs, outputs and processors as resources which allows them to be reused throughout a configuration with the [`resource` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/resource/), [`resource` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/resource/) and [`resource` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/resource/) respectively. For more information about any of these component types check out their sections: - [inputs](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/about/) - [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) - [outputs](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about/) - [buffers](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/buffers/about/) - [metrics](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/) - [tracers](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/about/) - [logger](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/logger/about/) - [caches](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) - [rate limits](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) --- # Page 18: Buffers **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/buffers/about.md --- # Buffers > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Buffers latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/buffers/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/buffers/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/buffers/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Redpanda Connect uses a transaction model internally for guaranteeing delivery of messages, this means that a message from an input is not acknowledged (or its offset committed, etc) until that message has been processed and either intentionally deleted or successfully delivered to all outputs. This transaction model makes Redpanda Connect safe to deploy in scenarios where data loss is unacceptable. However, sometimes it’s useful to customize the way in which messages are delivered, and this is where buffers come in. A buffer is an optional component type that comes immediately after the input layer and can be used as a way of decoupling the transaction model from components downstream such as the processing layer and outputs. This is considered an advanced component as most users will likely not benefit from a buffer, but they enable you to do things like group messages using window algorithms or intentionally weaken the delivery guarantees of the pipeline depending on the buffer you choose. Since buffers are able to modify (or disable) the transaction model within Redpanda Connect it is important that when you choose a buffer you read its documentation to understand the implication it will have on delivery guarantees. --- # Page 19: memory **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/buffers/memory.md --- # memory > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: memory latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/buffers/memory page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/buffers/memory.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/buffers/memory.adoc description: Stores consumed messages in memory and acknowledges them at the input level. During shutdown Redpanda Connect will make a best attempt at flushing all remaining messages before exiting cleanly. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores consumed messages in memory and acknowledges them at the input level. During shutdown Redpanda Connect will make a best attempt at flushing all remaining messages before exiting cleanly. #### Common ```yml buffers: memory: limit: 524288000 batch_policy: enabled: false count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml buffers: memory: limit: 524288000 batch_policy: enabled: false count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` This buffer is appropriate when consuming messages from inputs that do not gracefully handle back pressure and where delivery guarantees aren’t critical. This buffer has a configurable limit, where consumption will be stopped with back pressure upstream if the total size of messages in the buffer reaches this amount. Since this calculation is only an estimate, and the real size of messages in RAM is always higher, it is recommended to set the limit significantly below the amount of RAM available. ## [](#delivery-guarantees)Delivery guarantees This buffer intentionally weakens the delivery guarantees of the pipeline and therefore should never be used in places where data loss is unacceptable. ## [](#batching)Batching It is possible to batch up messages sent from this buffer using a [batch policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/#batch-policy). ## [](#fields)Fields ### [](#batch_policy)`batch_policy` Optionally configure a policy to flush buffered messages in batches. **Type**: `object` ### [](#batch_policy-byte_size)`batch_policy.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batch_policy-check)`batch_policy.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batch_policy-count)`batch_policy.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batch_policy-enabled)`batch_policy.enabled` Whether to batch messages as they are flushed. **Type**: `bool` **Default**: `false` ### [](#batch_policy-period)`batch_policy.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batch_policy-processors)`batch_policy.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#limit)`limit` The maximum buffer size (in bytes) to allow before applying backpressure upstream. **Type**: `int` **Default**: `524288000` --- # Page 20: none **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/buffers/none.md --- # none > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: none latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/buffers/none page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/buffers/none.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/buffers/none.adoc description: Do not buffer messages. This is the default and most resilient configuration. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Do not buffer messages. This is the default and most resilient configuration. ```yml # Config fields, showing default values buffer: none: {} ``` Selecting no buffer means the output layer is directly coupled with the input layer. This is the safest and lowest latency option since acknowledgements from at-least-once protocols can be propagated all the way from the output protocol to the input protocol. If the output layer is hit with back pressure it will propagate all the way to the input layer, and further up the data stream. If you need to relieve your pipeline of this back pressure consider using a more robust buffering solution such as Kafka before resorting to alternatives. --- # Page 21: system_window **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/buffers/system_window.md --- # system_window > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: system_window latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/buffers/system_window page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/buffers/system_window.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/buffers/system_window.adoc description: Chops a stream of messages into tumbling or sliding windows of fixed temporal size, following the system clock. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Chops a stream of messages into tumbling or sliding windows of fixed temporal size, following the system clock. #### Common ```yml buffers: system_window: timestamp_mapping: root = now() size: "" # No default (required) slide: "" offset: "" allowed_lateness: "" ``` #### Advanced ```yml buffers: system_window: timestamp_mapping: root = now() size: "" # No default (required) slide: "" offset: "" allowed_lateness: "" ``` A window is a grouping of messages that fit within a discrete measure of time following the system clock. Messages are allocated to a window either by the processing time (the time at which they’re ingested) or by the event time, and this is controlled via the [`timestamp_mapping` field](#timestamp_mapping). In tumbling mode (default) the beginning of a window immediately follows the end of a prior window. When the buffer is initialized the first window to be created and populated is aligned against the zeroth minute of the zeroth hour of the day by default, and may therefore be open for a shorter period than the specified size. A window is flushed only once the system clock surpasses its scheduled end. If an [`allowed_lateness`](#allowed_lateness) is specified then the window will not be flushed until the scheduled end plus that length of time. When a message is added to a window it has a metadata field `window_end_timestamp` added to it containing the timestamp of the end of the window as an RFC3339 string. ## [](#sliding-windows)Sliding windows Sliding windows begin from an offset of the prior windows' beginning rather than its end, and therefore messages may belong to multiple windows. In order to produce sliding windows specify a [`slide` duration](#slide). ## [](#back-pressure)Back pressure If back pressure is applied to this buffer either due to output services being unavailable or resources being saturated, windows older than the current and last according to the system clock will be dropped in order to prevent unbounded resource usage. This means you should ensure that under the worst case scenario you have enough system memory to store two windows' worth of data at a given time (plus extra for redundancy and other services). If messages could potentially arrive with event timestamps in the future (according to the system clock) then you should also factor in these extra messages in memory usage estimates. ## [](#delivery-guarantees)Delivery guarantees This buffer honours the transaction model within Redpanda Connect in order to ensure that messages are not acknowledged until they are either intentionally dropped or successfully delivered to outputs. However, since messages belonging to an expired window are intentionally dropped there are circumstances where not all messages entering the system will be delivered. When this buffer is configured with a slide duration it is possible for messages to belong to multiple windows, and therefore be delivered multiple times. In this case the first time the message is delivered it will be acked (or nacked) and subsequent deliveries of the same message will be a "best attempt". During graceful termination if the current window is partially populated with messages they will be nacked such that they are re-consumed the next time the service starts. ## [](#examples)Examples ### Counting Passengers at Traffic Given a stream of messages relating to cars passing through various traffic lights of the form: ```json { "traffic_light": "cbf2eafc-806e-4067-9211-97be7e42cee3", "created_at": "2021-08-07T09:49:35Z", "registration_plate": "AB1C DEF", "passengers": 3 } ``` We can use a window buffer in order to create periodic messages summarizing the traffic for a period of time of this form: ```json { "traffic_light": "cbf2eafc-806e-4067-9211-97be7e42cee3", "created_at": "2021-08-07T10:00:00Z", "total_cars": 15, "passengers": 43 } ``` With the following config: ```yaml buffer: system_window: timestamp_mapping: root = this.created_at size: 1h pipeline: processors: # Group messages of the window into batches of common traffic light IDs - group_by_value: value: '${! json("traffic_light") }' # Reduce each batch to a single message by deleting indexes > 0, and # aggregate the car and passenger counts. - mapping: | root = if batch_index() == 0 { { "traffic_light": this.traffic_light, "created_at": meta("window_end_timestamp"), "total_cars": json("registration_plate").from_all().unique().length(), "passengers": json("passengers").from_all().sum(), } } else { deleted() } ``` ## [](#fields)Fields ### [](#timestamp_mapping)`timestamp_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) applied to each message during ingestion that provides the timestamp to use for allocating it a window. By default the function `now()` is used in order to generate a fresh timestamp at the time of ingestion (the processing time), whereas this mapping can instead extract a timestamp from the message itself (the event time). The timestamp value assigned to `root` must either be a numerical unix time in seconds (with up to nanosecond precision via decimals), or a string in ISO 8601 format. If the mapping fails or provides an invalid result the message will be dropped (with logging to describe the problem). **Type**: `string` **Default**: `"root = now()"` ```yml # Examples timestamp_mapping: root = this.created_at timestamp_mapping: root = meta("kafka_timestamp_unix").number() ``` ### [](#size)`size` A duration string describing the size of each window. By default windows are aligned to the zeroth minute and zeroth hour on the UTC clock, meaning windows of 1 hour duration will match the turn of each hour in the day, this can be adjusted with the `offset` field. **Type**: `string` ```yml # Examples size: 30s size: 10m ``` ### [](#slide)`slide` An optional duration string describing by how much time the beginning of each window should be offset from the beginning of the previous, and therefore creates sliding windows instead of tumbling. When specified this duration must be smaller than the `size` of the window. **Type**: `string` **Default**: `""` ```yml # Examples slide: 30s slide: 10m ``` ### [](#offset)`offset` An optional duration string to offset the beginning of each window by, otherwise they are aligned to the zeroth minute and zeroth hour on the UTC clock. The offset cannot be a larger or equal measure to the window size or the slide. **Type**: `string` **Default**: `""` ```yml # Examples offset: -6h offset: 30m ``` ### [](#allowed_lateness)`allowed_lateness` An optional duration string describing the length of time to wait after a window has ended before flushing it, allowing late arrivals to be included. Since this windowing buffer uses the system clock an allowed lateness can improve the matching of messages when using event time. **Type**: `string` **Default**: `""` ```yml # Examples allowed_lateness: 10s allowed_lateness: 1m ``` --- # Page 22: Caches **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about.md --- # Caches > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Caches latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A cache is a key/value store which can be used by certain components for applications such as deduplication or data joins. Caches are configured as a named resource: ```yaml cache_resources: - label: foobar memcached: addresses: - localhost:11211 default_ttl: 60s ``` > It’s possible to layer caches with read-through and write-through behavior using the [`multilevel` cache](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/multilevel/). And then any components that use caches have a field `resource` that specifies the cache resource: ```yaml pipeline: processors: - cache: resource: foobar operator: add key: '${! json("message.id") }' value: "storeme" - mapping: root = if errored() { deleted() } ``` For the simple case where you wish to store messages in a cache as an output destination for your pipeline check out the [`cache` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/cache/). To see examples of more advanced uses of caches such as hydration and deduplication check out the [`cache` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cache/). --- # Page 23: aws_dynamodb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/aws_dynamodb.md --- # aws_dynamodb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_dynamodb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/aws_dynamodb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/aws_dynamodb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/aws_dynamodb.adoc description: Stores key/value pairs as a single document in a DynamoDB table. The key is stored as a string value and used as the table hash key. The value is stored as a binary value using the data_key field name. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores key/value pairs as a single document in a DynamoDB table. The key is stored as a string value and used as the table hash key. The value is stored as a binary value using the `data_key` field name. #### Common ```yml caches: aws_dynamodb: table: "" # No default (required) hash_key: "" # No default (required) data_key: "" # No default (required) ``` #### Advanced ```yml caches: aws_dynamodb: table: "" # No default (required) hash_key: "" # No default (required) data_key: "" # No default (required) consistent_read: false default_ttl: "" # No default (optional) ttl_key: "" # No default (optional) retries: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` A prefix can be specified to allow multiple cache types to share a single DynamoDB table. An optional TTL duration (`ttl`) and field (`ttl_key`) can be specified if the backing table has TTL enabled. Strong read consistency can be enabled using the `consistent_read` configuration field. ## [](#fields)Fields ### [](#consistent_read)`consistent_read` Whether to use strongly consistent reads on Get commands. **Type**: `bool` **Default**: `false` ### [](#credentials)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#data_key)`data_key` The key of the table column to store item values within. **Type**: `string` ### [](#default_ttl)`default_ttl` An optional default TTL to set for items, calculated from the moment the item is cached. A `ttl_key` must be specified in order to set item TTLs. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#hash_key)`hash_key` The key of the table column to store item keys within. **Type**: `string` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#retries)`retries` Determine time intervals and cut offs for retry attempts. **Type**: `object` ### [](#retries-initial_interval)`retries.initial_interval` The initial period to wait between retry attempts. **Type**: `string` **Default**: `1s` ```yaml # Examples: initial_interval: 50ms # --- initial_interval: 1s ``` ### [](#retries-max_elapsed_time)`retries.max_elapsed_time` The maximum overall period of time to spend on retry attempts before the request is aborted. **Type**: `string` **Default**: `30s` ```yaml # Examples: max_elapsed_time: 1m # --- max_elapsed_time: 1h ``` ### [](#retries-max_interval)`retries.max_interval` The maximum period to wait between retry attempts **Type**: `string` **Default**: `5s` ```yaml # Examples: max_interval: 5s # --- max_interval: 1m ``` ### [](#table)`table` The table to store items in. **Type**: `string` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#ttl_key)`ttl_key` The column key to place the TTL value within. **Type**: `string` --- # Page 24: aws_s3 **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/aws_s3.md --- # aws_s3 > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_s3 latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/aws_s3 page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/aws_s3.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/aws_s3.adoc description: Stores each item in an S3 bucket as a file, where an item ID is the path of the item within the bucket. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores each item in an S3 bucket as a file, where an item ID is the path of the item within the bucket. #### Common ```yml caches: aws_s3: bucket: "" # No default (required) content_type: application/octet-stream ``` #### Advanced ```yml caches: aws_s3: bucket: "" # No default (required) content_type: application/octet-stream force_path_style_urls: false retries: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` It is not possible to atomically upload S3 objects exclusively when the target does not already exist, therefore this cache is not suitable for deduplication. ## [](#fields)Fields ### [](#bucket)`bucket` The S3 bucket to store items in. **Type**: `string` ### [](#content_type)`content_type` The content type to set for each item. **Type**: `string` **Default**: `application/octet-stream` ### [](#credentials)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#force_path_style_urls)`force_path_style_urls` Forces the client API to use path style URLs, which helps when connecting to custom endpoints. **Type**: `bool` **Default**: `false` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#retries)`retries` Determine time intervals and cut offs for retry attempts. **Type**: `object` ### [](#retries-initial_interval)`retries.initial_interval` The initial period to wait between retry attempts. **Type**: `string` **Default**: `1s` ```yaml # Examples: initial_interval: 50ms # --- initial_interval: 1s ``` ### [](#retries-max_elapsed_time)`retries.max_elapsed_time` The maximum overall period of time to spend on retry attempts before the request is aborted. **Type**: `string` **Default**: `30s` ```yaml # Examples: max_elapsed_time: 1m # --- max_elapsed_time: 1h ``` ### [](#retries-max_interval)`retries.max_interval` The maximum period to wait between retry attempts **Type**: `string` **Default**: `5s` ```yaml # Examples: max_interval: 5s # --- max_interval: 1m ``` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` --- # Page 25: gcp_cloud_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/gcp_cloud_storage.md --- # gcp_cloud_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_cloud_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/gcp_cloud_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/gcp_cloud_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/gcp_cloud_storage.adoc description: Use a Google Cloud Storage bucket as a cache. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Use a Google Cloud Storage bucket as a cache. ```yml caches: gcp_cloud_storage: bucket: "" # No default (required) content_type: "" # No default (optional) credentials_json: "" ``` It is not possible to atomically upload cloud storage objects exclusively when the target does not already exist, therefore this cache is not suitable for deduplication. ## [](#fields)Fields ### [](#bucket)`bucket` The Google Cloud Storage bucket to store items in. **Type**: `string` ### [](#content_type)`content_type` Optional field to explicitly set the Content-Type. **Type**: `string` ### [](#credentials_json)`credentials_json` An optional field to set Google Service Account Credentials json. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` --- # Page 26: lru **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/lru.md --- # lru > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: lru latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/lru page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/lru.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/lru.adoc description: Stores key/value pairs in a lru in-memory cache. This cache is therefore reset every time the service restarts. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores key/value pairs in a lru in-memory cache. This cache is therefore reset every time the service restarts. #### Common ```yml caches: lru: cap: 1000 init_values: {} ``` #### Advanced ```yml caches: lru: cap: 1000 init_values: {} algorithm: standard two_queues_recent_ratio: 0.25 two_queues_ghost_ratio: 0.5 optimistic: false ``` This provides the lru package which implements a fixed-size thread safe LRU cache. It uses the package [`lru`](https://github.com/hashicorp/golang-lru/v2) The field init\_values can be used to pre-populate the memory cache with any number of key/value pairs: ```yaml cache_resources: - label: foocache lru: cap: 1024 init_values: foo: bar ``` These values can be overridden during execution. ## [](#fields)Fields ### [](#algorithm)`algorithm` the lru cache implementation **Type**: `string` **Default**: `standard` | Option | Summary | | --- | --- | | arc | is an adaptive replacement cache. It tracks recent evictions as well as recent usage in both the frequent and recent caches. Its computational overhead is comparable to two_queues, but the memory overhead is linear with the size of the cache. ARC has been patented by IBM. | | standard | is a simple LRU cache. It is based on the LRU implementation in groupcache | | two_queues | tracks frequently used and recently used entries separately. This avoids a burst of accesses from taking out frequently used entries, at the cost of about 2x computational overhead and some extra bookkeeping. | ### [](#cap)`cap` The cache maximum capacity (number of entries) **Type**: `int` **Default**: `1000` ### [](#init_values)`init_values` A table of key/value pairs that should be present in the cache on initialization. This can be used to create static lookup tables. **Type**: `object` **Default**: `{}` ```yaml # Examples: init_values: Nickelback: "1995" Spice Girls: "1994" The Human League: "1977" ``` ### [](#optimistic)`optimistic` If true, we do not lock on read/write events. The lru package is thread-safe, however the ADD operation is not atomic. **Type**: `bool` **Default**: `false` ### [](#two_queues_ghost_ratio)`two_queues_ghost_ratio` is the default ratio of ghost entries kept to track entries recently evicted on two\_queues cache. **Type**: `float` **Default**: `0.5` ### [](#two_queues_recent_ratio)`two_queues_recent_ratio` is the ratio of the two\_queues cache dedicated to recently added entries that have only been accessed once. **Type**: `float` **Default**: `0.25` --- # Page 27: memcached **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/memcached.md --- # memcached > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: memcached latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/memcached page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/memcached.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/memcached.adoc description: Connects to a cluster of memcached services, a prefix can be specified to allow multiple cache types to share a memcached cluster under different namespaces. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Connects to a cluster of memcached services, a prefix can be specified to allow multiple cache types to share a memcached cluster under different namespaces. #### Common ```yml caches: memcached: addresses: [] # No default (required) prefix: "" # No default (optional) default_ttl: 300s ``` #### Advanced ```yml caches: memcached: addresses: [] # No default (required) prefix: "" # No default (optional) default_ttl: 300s retries: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s ``` ## [](#fields)Fields ### [](#addresses)`addresses[]` A list of addresses of memcached servers to use. **Type**: `array` ### [](#default_ttl)`default_ttl` A default TTL to set for items, calculated from the moment the item is cached. **Type**: `string` **Default**: `300s` ### [](#prefix)`prefix` An optional string to prefix item keys with in order to prevent collisions with similar services. **Type**: `string` ### [](#retries)`retries` Determine time intervals and cut offs for retry attempts. **Type**: `object` ### [](#retries-initial_interval)`retries.initial_interval` The initial period to wait between retry attempts. **Type**: `string` **Default**: `1s` ```yaml # Examples: initial_interval: 50ms # --- initial_interval: 1s ``` ### [](#retries-max_elapsed_time)`retries.max_elapsed_time` The maximum overall period of time to spend on retry attempts before the request is aborted. **Type**: `string` **Default**: `30s` ```yaml # Examples: max_elapsed_time: 1m # --- max_elapsed_time: 1h ``` ### [](#retries-max_interval)`retries.max_interval` The maximum period to wait between retry attempts **Type**: `string` **Default**: `5s` ```yaml # Examples: max_interval: 5s # --- max_interval: 1m ``` --- # Page 28: memory **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/memory.md --- # memory > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: memory latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/memory page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/memory.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/memory.adoc description: Stores key/value pairs in a map held in memory. This cache is therefore reset every time the service restarts. Each item in the cache has a TTL set from the moment it was last edited, after which it will be removed during the next compaction. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores key/value pairs in a map held in memory. This cache is therefore reset every time the service restarts. Each item in the cache has a TTL set from the moment it was last edited, after which it will be removed during the next compaction. #### Common ```yml caches: memory: default_ttl: 5m compaction_interval: 60s init_values: {} ``` #### Advanced ```yml caches: memory: default_ttl: 5m compaction_interval: 60s init_values: {} shards: 1 ``` The compaction interval determines how often the cache is cleared of expired items, and this process is only triggered on writes to the cache. Access to the cache is blocked during this process. Item expiry can be disabled entirely by setting the `compaction_interval` to an empty string. The field `init_values` can be used to prepopulate the memory cache with any number of key/value pairs which are exempt from TTLs: ```yaml cache_resources: - label: foocache memory: default_ttl: 60s init_values: foo: bar ``` These values can be overridden during execution, at which point the configured TTL is respected as usual. ## [](#fields)Fields ### [](#compaction_interval)`compaction_interval` The period of time to wait before each compaction, at which point expired items are removed. This field can be set to an empty string in order to disable compactions/expiry entirely. **Type**: `string` **Default**: `60s` ### [](#default_ttl)`default_ttl` The default TTL of each item. After this period an item will be eligible for removal during the next compaction. **Type**: `string` **Default**: `5m` ### [](#init_values)`init_values` A table of key/value pairs that should be present in the cache on initialization. This can be used to create static lookup tables. **Type**: `object` **Default**: `{}` ```yaml # Examples: init_values: Nickelback: "1995" Spice Girls: "1994" The Human League: "1977" ``` ### [](#shards)`shards` A number of logical shards to spread keys across, increasing the shards can have a performance benefit when processing a large number of keys. **Type**: `int` **Default**: `1` --- # Page 29: mongodb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/mongodb.md --- # mongodb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mongodb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/mongodb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/mongodb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/mongodb.adoc description: Use a MongoDB instance as a cache. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- Use a MongoDB instance as a cache. #### Common ```yml caches: mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" collection: "" # No default (required) key_field: "" # No default (required) value_field: "" # No default (required) ``` #### Advanced ```yml caches: mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" app_name: benthos aws: enabled: false region: "" # No default (optional) session_duration: 1h id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) roles: [] # No default (optional) collection: "" # No default (required) key_field: "" # No default (required) value_field: "" # No default (required) ``` ## [](#fields)Fields ### [](#app_name)`app_name` The client application name. **Type**: `string` **Default**: `benthos` ### [](#aws)`aws` AWS IAM authentication using the `MONGODB-AWS` mechanism, for example against MongoDB Atlas. When enabled, IAM credentials are used instead of a static username and password. Role-derived session credentials are resolved when the component connects and are re-resolved whenever it reconnects. The `mongodb` processor and cache establish their client once at creation and cannot refresh expiring session credentials, so `role`, `roles` and session tokens are rejected for those components; use the ambient credential chain or long-lived access keys with them. For long-running pipelines, prefer the ambient credential chain (leave keys and roles unset), which the driver refreshes automatically. **Type**: `object` ### [](#aws-enabled)`aws.enabled` Enable AWS IAM authentication using the driver-native `MONGODB-AWS` mechanism. The MongoDB Atlas database user must be created with the AWS IAM authentication type, and connections require TLS. When no static credentials or roles are configured, the ambient AWS credential chain (environment variables, EC2 instance profile, EKS pod role) is used and expiring credentials are refreshed automatically. **Type**: `bool` **Default**: `false` ### [](#aws-id)`aws.id` The ID of credentials to use. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region used when assuming roles (for STS calls). Only used when `role` or `roles` are configured; the ambient and static-key paths ignore it. If no region is specified then the environment default is used. **Type**: `string` ### [](#aws-role)`aws.role` Optional AWS IAM role ARN to assume for authentication. Cannot be combined with `roles`; use the `roles` array instead when chaining multiple roles. **Type**: `string` ### [](#aws-role_external_id)`aws.role_external_id` Optional external ID for the role assumption. Only used with the `role` field, which cannot be combined with `roles`. **Type**: `string` ### [](#aws-roles)`aws.roles[]` Optional array of AWS IAM roles to assume for authentication. Roles can be assumed in sequence, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. Cannot be combined with `role`. **Type**: `array` ### [](#aws-roles-role)`aws.roles[].role` AWS IAM role ARN to assume. **Type**: `string` **Default**: `""` ### [](#aws-roles-role_external_id)`aws.roles[].role_external_id` Optional external ID for the role assumption. **Type**: `string` **Default**: `""` ### [](#aws-secret)`aws.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-session_duration)`aws.session_duration` The duration of the STS session requested when assuming roles. AWS requires at least 15 minutes and caps sessions created through role chaining at one hour. Only used when `role` or `roles` are configured. For long-running pipelines, prefer the ambient credential chain over a fixed session duration, since the driver refreshes ambient credentials automatically as they near expiry. **Type**: `string` **Default**: `1h` ### [](#aws-token)`aws.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#collection)`collection` The name of the target collection. **Type**: `string` ### [](#database)`database` The name of the target MongoDB database. **Type**: `string` ### [](#key_field)`key_field` The field in the document that is used as the key. **Type**: `string` ### [](#password)`password` The password to connect to the database. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#url)`url` The URL of the target MongoDB server. **Type**: `string` ```yaml # Examples: url: mongodb://localhost:27017 ``` ### [](#username)`username` The username to connect to the database. **Type**: `string` **Default**: `""` ### [](#value_field)`value_field` The field in the document that is used as the value. **Type**: `string` --- # Page 30: multilevel **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/multilevel.md --- # multilevel > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: multilevel latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/multilevel page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/multilevel.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/multilevel.adoc description: Combines multiple caches as levels, performing read-through and write-through operations across them. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Combines multiple caches as levels, performing read-through and write-through operations across them. ```yml caches: multilevel: - label: "" memory: default_ttl: 5m compaction_interval: 60s - label: "" redis: url: redis://localhost:6379 expiration: 24h ``` ## [](#examples)Examples ### [](#hot-and-cold-cache)Hot and cold cache The multilevel cache is useful for reducing traffic against a remote cache by routing it through a local cache. In the following example requests will only go through to the memcached server if the local memory cache is missing the key. ```yaml pipeline: processors: - branch: processors: - cache: resource: leveled operator: get key: ${! json("key") } - catch: - mapping: 'root = {"err":error()}' result_map: 'root.result = this' cache_resources: - label: leveled multilevel: [ hot, cold ] - label: hot memory: default_ttl: 60s - label: cold memcached: addresses: [ TODO:11211 ] default_ttl: 60s ``` --- # Page 31: nats_kv **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/nats_kv.md --- # nats_kv > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats_kv latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/nats_kv page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/nats_kv.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/nats_kv.adoc description: Cache key/values in a NATS key-value bucket. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Cache key/value pairs in a NATS key-value bucket. #### Common ```yml caches: nats_kv: urls: [] # No default (required) bucket: "" # No default (required) ``` #### Advanced ```yml caches: nats_kv: urls: [] # No default (required) max_reconnects: "" # No default (optional) bucket: "" # No default (required) tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) ``` ## [](#connection-name)Connection name When monitoring and managing a production [NATS system](https://docs.nats.io/nats-concepts/overview), it is often useful to know which connection a message was sent or received from. To achieve this, set the connection name option when creating a NATS connection. Redpanda Connect can then automatically set the connection name to the NATS component label, so that monitoring tools between NATS and Redpanda Connect can stay in sync. ## [](#authentication)Authentication A number of Redpanda Connect components use NATS services. Each of these components support optional, advanced authentication parameters for [NKeys](https://docs.nats.io/nats-server/configuration/securing_nats/auth_intro/nkey_auth) and [user credentials](https://docs.nats.io/using-nats/developer/connecting/creds). For an in-depth guide, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/nats_admin/security/jwt). ### [](#nkeys)NKeys NATS server can use NKeys in several ways for authentication. The simplest approach is to configure the server with a list of user’s public keys. The server can then generate a challenge for each connection request from a client, and the client must respond to the challenge by signing it with its private NKey, configured in the `nkey_file` or `nkey` field. For more details, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth). ### [](#user-credentials)User credentials NATS server also supports decentralized authentication based on JSON Web Tokens (JWTs). When a server is configured to use this authentication scheme, clients need a [user JWT](https://docs.nats.io/nats-server/configuration/securing_nats/jwt#json-web-tokens) and a corresponding [NKey secret](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth) to connect. You can use either of the following methods to supply the user JWT and NKey secret: - In the `user_credentials_file` field, enter the path to a file containing both the private key and the JWT. You can generate the file using the [nsc tool](https://docs.nats.io/nats-tools/nsc). - In the `user_jwt` field, enter a plain text JWT, and in the `user_nkey_seed` field, enter the plain text NKey seed or private key. For more details about authentication using JWTs, see the [NATS documentation](https://docs.nats.io/using-nats/developer/connecting/creds). ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` The NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of an user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plain text user JWT (given along with the corresponding user NKey Seed). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plain text user NKey Seed (given along with the corresponding user JWT). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#bucket)`bucket` The name of the KV bucket. **Type**: `string` ```yaml # Examples: bucket: my_kv_bucket ``` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If an item of the list contains commas it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 32: noop **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/noop.md --- # noop > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: noop latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/noop page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/noop.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/noop.adoc description: Noop is a cache that stores nothing, all gets returns not found. Why? Sometimes doing nothing is the braver option. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Noop is a cache that stores nothing, all gets returns not found. Why? Sometimes doing nothing is the braver option. Introduced in version 4.27.0. ```yml caches: noop: {} ``` --- # Page 33: redis **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/redis.md --- # redis > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/redis page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/redis.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/redis.adoc description: Use a Redis instance as a cache. The expiration can be set to zero or an empty string in order to set no expiration. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Use a Redis instance as a cache. The expiration can be set to zero or an empty string in order to set no expiration. #### Common ```yml caches: redis: url: "" # No default (required) prefix: "" # No default (optional) ``` #### Advanced ```yml caches: redis: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] prefix: "" # No default (optional) default_ttl: "" # No default (optional) retries: initial_interval: 500ms max_interval: 1s max_elapsed_time: 5s ``` ## [](#fields)Fields ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#default_ttl)`default_ttl` An optional default TTL to set for items, calculated from the moment the item is cached. **Type**: `string` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#prefix)`prefix` An optional string to prefix item keys with in order to prevent collisions with similar services. **Type**: `string` ### [](#retries)`retries` Determine time intervals and cut offs for retry attempts. **Type**: `object` ### [](#retries-initial_interval)`retries.initial_interval` The initial period to wait between retry attempts. **Type**: `string` **Default**: `500ms` ```yaml # Examples: initial_interval: 50ms # --- initial_interval: 1s ``` ### [](#retries-max_elapsed_time)`retries.max_elapsed_time` The maximum overall period of time to spend on retry attempts before the request is aborted. **Type**: `string` **Default**: `5s` ```yaml # Examples: max_elapsed_time: 1m # --- max_elapsed_time: 1h ``` ### [](#retries-max_interval)`retries.max_interval` The maximum period to wait between retry attempts **Type**: `string` **Default**: `1s` ```yaml # Examples: max_interval: 5s # --- max_interval: 1m ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 34: redpanda **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/redpanda.md --- # redpanda > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/redpanda page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/redpanda.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/redpanda.adoc description: A Kafka cache using the Franz Kafka client library. page-git-created-date: "2025-07-08" page-git-modified-date: "2026-08-11" --- A Kafka cache implemented using the [Franz Kafka client library](https://github.com/twmb/franz-go). #### Common ```yaml caches: redpanda: seed_brokers: [] # No default (required) topic: "" # No default (required) ``` #### Advanced ```yaml caches: redpanda: seed_brokers: [] # No default (required) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s topic: "" # No default (required) allow_auto_topic_creation: true ``` A cache that stores data in a Kafka topic. This cache is useful for data that is written frequently and queried infrequently. Reads from the cache require scanning the entire topic partition. If you expect frequent access, consider placing an in-memory caching layer in front of this one. Because only the latest values are needed, configure compaction for topics used as caches so that reads are less expensive when topics are rescanned. See [Compaction Settings](https://docs.redpanda.com/streaming/current/manage/cluster-maintenance/compaction-settings/). The cache does not have any TTL mechanisms. Use the Kafka topic retention policies to manage TTL. ## [](#fields)Fields ### [](#allow_auto_topic_creation)`allow_auto_topic_creation` Enables topics to be auto created if they do not exist when fetching their metadata. **Type**: `bool` **Default**: `true` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#conn_idle_timeout)`conn_idle_timeout` The amount of time that connections can remain idle before they are closed. **Type**: `string` **Default**: `20s` ### [](#metadata_max_age)`metadata_max_age` The maximum age of metadata before it is refreshed. This interval also controls how frequently regex topic patterns are re-evaluated to discover new matching topics. **Type**: `string` **Default**: `1m` ### [](#request_timeout_overhead)`request_timeout_overhead` Additional time to apply as overhead when calculating request deadlines. This buffer helps prevent premature timeouts, especially for requests that already define their own timeout values. **Type**: `string` **Default**: `10s` ### [](#sasl)`sasl[]` Specify one or more SASL authentication methods. Each method is tried in the order specified. If the broker supports the first mechanism, outgoing client connections use that mechanism. If the first mechanism fails, the client will use the first supported mechanism. If the broker does not support any client mechanisms, connections will fail. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS-specific fields for when [`sasl.mechanism`](#sasl-mechanism) is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. For more information, see the [credentials for AWS](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/) guide. **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` The credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` The ARN of the role to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming the specified role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used. Required only when using short-term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` A custom endpoint URL for AWS API requests. Use this to connect to AWS-compatible services or local testing environments instead of the standard AWS endpoints. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use for authentication. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM-based authentication as specified by the aws-msk-iam-auth Java library. | | OAUTHBEARER | OAuth Bearer authentication. | | PLAIN | PLAIN mechanism for plaintext password authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM authentication as specified in RFC5802. | | none | Disable SASL authentication. | ### [](#sasl-password)`sasl[].password` The password to use for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` The username to use for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to. Items containing commas are expanded into multiple addresses. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` The plaintext certificate to use for TLS authentication. Must be paired with the corresponding private key in the `key` field when using inline PEM data for mTLS client certificates. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path to a file containing the certificate to use for TLS authentication. Must be paired with the corresponding private key file in the `key_file` field when using file-based configuration for mTLS client certificates. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` Private key for mTLS client certificate as inline PEM data. Must correspond to the client certificate specified in the `cert` field. Use this field together with `cert` when providing certificate data inline rather than through files. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` Path to private key file for mTLS client certificate in PEM format. Must correspond to the client certificate specified in the `cert_file` field. Use this field together with `cert_file` when loading certificate data from files. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` The password to use for the private key (specified in the `key` or `key_file` fields), if it is password-protected. The PKCS#1 and PKCS#8 formats are supported. Supports environment variable interpolation for secure password management. The `pbeWithMD5AndDES-CBC` algorithm is obsolete and not supported for the PKCS#8 format. This algorithm does not authenticate the ciphertext, making it vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic)`topic` The topic to store data in. **Type**: `string` --- # Page 35: ristretto **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/ristretto.md --- # ristretto > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: ristretto latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/ristretto page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/ristretto.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/ristretto.adoc description: Stores key/value pairs in a map held in the memory-bound Ristretto cache. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores key/value pairs in a map held in the memory-bound [Ristretto cache](https://github.com/dgraph-io/ristretto). #### Common ```yml caches: ristretto: default_ttl: "" ``` #### Advanced ```yml caches: ristretto: default_ttl: "" get_retries: enabled: false initial_interval: 1s max_interval: 5s max_elapsed_time: 30s ``` This cache is more efficient and appropriate for high-volume use cases than the standard memory cache. However, the add command is non-atomic, and therefore this cache is not suitable for deduplication. ## [](#fields)Fields ### [](#default_ttl)`default_ttl` A default TTL to set for items, calculated from the moment the item is cached. Set to an empty string or zero duration to disable TTLs. **Type**: `string` **Default**: `""` ```yaml # Examples: default_ttl: 5m # --- default_ttl: 60s ``` ### [](#get_retries)`get_retries` Determines how and whether get attempts should be retried if the key is not found. Ristretto is a concurrent cache that does not immediately reflect writes, and so it can sometimes be useful to enable retries at the cost of speed in cases where the key is expected to exist. **Type**: `object` ### [](#get_retries-enabled)`get_retries.enabled` Whether retries should be enabled. **Type**: `bool` **Default**: `false` ### [](#get_retries-initial_interval)`get_retries.initial_interval` The initial period to wait between retry attempts. **Type**: `string` **Default**: `1s` ```yaml # Examples: initial_interval: 50ms # --- initial_interval: 1s ``` ### [](#get_retries-max_elapsed_time)`get_retries.max_elapsed_time` The maximum overall period of time to spend on retry attempts before the request is aborted. **Type**: `string` **Default**: `30s` ```yaml # Examples: max_elapsed_time: 1m # --- max_elapsed_time: 1h ``` ### [](#get_retries-max_interval)`get_retries.max_interval` The maximum period to wait between retry attempts **Type**: `string` **Default**: `5s` ```yaml # Examples: max_interval: 5s # --- max_interval: 1m ``` --- # Page 36: sql **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/sql.md --- # sql > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/sql page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/sql.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/sql.adoc description: Uses an SQL database table as a destination for storing cache key/value items. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- Uses an SQL database table as a destination for storing cache key/value items. #### Common ```yml caches: sql: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) key_column: "" # No default (required) value_column: "" # No default (required) set_suffix: "" # No default (optional) ``` #### Advanced ```yml caches: sql: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) key_column: "" # No default (required) value_column: "" # No default (required) set_suffix: "" # No default (optional) init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) ``` Each cache key/value pair will exist as a row within the specified table. Currently only the key and value columns are set, and therefore any other columns present within the target table must allow NULL values if this cache is going to be used for set and add operations. Cache operations are translated into SQL statements as follows: ## [](#get)Get All `get` operations are performed with a traditional `select` statement. ## [](#delete)Delete All `delete` operations are performed with a traditional `delete` statement. ## [](#set)Set The `set` operation is performed with a traditional `insert` statement. This will behave as an `add` operation by default, and so ideally needs to be adapted in order to provide updates instead of failing on collision s. Since different SQL engines implement upserts differently it is necessary to specify a `set_suffix` that modifies an `insert` statement in order to perform updates on conflict. ## [](#add)Add The `add` operation is performed with a traditional `insert` statement. ## [](#fields)Fields ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#key_column)`key_column` The name of a column to be used for storing cache item keys. This column should support strings of arbitrary size. **Type**: `string` ```yaml # Examples: key_column: foo ``` ### [](#set_suffix)`set_suffix` An optional suffix to append to each insert query for a cache `set` operation. This should modify an insert statement into an upsert appropriate for the given SQL engine. **Type**: `string` ```yaml # Examples: set_suffix: ON DUPLICATE KEY UPDATE bar=VALUES(bar) # --- set_suffix: ON CONFLICT (foo) DO UPDATE SET bar=excluded.bar # --- set_suffix: ON CONFLICT (foo) DO NOTHING ``` ### [](#table)`table` The table to insert/read/delete cache items. **Type**: `string` ```yaml # Examples: table: foo ``` ### [](#value_column)`value_column` The name of a column to be used for storing cache item values. This column should support strings of arbitrary size. **Type**: `string` ```yaml # Examples: value_column: bar ``` --- # Page 37: ttlru **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/ttlru.md --- # ttlru > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: ttlru latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/caches/ttlru page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/caches/ttlru.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/caches/ttlru.adoc description: Stores key/value pairs in a ttlru in-memory cache. This cache is therefore reset every time the service restarts. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores key/value pairs in a ttlru in-memory cache. This cache is therefore reset every time the service restarts. #### Common ```yml caches: ttlru: cap: 1024 default_ttl: 5m0s init_values: {} ``` #### Advanced ```yml caches: ttlru: cap: 1024 default_ttl: 5m0s ttl: "" # No default (optional) init_values: {} optimistic: false ``` The cache ttlru provides a simple, goroutine safe, cache with a fixed number of entries. Each entry has a per-cache defined TTL. This TTL is reset on both modification and access of the value. As a result, if the cache is full, and no items have expired, when adding a new item, the item with the soonest expiration will be evicted. It uses the package [`expirable`](https://github.com/hashicorp/golang-lru/tree/main/expirable) The field init\_values can be used to pre-populate the memory cache with any number of key/value pairs: ```yaml cache_resources: - label: foocache ttlru: default_ttl: '5m' cap: 1024 init_values: foo: bar ``` These values can be overridden during execution. ## [](#fields)Fields ### [](#cap)`cap` The cache maximum capacity (number of entries) **Type**: `int` **Default**: `1024` ### [](#default_ttl)`default_ttl` The cache ttl of each element **Type**: `string` **Default**: `5m0s` ### [](#init_values)`init_values` A table of key/value pairs that should be present in the cache on initialization. This can be used to create static lookup tables. **Type**: `object` **Default**: `{}` ```yaml # Examples: init_values: Nickelback: "1995" Spice Girls: "1994" The Human League: "1977" ``` ### [](#optimistic)`optimistic` If true, we do not lock on read/write events. The ttlru package is thread-safe, however the ADD operation is not atomic. **Type**: `bool` **Default**: `false` ### [](#ttl)`ttl` Deprecated. Please use `default_ttl` field **Type**: `string` --- # Page 38: Inputs **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/about.md --- # Inputs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Inputs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- An input is a source of data piped through an array of optional [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/): ```yaml input: label: my_redis_input redis_streams: url: tcp://localhost:6379 streams: - benthos_stream body_key: body consumer_group: benthos_group # Optional list of processing steps processors: - mapping: | root.document = this.without("links") root.link_count = this.links.length() ``` Some inputs have a logical end, when this happens the input gracefully terminates and Redpanda Connect will shut itself down once all messages have been processed fully. It’s also possible to specify a logical end for an input that otherwise doesn’t have one with the [`read_until` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/read_until/), which checks a condition against each consumed message in order to determine whether it should be the last. ## [](#brokering)Brokering Only one input is configured at the root of a Redpanda Connect config. However, the root input can be a [broker](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/broker/) which combines multiple inputs and merges the streams: ```yaml input: broker: inputs: - kafka: addresses: [ TODO ] topics: [ foo, bar ] consumer_group: foogroup - redis_streams: url: tcp://localhost:6379 streams: - benthos_stream body_key: body consumer_group: benthos_group ``` ## [](#labels)Labels Inputs have an optional field `label` that can uniquely identify them in observability data such as metrics and logs. This can be useful when running configs with multiple inputs, otherwise their metrics labels will be generated based on their composition. For more information check out the [metrics documentation](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/). ### [](#sequential-reads)Sequential reads Sometimes it’s useful to consume a sequence of inputs, where an input is only consumed once its predecessor is drained fully, you can achieve this with the [`sequence` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/). ## [](#generating-messages)Generating messages It’s possible to generate data with Redpanda Connect using the [`generate` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/generate/), which is also a convenient way to trigger scheduled pipelines. --- # Page 39: amqp_0_9 **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/amqp_0_9.md --- # amqp_0_9 > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: amqp_0_9 latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/amqp_0_9 page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/amqp_0_9.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/amqp_0_9.adoc description: Connects to an AMQP (0.91) queue. AMQP is a messaging protocol used by various message brokers, including RabbitMQ. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Connects to an AMQP (0.91) queue. AMQP is a messaging protocol used by various message brokers, including RabbitMQ. #### Common ```yml inputs: label: "" amqp_0_9: urls: [] # No default (required) queue: "" # No default (required) consumer_tag: "" prefetch_count: 10 ``` #### Advanced ```yml inputs: label: "" amqp_0_9: urls: [] # No default (required) queue: "" # No default (required) queue_declare: enabled: false durable: true auto_delete: false arguments: "" # No default (optional) bindings_declare: [] # No default (optional) consumer_tag: "" auto_ack: false nack_reject_patterns: [] prefetch_count: 10 prefetch_size: 0 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] ``` TLS is automatically enabled when connecting to an `amqps` URL. However, you can customize [TLS settings](#tls) if required. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `amqp_content_type` - `amqp_content_encoding` - `amqp_delivery_mode` - `amqp_priority` - `amqp_correlation_id` - `amqp_reply_to` - `amqp_expiration` - `amqp_message_id` - `amqp_timestamp` - `amqp_type` - `amqp_user_id` - `amqp_app_id` - `amqp_consumer_tag` - `amqp_delivery_tag` - `amqp_redelivered` - `amqp_exchange` - `amqp_routing_key` - All existing message headers, including nested headers prefixed with the key of their respective parent. You can access these metadata fields using [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#auto_ack)`auto_ack` Set to `true` to automatically acknowledge messages as soon as they are consumed rather than waiting for acknowledgments from downstream. This can improve throughput and prevent the pipeline from becoming blocked, but delivery guarantees are lost. **Type**: `bool` **Default**: `false` ### [](#bindings_declare)`bindings_declare[]` Passively declares the bindings of the target queue to make sure they exist and are configured correctly. If the bindings exist, then the passive declaration verifies that fields specified in this object match them. **Type**: `array` ```yaml # Examples: bindings_declare: - exchange: foo key: bar ``` ### [](#bindings_declare-exchange)`bindings_declare[].exchange` The exchange of the declared binding. **Type**: `string` **Default**: `""` ### [](#bindings_declare-key)`bindings_declare[].key` The key of the declared binding. **Type**: `string` **Default**: `""` ### [](#consumer_tag)`consumer_tag` A consumer tag to uniquely identify the consumer. **Type**: `string` **Default**: `""` ### [](#nack_reject_patterns)`nack_reject_patterns[]` A list of regular expression patterns to match against errors in messages that Redpanda Connect fails to deliver. When a message has an error that matches a pattern, it is dropped or delivered to a dead-letter queue (if a queue has been configured). By default, failed messages are negatively acknowledged (nacked) and requeued. **Type**: `array` **Default**: `[]` ```yaml # Examples: nack_reject_patterns: - "^reject me please:.+$" ``` ### [](#prefetch_count)`prefetch_count` The maximum number of pending messages at a given time. **Type**: `int` **Default**: `10` ### [](#prefetch_size)`prefetch_size` The maximum size of pending messages (in bytes) at a given time. **Type**: `int` **Default**: `0` ### [](#queue)`queue` An AMQP queue to consume from. **Type**: `string` ### [](#queue_declare)`queue_declare` Passively declares the [target queue](#queue) to make sure a queue with the specified name exists and is configured correctly. If the queue exists, then the passive declaration verifies that fields specified in this object match the its properties. **Type**: `object` ### [](#queue_declare-arguments)`queue_declare.arguments` Arguments for server-specific implementations of the queue (optional). You can use arguments to configure additional parameters for queue types that require them. For more information about available arguments, see the [RabbitMQ Client Library](https://github.com/rabbitmq/amqp091-go/blob/b3d409fe92c34bea04d8123a136384c85e8dc431/types.go#L282-L362). | Argument | Description | Accepted values | | --- | --- | --- | | x-queue-type | Declares the type of queue. | Options: classic (default), quorum, stream, drop-head, reject-publish, and reject-publish-dlx. | | x-max-length | The maximum number of messages in the queue. | A non-negative integer. | | x-max-length-bytes | The maximum size of messages (in bytes) in the queue. | A non-negative integer. | | x-overflow | Sets the queue’s overflow behavior. | Options: drop-head (default), reject-publish, reject-publish-dlx. | | x-message-ttl | The duration (in milliseconds) that messages remain in the queue before they expire and are discarded. | A string that represents the number of milliseconds. For example, 60000 retains messages for one minute. | | x-expires | The duration after which the queue automatically expires. | A positive integer. | | x-max-age | The duration (in configurable units) that streamed messages are retained on disk before they are discarded. | Options: Y, M, D, h, m, s. For example, 7D retains messages for a week. | | x-stream-max-segment-size-bytes | The maximum size (in bytes) of the segment files held on disk. | A positive integer. Default: 500000000 (approximately 500 MB). | | x-queue-version | The version of the classic queue to use. | Options: 1 or 2. | | x-consumer-timeout | The duration (in milliseconds) that a consumer can remain idle before it is automatically canceled. | A positive integer that represents the number of milliseconds. For example, 60000 sets a timeout duration of one minute. | | x-single-active-consumer | When set to true, a single consumer receives messages from the queue even when multiple consumers are subscribed to it. | A boolean. | **Type**: `object` ```yaml # Examples: arguments: x-max-length: 1000 x-max-length-bytes: 4096 x-queue-type: quorum ``` ### [](#queue_declare-auto_delete)`queue_declare.auto_delete` Whether the declared queue auto-deletes when there are no active consumers. **Type**: `bool` **Default**: `false` ### [](#queue_declare-durable)`queue_declare.durable` Whether the declared queue is durable. **Type**: `bool` **Default**: `true` ### [](#queue_declare-enabled)`queue_declare.enabled` Whether to enable queue declaration. **Type**: `bool` **Default**: `false` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. This input attempts to connect to each URL in the list, in order, until a successful connection is established. It then continues to use that URL until the connection is closed. If an item in the list contains commas, it is split into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "amqp://guest:guest@127.0.0.1:5672/" # --- urls: - "amqp://127.0.0.1:5672/,amqp://127.0.0.2:5672/" # --- urls: - "amqp://127.0.0.1:5672/" - "amqp://127.0.0.2:5672/" ``` --- # Page 40: aws_cloudwatch_logs **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_cloudwatch_logs.md --- # aws_cloudwatch_logs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_cloudwatch_logs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/aws_cloudwatch_logs page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/aws_cloudwatch_logs.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/aws_cloudwatch_logs.adoc description: Consumes log events from AWS CloudWatch Logs. page-git-created-date: "2026-03-13" page-git-modified-date: "2026-08-11" --- Consumes log events from AWS CloudWatch Logs. #### Common ```yml inputs: label: "" aws_cloudwatch_logs: log_group_name: "" # No default (required) log_stream_names: [] # No default (optional) log_stream_prefix: "" # No default (optional) filter_pattern: "" # No default (optional) start_time: "" # No default (optional) poll_interval: 5s auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" aws_cloudwatch_logs: log_group_name: "" # No default (required) log_stream_names: [] # No default (optional) log_stream_prefix: "" # No default (optional) filter_pattern: "" # No default (optional) start_time: "" # No default (optional) poll_interval: 5s limit: 1000 structured_log: true api_timeout: 30s auto_replay_nacks: true region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` Polls CloudWatch Log Groups for log events. Supports filtering by log streams, CloudWatch filter patterns, and configurable start times. Each log event becomes a separate message with metadata including the log group name, log stream name, timestamp, and ingestion time. > ❗ **IMPORTANT** > > This input provides at-least-once delivery. It tracks its position in memory only, so if the process restarts, it resumes from the configured `start_time` (or the beginning if not set). Duplicates can occur across restarts. For exactly-once outcomes, implement idempotent or deduplicated downstream processing. ## [](#credentials)Credentials By default, Redpanda Connect uses a shared credentials file when connecting to AWS services. You can also set credentials explicitly at the component level to transfer data across accounts. You can find out more in [AWS credentials](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `cloudwatch_log_group`: The name of the log group. - `cloudwatch_log_stream`: The name of the log stream. - `cloudwatch_timestamp`: The timestamp of the log event (Unix milliseconds). - `cloudwatch_ingestion_time`: The ingestion timestamp (Unix milliseconds). - `cloudwatch_event_id`: The unique event ID. You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#api_timeout)`api_timeout` The maximum time to wait for an API request to complete. **Type**: `string` **Default**: `30s` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#filter_pattern)`filter_pattern` An optional CloudWatch Logs filter pattern to apply when querying log events. For syntax details, see the [CloudWatch Logs filter and pattern syntax](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/FilterAndPatternSyntax.html) documentation. **Type**: `string` ```yaml # Examples: filter_pattern: [ERROR] ``` ### [](#limit)`limit` The maximum number of log events to return in a single API call. Valid range: 1-10000. **Type**: `int` **Default**: `1000` ### [](#log_group_name)`log_group_name` The name of the CloudWatch Log Group to consume from. **Type**: `string` ```yaml # Examples: log_group_name: my-app-logs ``` ### [](#log_stream_names)`log_stream_names[]` An optional list of log stream names to consume from. If not set, events from all streams in the log group will be consumed. **Type**: `array` ```yaml # Examples: log_stream_names: - stream-1 - stream-2 ``` ### [](#log_stream_prefix)`log_stream_prefix` An optional log stream name prefix to filter streams. Only streams starting with this prefix will be consumed. **Type**: `string` ```yaml # Examples: log_stream_prefix: prod- ``` ### [](#poll_interval)`poll_interval` The interval at which to poll for new log events. **Type**: `string` **Default**: `5s` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#start_time)`start_time` The time to start consuming log events from. Can be an RFC3339 timestamp (for example, `2024-01-01T00:00:00Z`) or the string `now` to start consuming from the current time. If not set, starts from the beginning of available logs. **Type**: `string` ```yaml # Examples: start_time: 2024-01-01T00:00:00Z # --- start_time: now ``` ### [](#structured_log)`structured_log` Whether to output log events as structured JSON objects with all metadata fields, or as plain text messages with metadata stored in Redpanda Connect message metadata. **Type**: `bool` **Default**: `true` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` --- # Page 41: aws_dynamodb_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_dynamodb_cdc.md --- # aws_dynamodb_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_dynamodb_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/aws_dynamodb_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/aws_dynamodb_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/aws_dynamodb_cdc.adoc description: Reads change data capture (CDC) events from DynamoDB Streams. page-topic-type: reference personas: data_engineer, streaming_developer, platform_operator learning-objective-1: Look up configuration options for DynamoDB CDC streaming learning-objective-2: Find metadata fields available for message processing learning-objective-3: Identify checkpointing and performance tuning settings page-git-created-date: "2026-03-04" page-git-modified-date: "2026-08-11" --- Stream item-level changes from DynamoDB tables using DynamoDB Streams. This input automatically manages shards, checkpoints progress for recovery, and processes multiple shards concurrently. Use this reference to: - Look up configuration options for DynamoDB CDC streaming - Find metadata fields available for message processing - Identify checkpointing and performance tuning settings ### Common ```yml inputs: label: "" aws_dynamodb_cdc: tables: [] checkpoint_table: redpanda_dynamodb_checkpoints checkpoint_namespace: "" start_from: trim_horizon auto_replay_nacks: true snapshot_mode: none ``` ### Advanced ```yml inputs: label: "" aws_dynamodb_cdc: tables: [] table_discovery_mode: single table_tag_filter: "" table_discovery_interval: 5m checkpoint_table: redpanda_dynamodb_checkpoints checkpoint_namespace: "" global_table: false global_table_replicas: [] batch_size: 1000 poll_interval: 1s start_from: trim_horizon checkpoint_limit: 1000 max_tracked_shards: 10000 auto_replay_nacks: true throttle_backoff: 100ms snapshot_mode: none snapshot_segments: 1 snapshot_batch_size: 100 snapshot_throttle: 100ms snapshot_deduplicate: true snapshot_buffer_size: 100000 region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` ## [](#prerequisites)Prerequisites The source DynamoDB table must have [DynamoDB Streams](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html) enabled. You can enable streams with one of these view types: - `KEYS_ONLY`: Only the key attributes of the modified item - `NEW_IMAGE`: The entire item as it appears after the modification - `OLD_IMAGE`: The entire item as it appeared before the modification - `NEW_AND_OLD_IMAGES`: Both the new and old item images ## [](#checkpointing)Checkpointing Checkpoints are stored in a separate DynamoDB table (configured via `checkpoint_table`). This table is created automatically if it does not exist. On restart, the input resumes from the last checkpointed position for each shard. ### [](#share-a-checkpoint-table-across-pipelines)Share a checkpoint table across pipelines By default, checkpoints are keyed by stream and shard only, so multiple pipelines that read the same stream and share one `checkpoint_table` overwrite each other’s positions, which causes skipped or duplicated events. This is common when several developers or environments test against the same table. Set `checkpoint_namespace` to give each pipeline its own isolated set of checkpoints within a shared table. Redpanda Connect prefixes the namespace to the checkpoint key, so pipelines with different namespaces never collide: ```yaml # Alice's pipeline input: aws_dynamodb_cdc: tables: [ orders ] region: us-east-1 checkpoint_table: shared_dynamodb_checkpoints checkpoint_namespace: dev-alice start_from: trim_horizon ``` ```yaml # Bob's pipeline: same stream and checkpoint table, isolated by namespace input: aws_dynamodb_cdc: tables: [ orders ] region: us-east-1 checkpoint_table: shared_dynamodb_checkpoints checkpoint_namespace: dev-bob start_from: trim_horizon ``` The `checkpoint_namespace` field requires Redpanda Connect 4.101.0 or later. It is backward compatible: leaving it unset (the default) keeps the existing checkpoint keys unchanged, so existing deployments are unaffected. The value cannot contain a `#` character. > 📝 **NOTE** > > A namespace provides isolation, not coordination. Two pipelines that use the same `checkpoint_namespace` and `checkpoint_table` still overwrite each other’s checkpoints. To run readers independently, give each one a distinct namespace. > ⚠️ **WARNING** > > Changing or removing `checkpoint_namespace` makes the pipeline read checkpoints under the new key. If no checkpoints exist yet for that namespace, for example the first time you set it or when you switch to a namespace that has never been used, the pipeline has nothing to resume from and starts at `start_from` (with `start_from: trim_horizon`, this replays all records still available in the stream’s 24-hour retention window). Switching back to a previously used namespace resumes from that namespace’s last checkpoints. ## [](#alternative-components)Alternative components For better performance and longer retention (up to 1 year vs 24 hours), consider using Kinesis Data Streams for DynamoDB with the `aws_kinesis` input instead. ## [](#message-structure)Message structure Each CDC event is delivered as a JSON message with the following structure. Use these fields in your Bloblang mappings with `this.`: ```json { "eventID": "abc123-", (1) "eventName": "INSERT | MODIFY | REMOVE", (2) "eventSource": "aws:dynamodb", "awsRegion": "us-east-1", "tableName": "my-table", (3) "dynamodb": { "keys": { (4) "pk": "user#123", "sk": "profile" }, "newImage": { (5) "pk": "user#123", "sk": "profile", "name": "Alice", "email": "alice@example.com" }, "oldImage": { (6) "pk": "user#123", "sk": "profile", "name": "Alice Smith" }, "sequenceNumber": "12345678901234567890", (7) "sizeBytes": 256, "streamViewType": "NEW_AND_OLD_IMAGES" } } ``` | 1 | Unique identifier for this change event. | | --- | --- | | 2 | Type of change: INSERT (new item), MODIFY (updated item), or REMOVE (deleted item). | | 3 | Name of the source DynamoDB table. | | 4 | Primary key attributes of the changed item. Always present. | | 5 | Item state after the change. Present for INSERT and MODIFY events (requires NEW_IMAGE or NEW_AND_OLD_IMAGES stream view type). | | 6 | Item state before the change. Present for MODIFY and REMOVE events (requires OLD_IMAGE or NEW_AND_OLD_IMAGES stream view type). | | 7 | Position of this record in the shard, used for ordering and checkpointing. | > 📝 **NOTE** > > DynamoDB attribute values are automatically unmarshalled from DynamoDB’s type format (`{"S": "value"}`) to plain values (`"value"`). ### [](#example-mapping)Example mapping ```yaml pipeline: processors: - mapping: | root.event_type = this.eventName root.table = this.tableName root.keys = this.dynamodb.keys root.new_data = this.dynamodb.newImage root.old_data = this.dynamodb.oldImage ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `dynamodb_shard_id`: The shard ID from which the record was read - `dynamodb_sequence_number`: The sequence number of the record in the stream - `dynamodb_event_name`: The type of change: INSERT, MODIFY, or REMOVE - `dynamodb_table`: The name of the DynamoDB table ## [](#metrics)Metrics This input emits the following metrics: - `dynamodb_cdc_shards_tracked`: Total number of shards being tracked (gauge) - `dynamodb_cdc_shards_active`: Number of shards currently being read from (gauge) ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batch_size)`batch_size` Maximum number of records to read per shard in a single request. Valid range: 1-1000. **Type**: `int` **Default**: `1000` ### [](#checkpoint_limit)`checkpoint_limit` Maximum number of unacknowledged messages before forcing a checkpoint update. Lower values provide better recovery guarantees but increase write overhead. **Type**: `int` **Default**: `1000` ### [](#checkpoint_namespace)`checkpoint_namespace` Isolates this pipeline’s checkpoints within a shared `checkpoint_table` by prefixing the namespace to the checkpoint key. Use this so that multiple pipelines reading the same stream can share one checkpoint table without overwriting each other’s positions, for example per-developer or per-environment test pipelines. Leave empty (the default) to keep the original checkpoint keys unchanged. A namespace isolates readers but does not coordinate them: pipelines that share the same namespace still collide. Changing or removing the namespace changes the checkpoint key. If no checkpoints exist yet under the new key, the pipeline starts from `start_from`. Switching back to a previously used namespace resumes from that namespace’s last checkpoints. The value cannot contain a `#` character. **Type**: `string` **Default**: `""` ### [](#checkpoint_table)`checkpoint_table` DynamoDB table name for storing checkpoints. Will be created if it doesn’t exist. **Type**: `string` **Default**: `redpanda_dynamodb_checkpoints` ### [](#credentials)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#global_table)`global_table` Provision the checkpoint table as a DynamoDB Global Table (v2) so checkpoints replicate across regions. Requires `global_table_replicas`. When the table is auto-created it is created as a global table; when it already exists, its replicas are reconciled (missing regions are added by calling `UpdateTable`). The existing table must have been created in global mode (`TableId` hash key). Enabling this against a pre-existing non-global checkpoint table fails fast with a clear error. **Type**: `bool` **Default**: `false` ### [](#global_table_replicas)`global_table_replicas[]` Regions other than this pipeline’s own region to replicate the checkpoint table to. The pipeline’s own region is always included. Required when `global_table` is true. Applied both when the checkpoint table is created and, for an existing global table, when reconciling replicas (missing regions are added; this list is not used to remove regions). **Type**: `array` **Default**: `[]` ### [](#max_tracked_shards)`max_tracked_shards` Maximum number of shards to track simultaneously. Prevents memory issues with extremely large tables. **Type**: `int` **Default**: `10000` ### [](#poll_interval)`poll_interval` Time to wait between polling attempts when no records are available. **Type**: `string` **Default**: `1s` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#snapshot_batch_size)`snapshot_batch_size` Records per scan request during snapshot. Maximum 1000. Lower values provide better backpressure control but require more API calls. **Type**: `int` **Default**: `100` ### [](#snapshot_buffer_size)`snapshot_buffer_size` Maximum CDC events to buffer for deduplication (approximately 100 bytes per entry). If exceeded, deduplication is disabled and duplicates may be emitted. **Type**: `int` **Default**: `100000` ### [](#snapshot_deduplicate)`snapshot_deduplicate` Deduplicate records that appear in both snapshot and CDC stream. Requires buffering CDC events during snapshot. If buffer is exceeded, deduplication is disabled to prevent data loss. **Type**: `bool` **Default**: `true` ### [](#snapshot_mode)`snapshot_mode` `none`: Streams CDC events only (default). `snapshot_only`: Performs a one-time full table scan with no ongoing streaming. `snapshot_and_cdc`: Scans the entire table, then streams changes. **Type**: `string` **Default**: `none` **Options**: `none`, `snapshot_only`, `snapshot_and_cdc` ### [](#snapshot_segments)`snapshot_segments` Number of parallel scan segments (1-10). Higher parallelism scans faster but consumes more Read Capacity Units (RCUs). A lower value is safer to start with. **Type**: `int` **Default**: `1` ### [](#snapshot_throttle)`snapshot_throttle` Minimum time between scan requests per segment. Use this to limit Read Capacity Unit (RCU) consumption during snapshot. **Type**: `string` **Default**: `100ms` ### [](#start_from)`start_from` Where to start reading on a genuinely fresh pipeline (no checkpoint state exists yet under this `checkpoint_namespace` for the stream). `trim_horizon` starts from the oldest available record, `latest` starts from new records. `latest` is honoured only on that first discovery: once any checkpoint state exists, shards discovered later - rotation children found by the periodic refresh, and any checkpoint-less shard after a restart - always start at `trim_horizon` so their backlog is never skipped. In practice a restart under `latest` therefore replays from each shard’s oldest retained record rather than only new records; at-least-once delivery takes precedence over the configured start position. **Type**: `string` **Default**: `trim_horizon` **Options**: `trim_horizon`, `latest` ### [](#table_discovery_interval)`table_discovery_interval` Interval for rescanning and discovering new tables when using `tag` or `includelist` mode. Set to 0 to disable periodic rescanning. **Type**: `string` **Default**: `5m` ### [](#table_discovery_mode)`table_discovery_mode` `single`: Streams from tables specified in the `tables` list. `tag`: Auto-discovers tables by tags (ignores the `tables` field). `includelist`: Streams from tables in the `tables` list. Use `single` instead; `includelist` is kept for backward compatibility. **Type**: `string` **Default**: `single` **Options**: `single`, `tag`, `includelist` ### [](#table_tag_filter)`table_tag_filter` Multi-tag filter in the format `key1:v1,v2;key2:v3,v4`. Matches tables where (key1=v1 OR key1=v2) AND (key2=v3 OR key2=v4). Required when `table_discovery_mode` is `tag`. **Type**: `string` **Default**: `""` ### [](#tables)`tables[]` List of table names to stream from. For single table mode, provide one table. For multi-table mode, provide multiple tables. **Type**: `array` **Default**: `[]` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#throttle_backoff)`throttle_backoff` Time to wait when applying backpressure due to too many in-flight messages. **Type**: `string` **Default**: `100ms` ## [](#examples)Examples ### [](#consume-cdc-events)Consume CDC events Read change events from a DynamoDB table with streams enabled. ```yaml input: aws_dynamodb_cdc: tables: [my-table] region: us-east-1 ``` ### [](#start-from-latest)Start from latest Only process new changes, ignoring existing stream data. ```yaml input: aws_dynamodb_cdc: tables: [orders] start_from: latest region: us-west-2 ``` ### [](#snapshot-and-cdc)Snapshot and CDC Scan all existing records, then stream ongoing changes. ```yaml input: aws_dynamodb_cdc: tables: [products] snapshot_mode: snapshot_and_cdc snapshot_segments: 5 region: us-east-1 ``` ### [](#auto-discover-tables-by-tag)Auto-discover tables by tag Automatically discover and stream from all tables with a specific tag. ```yaml input: aws_dynamodb_cdc: table_discovery_mode: tag table_tag_filter: "stream-enabled:true" table_discovery_interval: 5m region: us-east-1 ``` ### [](#auto-discover-tables-by-multiple-tags)Auto-discover tables by multiple tags Discover tables matching multiple tag criteria with OR logic per key, AND logic across keys. ```yaml input: aws_dynamodb_cdc: table_discovery_mode: tag table_tag_filter: "environment:prod,staging;team:data,analytics" table_discovery_interval: 5m region: us-east-1 # Matches tables with: (environment=prod OR environment=staging) AND (team=data OR team=analytics) ``` ### [](#stream-from-multiple-specific-tables)Stream from multiple specific tables Stream from an explicit list of tables simultaneously. ```yaml input: aws_dynamodb_cdc: table_discovery_mode: includelist tables: - orders - customers - products region: us-west-2 ``` ## [](#suggested-reading)Suggested reading For common patterns including filtering events, routing to Kafka or S3, and detecting changed fields, see the [DynamoDB CDC Patterns](https://docs.redpanda.com/connect/cookbooks/dynamodb_cdc/) cookbook. --- # Page 42: aws_kinesis **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_kinesis.md --- # aws_kinesis > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_kinesis latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/aws_kinesis page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/aws_kinesis.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/aws_kinesis.adoc description: Receive messages from one or more Kinesis streams. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Receive messages from one or more Kinesis streams. #### Common ```yml inputs: label: "" aws_kinesis: streams: [] # No default (required) dynamodb: table: "" create: false billing_mode: PAY_PER_REQUEST read_capacity_units: 0 write_capacity_units: 0 region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) checkpoint_limit: 1024 auto_replay_nacks: true commit_period: 5s steal_grace_period: 2s start_from_oldest: true batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml inputs: label: "" aws_kinesis: streams: [] # No default (required) dynamodb: table: "" create: false billing_mode: PAY_PER_REQUEST read_capacity_units: 0 write_capacity_units: 0 region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) checkpoint_limit: 1024 auto_replay_nacks: true commit_period: 5s steal_grace_period: 2s rebalance_period: 30s lease_period: 30s start_from_oldest: true region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` Consumes messages from one or more Kinesis streams either by automatically balancing shards across other instances of this input, or by consuming shards listed explicitly. The latest message sequence consumed by this input is stored within a [DynamoDB table](#table-schema), which allows it to resume at the correct sequence of the shard during restarts. This table is also used for coordination across distributed inputs when shard balancing. Redpanda Connect will not store a consumed sequence unless it is acknowledged at the output level, which ensures at-least-once delivery guarantees. ## [](#ordering)Ordering By default messages of a shard can be processed in parallel, up to a limit determined by the field `checkpoint_limit`. However, if strict ordered processing is required then this value must be set to 1 in order to process shard messages in lock-step. When doing so it is recommended that you perform batching at this component for performance as it will not be possible to batch lock-stepped messages at the output level. ## [](#table-schema)Table schema It’s possible to configure Redpanda Connect to create the DynamoDB table required for coordination if it does not already exist. However, if you wish to create this yourself (recommended) then create a table with a string HASH key `StreamID` and a string RANGE key `ShardID`. ## [](#batching)Batching Use the `batching` fields to configure an optional [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/#batch-policy). Each stream shard will be batched separately in order to ensure that acknowledgements aren’t contaminated. ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batching-2)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_limit)`checkpoint_limit` The maximum gap between the in flight sequence versus the latest acknowledged sequence at a given time. Increasing this limit enables parallel processing and batching at the output level to work on individual shards. Any given sequence will not be committed unless all messages under that offset are delivered in order to preserve at least once delivery guarantees. **Type**: `int` **Default**: `1024` ### [](#commit_period)`commit_period` The period of time between each update to the checkpoint table. **Type**: `string` **Default**: `5s` ### [](#credentials)`credentials` Manually configure the AWS credentials to use (optional). For more information, see the [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of the AWS credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` The profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` The role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to use when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the AWS credentials in use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the AWS credentials in use. This is a required value for short-term credentials. **Type**: `string` ### [](#dynamodb)`dynamodb` Determines the table used for storing and accessing the latest consumed sequence for shards, and for coordinating balanced consumers of streams. **Type**: `object` ### [](#dynamodb-billing_mode)`dynamodb.billing_mode` When creating the table determines the billing mode. **Type**: `string` **Default**: `PAY_PER_REQUEST` **Options**: `PROVISIONED`, `PAY_PER_REQUEST` ### [](#dynamodb-create)`dynamodb.create` Whether, if the table does not exist, it should be created. **Type**: `bool` **Default**: `false` ### [](#dynamodb-credentials)`dynamodb.credentials` Manually configure the AWS credentials to use (optional). For more information, see the [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#dynamodb-credentials-from_ec2_role)`dynamodb.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#dynamodb-credentials-id)`dynamodb.credentials.id` The ID of the AWS credentials to use. **Type**: `string` ### [](#dynamodb-credentials-profile)`dynamodb.credentials.profile` The profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#dynamodb-credentials-role)`dynamodb.credentials.role` The role ARN to assume. **Type**: `string` ### [](#dynamodb-credentials-role_external_id)`dynamodb.credentials.role_external_id` An external ID to use when assuming a role. **Type**: `string` ### [](#dynamodb-credentials-secret)`dynamodb.credentials.secret` The secret for the AWS credentials in use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#dynamodb-credentials-token)`dynamodb.credentials.token` The token for the AWS credentials in use. This is a required value for short-term credentials. **Type**: `string` ### [](#dynamodb-endpoint)`dynamodb.endpoint` A custom endpoint URL for AWS API requests. Use this to connect to AWS-compatible services or local testing environments instead of the standard AWS endpoints. **Type**: `string` ### [](#dynamodb-read_capacity_units)`dynamodb.read_capacity_units` Set the provisioned read capacity when creating the table with a `billing_mode` of `PROVISIONED`. **Type**: `int` **Default**: `0` ### [](#dynamodb-region)`dynamodb.region` The AWS region to target. **Type**: `string` ### [](#dynamodb-table)`dynamodb.table` The name of the table to access. **Type**: `string` **Default**: `""` ### [](#dynamodb-tcp)`dynamodb.tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#dynamodb-tcp-connect_timeout)`dynamodb.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#dynamodb-tcp-keep_alive)`dynamodb.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#dynamodb-tcp-keep_alive-count)`dynamodb.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#dynamodb-tcp-keep_alive-idle)`dynamodb.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#dynamodb-tcp-keep_alive-interval)`dynamodb.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#dynamodb-tcp-tcp_user_timeout)`dynamodb.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#dynamodb-write_capacity_units)`dynamodb.write_capacity_units` Set the provisioned write capacity when creating the table with a `billing_mode` of `PROVISIONED`. **Type**: `int` **Default**: `0` ### [](#endpoint)`endpoint` A custom endpoint URL for AWS API requests. Use this to connect to AWS-compatible services or local testing environments instead of the standard AWS endpoints. **Type**: `string` ### [](#lease_period)`lease_period` The period of time after which a client that has failed to update a shard checkpoint is assumed to be inactive. **Type**: `string` **Default**: `30s` ### [](#rebalance_period)`rebalance_period` The period of time between each attempt to rebalance shards across clients. **Type**: `string` **Default**: `30s` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#start_from_oldest)`start_from_oldest` Whether to consume from the oldest message when a sequence does not yet exist for the stream. **Type**: `bool` **Default**: `true` ### [](#steal_grace_period)`steal_grace_period` Determines how long beyond the next commit period a client will wait when stealing a shard for the current owner to store a checkpoint. A longer value increases the time taken to balance shards but reduces the likelihood of processing duplicate messages. **Type**: `string` **Default**: `2s` ### [](#streams)`streams[]` One or more Kinesis data streams to consume from. Streams can either be specified by their name or full ARN. Shards of a stream are automatically balanced across consumers by coordinating through the provided DynamoDB table. Multiple comma separated streams can be listed in a single element. Shards are automatically distributed across consumers of a stream by coordinating through the provided DynamoDB table. Alternatively, it’s possible to specify an explicit shard to consume from with a colon after the stream name, e.g. `foo:0` would consume the shard `0` of the stream `foo`. **Type**: `array` ```yaml # Examples: streams: - foo - "arn:aws:kinesis:*:111122223333:stream/my-stream" ``` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` --- # Page 43: aws_s3 **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_s3.md --- # aws_s3 > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_s3 latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/aws_s3 page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/aws_s3.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/aws_s3.adoc description: Downloads objects within an Amazon S3 bucket, optionally filtered by a prefix, either by walking the items in the bucket or by streaming upload notifications in realtime. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Downloads objects within an Amazon S3 bucket, optionally filtered by a prefix, either by walking the items in the bucket or by streaming upload notifications in real time. #### Common ```yml inputs: label: "" aws_s3: bucket: "" prefix: "" scanner: to_the_end: {} sqs: url: "" endpoint: "" key_path: Records.*.s3.object.key bucket_path: Records.*.s3.bucket.name envelope_path: "" delay_period: "" max_messages: 10 wait_time_seconds: 0 nack_visibility_timeout: 0 zero_key_warn_interval: 30s ``` #### Advanced ```yml inputs: label: "" aws_s3: bucket: "" prefix: "" region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) force_path_style_urls: false delete_objects: false scanner: to_the_end: {} sqs: url: "" endpoint: "" key_path: Records.*.s3.object.key bucket_path: Records.*.s3.bucket.name envelope_path: "" delay_period: "" max_messages: 10 wait_time_seconds: 0 nack_visibility_timeout: 0 zero_key_warn_interval: 30s ``` ## [](#stream-objects-on-upload-with-sqs)Stream objects on upload with SQS A common pattern for consuming S3 objects is to emit upload notification events from the bucket either directly to an SQS queue, or to an SNS topic that is consumed by an SQS queue, and then have your consumer listen for events that prompt it to download the newly uploaded objects. More information about this pattern and how to set it up can be found in the [Amazon S3 docs](https://docs.aws.amazon.com/AmazonS3/latest/dev/ways-to-add-notification-config-to-bucket.html). Redpanda Connect is able to follow this pattern when you configure an `sqs.url`, where it consumes events from SQS and downloads only the object keys contained in those events. For this to work, Redpanda Connect needs to know where within the event the key and bucket names can be found, specified as [dot paths](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) with the fields `sqs.key_path` and `sqs.bucket_path`. The default values for these fields should already be correct when following the guide above. If your notification events are being routed to SQS via an SNS topic, the events are enveloped by SNS, in which case you also need to specify the field `sqs.envelope_path`, which in the case of SNS to SQS will usually be `Message`. When using SQS, make sure you have sensible values for `sqs.max_messages` and also the visibility timeout of the queue itself. When Redpanda Connect consumes an S3 object the SQS message that triggered it is not deleted until the S3 object has been sent onwards. This ensures at-least-once crash resiliency, but also means that if the S3 object takes longer to process than the visibility timeout of your queue, then the same objects might be processed multiple times. ## [](#download-large-files)Download large files When downloading large files, process them in streamed parts to avoid loading the entire file into memory at once. To do this, specify a [`scanner`](#scanner) that determines how to break the input into smaller individual messages. ## [](#bucket-and-prefix)Bucket and prefix The `bucket` field accepts a bucket name only, not an ARN. For example, use `my-bucket`, not `arn:aws:s3:::my-bucket`. The `prefix` field accepts a single string. To consume from multiple prefixes in the same bucket, use multiple `aws_s3` inputs in a [`broker` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/broker/): ```yaml input: broker: inputs: - aws_s3: bucket: my-bucket prefix: logs/app1/ - aws_s3: bucket: my-bucket prefix: logs/app2/ ``` ## [](#credentials)Credentials By default, Redpanda Connect uses a shared credentials file when connecting to AWS services. You can also set credentials explicitly at the component level to transfer data across accounts. You can find out more in [AWS credentials](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#s3-compatible-storage)S3-compatible storage The `endpoint` and `force_path_style_urls` fields let you connect to S3-compatible storage services such as Cloudflare R2, MinIO, or DigitalOcean Spaces. For Cloudflare R2, set `endpoint` to your account endpoint URL and enable `force_path_style_urls`: ```yaml input: aws_s3: bucket: r2-bucket endpoint: https://.r2.cloudflarestorage.com force_path_style_urls: true region: auto credentials: id: secret: ``` Find your account ID in the Cloudflare dashboard under **R2 > Overview > Account Details**. Generate API credentials under **R2 > Manage R2 API Tokens**. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `s3_key` - `s3_bucket` - `s3_last_modified_unix` - `s3_last_modified` (RFC3339) - `s3_content_type` - `s3_content_encoding` - `s3_version_id` - All user defined metadata You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). Note that user defined metadata is case insensitive within AWS, and it is likely that the keys will be received in a capitalized form, if you wish to make them consistent you can map all metadata keys to lower or uppercase using a Bloblang mapping such as `meta = meta().map_each_key(key → key.lowercase())`. ## [](#fields)Fields ### [](#bucket)`bucket` The bucket to consume from. If the field `sqs.url` is specified this field is optional. **Type**: `string` **Default**: `""` ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#delete_objects)`delete_objects` Whether to delete downloaded objects from the bucket once they are processed. **Type**: `bool` **Default**: `false` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#force_path_style_urls)`force_path_style_urls` Forces the client API to use path style URLs for downloading keys, which is often required when connecting to custom endpoints. **Type**: `bool` **Default**: `false` ### [](#prefix)`prefix` An optional path prefix, if set only objects with the prefix are consumed when walking a bucket. **Type**: `string` **Default**: `""` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#scanner)`scanner` The [scanner](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/about/) by which the stream of bytes consumed will be broken out into individual messages. Scanners are useful for processing large sources of data without holding the entirety of it within memory. For example, the `csv` scanner allows you to process individual CSV rows without loading the entire CSV file in memory at once. **Type**: `scanner` **Default**: ```yaml to_the_end: {} ``` ### [](#sqs)`sqs` Consume SQS messages in order to trigger key downloads. **Type**: `object` ### [](#sqs-bucket_path)`sqs.bucket_path` A [dot path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) whereby the bucket name can be found in SQS messages. **Type**: `string` **Default**: `Records.*.s3.bucket.name` ### [](#sqs-delay_period)`sqs.delay_period` An optional period of time to wait from when a notification was originally sent to when the target key download is attempted. **Type**: `string` **Default**: `""` ```yaml # Examples: delay_period: 10s # --- delay_period: 5m ``` ### [](#sqs-endpoint)`sqs.endpoint` A custom endpoint to use when connecting to SQS. **Type**: `string` **Default**: `""` ### [](#sqs-envelope_path)`sqs.envelope_path` A [dot path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) of a field to extract an enveloped JSON payload for further extracting the key and bucket from SQS messages. This is specifically useful when subscribing an SQS queue to an SNS topic that receives bucket events. **Type**: `string` **Default**: `""` ```yaml # Examples: envelope_path: Message ``` ### [](#sqs-key_path)`sqs.key_path` A [dot path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) whereby object keys are found in SQS messages. **Type**: `string` **Default**: `Records.*.s3.object.key` ### [](#sqs-max_messages)`sqs.max_messages` The maximum number of SQS messages to consume from each request. **Type**: `int` **Default**: `10` ### [](#sqs-nack_visibility_timeout)`sqs.nack_visibility_timeout` Custom SQS Nack Visibility timeout in seconds. Default is 0 **Type**: `int` **Default**: `0` ### [](#sqs-url)`sqs.url` An optional SQS URL to connect to. When specified this queue will control which objects are downloaded. **Type**: `string` **Default**: `""` ### [](#sqs-wait_time_seconds)`sqs.wait_time_seconds` Whether to set the wait time. Enabling this activates long-polling. Valid values: 0 to 20. **Type**: `int` **Default**: `0` ### [](#sqs-zero_key_warn_interval)`sqs.zero_key_warn_interval` A message from which no target key can be extracted (for example due to a misconfigured `key_path`/`bucket_path`) is never deleted, so it’s redelivered and re-evaluated repeatedly until the underlying issue is fixed. This field limits how often that condition is logged as a warning, to avoid flooding the logs; every occurrence is still logged at debug level. Set to `0s` to warn on every occurrence. **Type**: `string` **Default**: `30s` ```yaml # Examples: zero_key_warn_interval: 10s # --- zero_key_warn_interval: 0s ``` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` --- # Page 44: aws_sqs **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_sqs.md --- # aws_sqs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_sqs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/aws_sqs page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/aws_sqs.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/aws_sqs.adoc description: Consume messages from an AWS SQS URL. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consume messages from an AWS SQS URL. #### Common ```yml inputs: label: "" aws_sqs: url: "" # No default (required) max_outstanding_messages: 1000 ``` #### Advanced ```yml inputs: label: "" aws_sqs: url: "" # No default (required) delete_message: true reset_visibility: true max_number_of_messages: 10 max_outstanding_messages: 1000 wait_time_seconds: 0 message_timeout: 30s region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` ## [](#credentials)Credentials By default, Redpanda Connect uses a shared credentials file when connecting to AWS services. You can also set credentials explicitly at the component level, which allows you to transfer data across accounts. To find out more, see [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `sqs_message_id` - `sqs_receipt_handle` - `sqs_approximate_receive_count` - All message attributes You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#delete_message)`delete_message` Whether to delete the consumed message when it’s acknowledged. Set to `false` to handle the deletion using a different mechanism. **Type**: `bool` **Default**: `true` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#max_number_of_messages)`max_number_of_messages` The maximum number of messages that Redpanda Connect can return each time it polls the SQS URL. Enter values from `1` to `10` only. **Type**: `int` **Default**: `10` ### [](#max_outstanding_messages)`max_outstanding_messages` The maximum number of pending messages that Redpanda Connect can have in flight at the same time. **Type**: `int` **Default**: `1000` ### [](#message_timeout)`message_timeout` The maximum time allowed to process a received message before Redpanda Connect refreshes the [receipt handle](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-queue-message-identifiers.html), and the message becomes visible in the queue again. Redpanda Connect attempts to refresh the receipt handle after half of the timeout has elapsed. **Type**: `string` **Default**: `30s` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#reset_visibility)`reset_visibility` Whether to set the visibility timeout of the consumed message to zero if Redpanda Connect receives a negative acknowledgement. Set to `false` to use the [queue’s visibility timeout](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html) for each message rather than releasing the message immediately for reprocessing. **Type**: `bool` **Default**: `true` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#url)`url` The SQS URL to consume from. **Type**: `string` ### [](#wait_time_seconds)`wait_time_seconds` Whether to set a wait time (in seconds). Enter values from `1` to `20` to enable wait times and to activate [log polling](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-short-and-long-polling.html) for queued messages. **Type**: `int` **Default**: `0` --- # Page 45: azure_blob_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/azure_blob_storage.md --- # azure_blob_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_blob_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/azure_blob_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/azure_blob_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/azure_blob_storage.adoc description: Downloads objects within an Azure Blob Storage container, optionally filtered by a prefix. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Downloads objects within an Azure Blob Storage container, optionally filtered by a prefix. #### Common ```yml inputs: label: "" azure_blob_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" container: "" # No default (required) prefix: "" scanner: to_the_end: {} targets_input: "" # No default (optional) ``` #### Advanced ```yml inputs: label: "" azure_blob_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" container: "" # No default (required) prefix: "" scanner: to_the_end: {} delete_objects: false targets_input: "" # No default (optional) ``` Supports multiple authentication methods but only one of the following is required: - `storage_connection_string` - `storage_account` and `storage_access_key` - `storage_account` and `storage_sas_token` - `storage_account` to access via [DefaultAzureCredential](https://pkg.go.dev/github.com/Azure/azure-sdk-for-go/sdk/azidentity#DefaultAzureCredential) If multiple are set then the `storage_connection_string` is given priority. If the `storage_connection_string` does not contain the `AccountName` parameter, please specify it in the `storage_account` field. ## [](#download-large-files)Download large files When downloading large files it’s often necessary to process it in streamed parts in order to avoid loading the entire file in memory at a given time. In order to do this a [`scanner`](#scanner) can be specified that determines how to break the input into smaller individual messages. ## [](#stream-new-files)Stream new files By default this input will consume all files found within the target container and will then gracefully terminate. This is referred to as a "batch" mode of operation. However, it’s possible to instead configure a container as [an Event Grid source](https://learn.microsoft.com/en-gb/azure/event-grid/event-schema-blob-storage) and then use this as a [`targets_input`](#targets_input), in which case new files are consumed as they’re uploaded and Redpanda Connect will continue listening for and downloading files as they arrive. This is referred to as a "streamed" mode of operation. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `blob_storage_key` - `blob_storage_container` - `blob_storage_last_modified` - `blob_storage_last_modified_unix` - `blob_storage_content_type` - `blob_storage_content_encoding` - All user defined metadata You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#container)`container` The name of the container from which to download blobs. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#delete_objects)`delete_objects` Whether to delete downloaded objects from the blob once they are processed. **Type**: `bool` **Default**: `false` ### [](#prefix)`prefix` An optional path prefix, if set only objects with the prefix are consumed. **Type**: `string` **Default**: `""` ### [](#scanner)`scanner` The [scanner](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/about/) by which the stream of bytes consumed will be broken out into individual messages. Scanners are useful for processing large sources of data without holding the entirety of it within memory. For example, the `csv` scanner allows you to process individual CSV rows without loading the entire CSV file in memory at once. **Type**: `scanner` **Default**: ```yaml to_the_end: {} ``` ### [](#storage_access_key)`storage_access_key` The storage account access key. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_account)`storage_account` The storage account to access. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_connection_string)`storage_connection_string` A storage account connection string. This field is required if `storage_account` and `storage_access_key` / `storage_sas_token` are not set. **Type**: `string` **Default**: `""` ### [](#storage_sas_token)`storage_sas_token` The storage account SAS token. This field is ignored if `storage_connection_string` or `storage_access_key` are set. **Type**: `string` **Default**: `""` ### [](#targets_input)`targets_input` > ⚠️ **CAUTION** > > This is an experimental field that provides an optional source of download targets, configured as a [regular Redpanda Connect input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/about/). Each message yielded by this input should be a single structured object containing a field `name`, which represents the blob to be downloaded. This requires setting up [Azure Blob Storage as an Event Grid source](https://learn.microsoft.com/en-gb/azure/event-grid/event-schema-blob-storage) and an associated event handler that a Redpanda Connect input can read from. For example, use either one of the following: - [Azure Event Hubs](https://learn.microsoft.com/en-gb/azure/event-grid/handler-event-hubs) using the `kafka` input - [Namespace topics](https://learn.microsoft.com/en-gb/azure/event-grid/handler-event-grid-namespace-topic) using the `mqtt` input **Type**: `input` ```yaml # Examples: targets_input: mqtt: topics: - some-topic urls: - example.westeurope-1.ts.eventgrid.azure.net:8883 processors: - unarchive: format: json_array - mapping: |- if this.eventType == "Microsoft.Storage.BlobCreated" { root.name = this.data.url.parse_url().path.trim_prefix("/foocontainer/") } else { root = deleted() } ``` --- # Page 46: azure_cosmosdb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/azure_cosmosdb.md --- # azure_cosmosdb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_cosmosdb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/azure_cosmosdb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/azure_cosmosdb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/azure_cosmosdb.adoc description: Executes a SQL query against Azure CosmosDB and creates a batch of messages from each page of items. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a SQL query against [Azure CosmosDB](https://learn.microsoft.com/en-us/azure/cosmos-db/introduction) and creates a batch of messages from each page of items. ### Common ```yml inputs: label: "" azure_cosmosdb: endpoint: "" # No default (optional) account_key: "" # No default (optional) connection_string: "" # No default (optional) database: "" # No default (required) container: "" # No default (required) partition_keys_map: "" # No default (required) query: "" # No default (required) args_mapping: "" # No default (optional) auto_replay_nacks: true ``` ### Advanced ```yml inputs: label: "" azure_cosmosdb: endpoint: "" # No default (optional) account_key: "" # No default (optional) connection_string: "" # No default (optional) database: "" # No default (required) container: "" # No default (required) partition_keys_map: "" # No default (required) query: "" # No default (required) args_mapping: "" # No default (optional) batch_count: -1 auto_replay_nacks: true ``` ## [](#cross-partition-queries)Cross-partition queries Cross-partition queries are currently not supported by the underlying driver. For every query, the PartitionKey values must be known in advance and specified in the config. [See details](https://github.com/Azure/azure-sdk-for-go/issues/18578#issuecomment-1222510989). ## [](#credentials)Credentials You can use one of the following authentication mechanisms: - Set the `endpoint` field and the `account_key` field - Set only the `endpoint` field to use [DefaultAzureCredential](https://pkg.go.dev/github.com/Azure/azure-sdk-for-go/sdk/azidentity#DefaultAzureCredential) - Set the `connection_string` field ## [](#metadata)Metadata This component adds the following metadata fields to each message: - `activity_id` - `request_charge` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#examples)Examples ### [](#query-container)Query container Execute a parametrized SQL query to select documents from a container. ```yaml input: azure_cosmosdb: endpoint: http://localhost:8080 account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== database: blobbase container: blobfish partition_keys_map: root = "AbyssalPlain" query: SELECT * FROM blobfish AS b WHERE b.species = @species args_mapping: | root = [ { "Name": "@species", "Value": "smooth-head" }, ] ``` ## [](#fields)Fields ### [](#account_key)`account_key` Account key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== ``` ### [](#args_mapping)`args_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that, for each message, creates a list of arguments to use with the query. **Type**: `string` ```yaml # Examples: args_mapping: |- root = [ { "Name": "@name", "Value": "benthos" }, ] ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batch_count)`batch_count` The maximum number of messages that should be accumulated into each batch. Use '-1' specify dynamic page size. **Type**: `int` **Default**: `-1` ### [](#connection_string)`connection_string` Connection string. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: connection_string: AccountEndpoint=https://localhost:8081/;AccountKey=C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw==; ``` ### [](#container)`container` Container. **Type**: `string` ```yaml # Examples: container: testcontainer ``` ### [](#database)`database` Database. **Type**: `string` ```yaml # Examples: database: testdb ``` ### [](#endpoint)`endpoint` CosmosDB endpoint. **Type**: `string` ```yaml # Examples: endpoint: https://localhost:8081 ``` ### [](#partition_keys_map)`partition_keys_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to a single partition key value or an array of partition key values of type string, integer or boolean. Currently, hierarchical partition keys are not supported so only one value may be provided. **Type**: `string` ```yaml # Examples: partition_keys_map: root = "blobfish" # --- partition_keys_map: root = 41 # --- partition_keys_map: root = true # --- partition_keys_map: root = null # --- partition_keys_map: root = now().ts_format("2006-01-02") ``` ### [](#query)`query` The query to execute **Type**: `string` ```yaml # Examples: query: SELECT c.foo FROM testcontainer AS c WHERE c.bar = "baz" AND c.timestamp < @timestamp ``` ## [](#cosmosdb-emulator)CosmosDB emulator If you wish to run the CosmosDB emulator that is referenced in the documentation [here](https://learn.microsoft.com/en-us/azure/cosmos-db/linux-emulator), the following Docker command should do the trick: ```bash > docker run --rm -it -p 8081:8081 --name=cosmosdb -e AZURE_COSMOS_EMULATOR_PARTITION_COUNT=10 -e AZURE_COSMOS_EMULATOR_ENABLE_DATA_PERSISTENCE=false mcr.microsoft.com/cosmosdb/linux/azure-cosmos-emulator ``` Note: `AZURE_COSMOS_EMULATOR_PARTITION_COUNT` controls the number of partitions that will be supported by the emulator. The bigger the value, the longer it takes for the container to start up. Additionally, instead of installing the container self-signed certificate which is exposed via `[https://localhost:8081/_explorer/emulator.pem](https://localhost:8081/_explorer/emulator.pem)`, you can run [mitmproxy](https://mitmproxy.org/) like so: ```bash > mitmproxy -k --mode "reverse:https://localhost:8081" ``` Then you can access the CosmosDB UI via `[http://localhost:8080/_explorer/index.html](http://localhost:8080/_explorer/index.html)` and use `[http://localhost:8080](http://localhost:8080)` as the CosmosDB endpoint. --- # Page 47: azure_queue_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/azure_queue_storage.md --- # azure_queue_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_queue_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/azure_queue_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/azure_queue_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/azure_queue_storage.adoc description: Dequeue objects from an Azure Storage Queue. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Dequeue objects from an Azure Storage Queue. #### Common ```yml inputs: label: "" azure_queue_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" queue_name: "" # No default (required) ``` #### Advanced ```yml inputs: label: "" azure_queue_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" queue_name: "" # No default (required) dequeue_visibility_timeout: 30s max_in_flight: 10 track_properties: false ``` This input adds the following metadata fields to each message: ```none - queue_storage_insertion_time - queue_storage_queue_name - queue_storage_message_lag (if 'track_properties' set to true) - All user defined queue metadata ``` Only one authentication method is required, `storage_connection_string` or `storage_account` and `storage_access_key`. If both are set then the `storage_connection_string` is given priority. ## [](#fields)Fields ### [](#dequeue_visibility_timeout)`dequeue_visibility_timeout` The timeout duration until a dequeued message gets visible again, 30s by default **Type**: `string` **Default**: `30s` ### [](#max_in_flight)`max_in_flight` The maximum number of unprocessed messages to fetch at a given time. **Type**: `int` **Default**: `10` ### [](#queue_name)`queue_name` The name of the source storage queue. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: queue_name: foo_queue # --- queue_name: ${! env("MESSAGE_TYPE").lowercase() } ``` ### [](#storage_access_key)`storage_access_key` The storage account access key. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_account)`storage_account` The storage account to access. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_connection_string)`storage_connection_string` A storage account connection string. This field is required if `storage_account` and `storage_access_key` / `storage_sas_token` are not set. **Type**: `string` **Default**: `""` ### [](#track_properties)`track_properties` If set to `true` the queue is polled on each read request for information such as the queue message lag. These properties are added to consumed messages as metadata, but will also have a negative performance impact. **Type**: `bool` **Default**: `false` --- # Page 48: azure_table_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/azure_table_storage.md --- # azure_table_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_table_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/azure_table_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/azure_table_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/azure_table_storage.adoc description: Queries an Azure Storage Account Table, optionally with multiple filters. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Queries an Azure Storage Account Table, optionally with multiple filters. #### Common ```yml inputs: label: "" azure_table_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" table_name: "" # No default (required) ``` #### Advanced ```yml inputs: label: "" azure_table_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" table_name: "" # No default (required) filter: "" select: "" page_size: 1000 ``` Queries an Azure Storage Account Table, optionally with multiple filters. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `table_storage_name` - `row_num` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#filter)`filter` OData filter expression. Is not set all rows are returned. Valid operators are `eq, ne, gt, lt, ge and le` **Type**: `string` **Default**: `""` ```yaml # Examples: filter: PartitionKey eq 'foo' and RowKey gt '1000' ``` ### [](#page_size)`page_size` Maximum number of records to return on each page. **Type**: `int` **Default**: `1000` ### [](#select)`select` Select expression using OData notation. Limits the columns on each record to just those requested. **Type**: `string` **Default**: `""` ```yaml # Examples: select: PartitionKey,RowKey,Foo,Bar,Timestamp ``` ### [](#storage_access_key)`storage_access_key` The storage account access key. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_account)`storage_account` The storage account to access. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_connection_string)`storage_connection_string` A storage account connection string. This field is required if `storage_account` and `storage_access_key` / `storage_sas_token` are not set. **Type**: `string` **Default**: `""` ### [](#storage_sas_token)`storage_sas_token` The storage account SAS token. This field is ignored if `storage_connection_string` or `storage_access_key` are set. **Type**: `string` **Default**: `""` ### [](#table_name)`table_name` The table to read messages from. **Type**: `string` ```yaml # Examples: table_name: Foo ``` --- # Page 49: batched **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/batched.md --- # batched > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: batched latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/batched page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/batched.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/batched.adoc description: Consumes data from a child input and applies a batching policy to the stream. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consumes data from a child input and applies a batching policy to the stream. #### Common ```yml inputs: label: "" batched: child: "" # No default (required) policy: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml inputs: label: "" batched: child: "" # No default (required) policy: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` Batching at the input level is sometimes useful for processing across micro-batches, and can also sometimes be a useful performance trick. However, most inputs are fine without it so unless you have a specific plan for batching this component is not worth using. ## [](#fields)Fields ### [](#child)`child` The child input. **Type**: `input` ### [](#policy)`policy` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: policy: byte_size: 5000 count: 0 period: 1s # --- policy: count: 10 period: 1s # --- policy: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#policy-byte_size)`policy.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#policy-check)`policy.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#policy-count)`policy.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#policy-period)`policy.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#policy-processors)`policy.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` --- # Page 50: broker **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/broker.md --- # broker > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: broker latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/broker page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/broker.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/broker.adoc description: Allows you to combine multiple inputs into a single stream of data, where each input will be read in parallel. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Allows you to combine multiple inputs into a single stream of data, where each input will be read in parallel. #### Common ```yml inputs: label: "" broker: inputs: [] # No default (required) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml inputs: label: "" broker: copies: 1 inputs: [] # No default (required) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` A broker type is configured with its own list of input configurations and a field to specify how many copies of the list of inputs should be created. Adding more input types allows you to combine streams from multiple sources into one. For example, reading from both RabbitMQ and Kafka: ```yaml input: broker: copies: 1 inputs: - amqp_0_9: urls: - amqp://guest:guest@localhost:5672/ consumer_tag: benthos-consumer queue: benthos-queue # Optional list of input specific processing steps processors: - mapping: | root.message = this root.meta.link_count = this.links.length() root.user.age = this.user.age.number() - kafka: addresses: - localhost:9092 client_id: benthos_kafka_input consumer_group: benthos_consumer_group topics: [ benthos_stream:0 ] ``` If the number of copies is greater than zero the list will be copied that number of times. For example, if your inputs were of type foo and bar, with 'copies' set to '2', you would end up with two 'foo' inputs and two 'bar' inputs. ## [](#batching)Batching It’s possible to configure a [batch policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/#batch-policy) with a broker using the `batching` fields. When doing this the feeds from all child inputs are combined. Some inputs do not support broker based batching and specify this in their documentation. ## [](#processors)Processors It is possible to configure [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) at the broker level, where they will be applied to _all_ child inputs, as well as on the individual child inputs. If you have processors at both the broker level _and_ on child inputs then the broker processors will be applied _after_ the child nodes processors. ## [](#fields)Fields ### [](#batching-2)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#copies)`copies` Whatever is specified within `inputs` will be created this many times. **Type**: `int` **Default**: `1` ### [](#inputs)`inputs[]` A list of inputs to create. **Type**: `array` --- # Page 51: gateway **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gateway.md --- # gateway > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gateway latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/gateway page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/gateway.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/gateway.adoc description: Receive messages delivered over HTTP. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-08-11" --- The `gateway` input is a Cloud-only component that receives messages over HTTP and injects them into a running Redpanda Connect pipeline. It’s ideal for: - Receiving webhook events from third-party services - Accepting real-time telemetry or sensor data over HTTP - Building lightweight ingest endpoints for client apps For on-premises or self-managed deployments, use the [`http_server`](https://docs.redpanda.com/connect/components/inputs/http_server/) input instead. This component is fully managed and available in the following Redpanda Cloud deployment types: - **Serverless** - **Dedicated** - **Bring Your Own Cloud (BYOC)** When a pipeline with a `gateway` input is deployed, Redpanda Cloud provisions a secure URL that you can use to send HTTP requests. You can post raw payloads, JSON messages, or stream events in real time. Authentication and access control are handled through standard Redpanda Cloud API tokens. For more information, see [Cloud API Authentication](https://docs.redpanda.com/api/doc/cloud-dataplane/authentication). Network access: - On **public clusters** (Serverless and Dedicated), the gateway URL is accessible over the public internet. - On **private clusters** (BYOC), the gateway is accessible only from within your configured VPC. #### Common ```yaml input: label: "" gateway: path: / rate_limit: "" ``` #### Advanced ```yaml input: label: "" gateway: path: / rate_limit: "" sync_response: status: "200" headers: Content-Type: application/octet-stream metadata_headers: include_prefixes: [] include_patterns: [] ``` The field `rate_limit` allows you to specify an optional [`rate_limit` resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) that applies to all HTTP requests. When the rate limit is breached, HTTP requests return a 429 response with a Retry-After header. ## [](#responses)Responses You can also return a response for each message received using [synchronous responses](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/sync_responses/). When doing so, you can customize headers using the `sync_response.headers` field, which supports [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) in the value based on the response message contents. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `http_server_user_agent` - `http_server_request_path` - `http_server_verb` - `http_server_remote_ip` - All headers (only first values are taken) - All query parameters - All path parameters - All cookies You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#path)`path` The endpoint path to listen for data delivery requests. **Type**: `string` **Default**: `/` ### [](#rate_limit)`rate_limit` An optional [rate limit](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) to throttle requests by. **Type**: `string` **Default**: `""` ### [](#sync_response)`sync_response` Customize messages returned using [synchronous responses](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/sync_responses/). **Type**: `object` ### [](#sync_response-headers)`sync_response.headers` Specify headers to return with synchronous responses. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: ```yaml Content-Type: "application/octet-stream" ``` ### [](#sync_response-metadata_headers)`sync_response.metadata_headers` Specify criteria for which metadata values are added to the response as headers. **Type**: `object` ### [](#sync_response-metadata_headers-include_patterns)`sync_response.metadata_headers.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#sync_response-metadata_headers-include_prefixes)`sync_response.metadata_headers.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#sync_response-status)`sync_response.status` Specify the status code to return with synchronous responses. This is a string value, which allows you to customize it based on resulting payloads and their metadata. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `200` ```yaml # Examples: status: ${! json("status") } # --- status: ${! meta("status") } ``` ### [](#tcp)`tcp` Customize messages returned via [synchronous responses](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/sync_responses/). **Type**: `object` ### [](#tcp-reuse_addr)`tcp.reuse_addr` Enable SO\_REUSEADDR, allowing binding to ports in TIME\_WAIT state. Useful for graceful restarts and config reloads where the server needs to rebind to the same port immediately after shutdown. **Type**: `bool` **Default**: `false` ### [](#tcp-reuse_port)`tcp.reuse_port` Enable SO\_REUSEPORT, allowing multiple sockets to bind to the same port for load balancing across multiple processes/threads. **Type**: `bool` **Default**: `false` ## [](#examples)Examples ### [](#ingest-a-real-time-stream-of-sensor-data)Ingest a real-time stream of sensor data Use the `gateway` input to stream telemetry data from edge devices or browser clients that connect over HTTP. Suppose a client connects and sends JSON-encoded sensor readings like this: ```json { "sensor_id": "temp-001", "value": 22.5, "unit": "C" } { "sensor_id": "temp-001", "value": 22.8, "unit": "C" } { "sensor_id": "temp-001", "value": 23.1, "unit": "C" } ``` Redpanda Connect treats each line as an individual message. The following pipeline sets up a `gateway` input to handle these connections and logs each message: ```yaml input: label: sensor_stream gateway: path: /ws/sensors rate_limit: "" pipeline: processors: - log: level: INFO message: "Received reading from ${! json(\"sensor_id\") }: ${! json(\"value\") } ${! json(\"unit\") }" ``` This configuration: - Accepts HTTP connections on `/ws/sensors` - Receives a stream of messages over a single connection - Logs each message using Bloblang interpolation You can replace the `log` processor with any downstream output, such as Redpanda or Amazon S3, to persist or analyze the data in real time. --- # Page 52: gcp_bigquery_select **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gcp_bigquery_select.md --- # gcp_bigquery_select > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_bigquery_select latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/gcp_bigquery_select page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/gcp_bigquery_select.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/gcp_bigquery_select.adoc description: Executes a SELECT query against BigQuery and creates a message for each row received. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a `SELECT` query against BigQuery and creates a message for each row received. ```yml inputs: label: "" gcp_bigquery_select: project: "" # No default (required) credentials_json: "" table: "" # No default (required) columns: [] # No default (required) where: "" # No default (optional) auto_replay_nacks: true job_labels: {} priority: "" args_mapping: "" # No default (optional) prefix: "" # No default (optional) suffix: "" # No default (optional) ``` Once the rows from the query are exhausted, this input shuts down, allowing the pipeline to gracefully terminate (or the next input in a [sequence](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/) to execute). ## [](#examples)Examples ### [](#word-counts)Word counts Here we query the public corpus of Shakespeare’s works to generate a stream of the top 10 words that are 3 or more characters long: ```yaml input: gcp_bigquery_select: project: sample-project table: bigquery-public-data.samples.shakespeare columns: - word - sum(word_count) as total_count where: length(word) >= ? suffix: | GROUP BY word ORDER BY total_count DESC LIMIT 10 args_mapping: | root = [ 3 ] ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of placeholder arguments in the field `where`. **Type**: `string` ```yaml # Examples: args_mapping: root = [ "article", now().ts_format("2006-01-02") ] ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#columns)`columns[]` A list of columns to query. **Type**: `array` ### [](#credentials_json)`credentials_json` Base64-encoded Google Service Account credentials in JSON format (optional). Use this field to authenticate with Google Cloud services. For more information about creating service account credentials, see [Google’s service account documentation](https://developers.google.com/workspace/guides/create-credentials#create_credentials_for_a_service_account). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#job_labels)`job_labels` A list of labels to add to the query job. **Type**: `object` **Default**: `{}` ### [](#prefix)`prefix` An optional prefix to prepend to the select query (before SELECT). **Type**: `string` ### [](#priority)`priority` The priority with which to schedule the query. **Type**: `string` **Default**: `""` ### [](#project)`project` GCP project where the query job will execute. **Type**: `string` ### [](#suffix)`suffix` An optional suffix to append to the select query. **Type**: `string` ### [](#table)`table` Fully-qualified BigQuery table name to query. **Type**: `string` ```yaml # Examples: table: bigquery-public-data.samples.shakespeare ``` ### [](#where)`where` An optional where clause to add. Placeholder arguments are populated with the `args_mapping` field. Placeholders should always be question marks (`?`). **Type**: `string` ```yaml # Examples: where: type = ? and created_at > ? # --- where: user_id = ? ``` --- # Page 53: gcp_cloud_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gcp_cloud_storage.md --- # gcp_cloud_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_cloud_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/gcp_cloud_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/gcp_cloud_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/gcp_cloud_storage.adoc description: Downloads objects within a Google Cloud Storage bucket, optionally filtered by a prefix. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Downloads objects within a Google Cloud Storage bucket, optionally filtered by a prefix. #### Common ```yml inputs: label: "" gcp_cloud_storage: bucket: "" # No default (required) prefix: "" credentials_json: "" scanner: to_the_end: {} ``` #### Advanced ```yml inputs: label: "" gcp_cloud_storage: bucket: "" # No default (required) prefix: "" credentials_json: "" scanner: to_the_end: {} delete_objects: false ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `gcs_key` - `gcs_bucket` - `gcs_last_modified` - `gcs_last_modified_unix` - `gcs_content_type` - `gcs_content_encoding` - All user defined metadata You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#bucket)`bucket` The name of the bucket from which to download objects. **Type**: `string` ### [](#credentials_json)`credentials_json` Base64-encoded Google Service Account credentials in JSON format (optional). Use this field to authenticate with Google Cloud services. For more information about creating service account credentials, see [Google’s service account documentation](https://developers.google.com/workspace/guides/create-credentials#create_credentials_for_a_service_account). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#delete_objects)`delete_objects` Whether to delete downloaded objects from the bucket once they are processed. **Type**: `bool` **Default**: `false` ### [](#prefix)`prefix` Optional path prefix, if set only objects with the prefix are consumed. **Type**: `string` **Default**: `""` ### [](#scanner)`scanner` The [scanner](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/about/) by which the stream of bytes consumed will be broken out into individual messages. Scanners are useful for processing large sources of data without holding the entirety of it within memory. For example, the `csv` scanner allows you to process individual CSV rows without loading the entire CSV file in memory at once. **Type**: `scanner` **Default**: ```yaml to_the_end: {} ``` --- # Page 54: gcp_pubsub **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gcp_pubsub.md --- # gcp_pubsub > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_pubsub latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/gcp_pubsub page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/gcp_pubsub.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/gcp_pubsub.adoc description: Consumes messages from a GCP Cloud Pub/Sub subscription. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consumes messages from a GCP Cloud Pub/Sub subscription. #### Common ```yml inputs: label: "" gcp_pubsub: project: "" # No default (required) credentials_json: "" subscription: "" # No default (required) endpoint: "" sync: false max_outstanding_messages: 1000 max_outstanding_bytes: 1000000000 ``` #### Advanced ```yml inputs: label: "" gcp_pubsub: project: "" # No default (required) credentials_json: "" subscription: "" # No default (required) endpoint: "" sync: false max_outstanding_messages: 1000 max_outstanding_bytes: 1000000000 create_subscription: enabled: false topic: "" ``` For information on how to set up credentials see [this guide](https://cloud.google.com/docs/authentication/production). ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `gcp_pubsub_publish_time_unix` - The time at which the message was published to the topic. - `gcp_pubsub_delivery_attempt` - When dead lettering is enabled, this is set to the number of times PubSub has attempted to deliver a message. - `gcp_pubsub_message_id` - The unique identifier of the message. - `gcp_pubsub_ordering_key` - The ordering key of the message. - All message attributes You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#create_subscription)`create_subscription` Allows you to configure the input subscription and creates if it doesn’t exist. **Type**: `object` ### [](#create_subscription-enabled)`create_subscription.enabled` Whether to configure subscription or not. **Type**: `bool` **Default**: `false` ### [](#create_subscription-topic)`create_subscription.topic` Defines the topic that the subscription should be vinculated to. **Type**: `string` **Default**: `""` ### [](#credentials_json)`credentials_json` Base64-encoded Google Service Account credentials in JSON format (optional). Use this field to authenticate with Google Cloud services. For more information about creating service account credentials, see [Google’s service account documentation](https://developers.google.com/workspace/guides/create-credentials#create_credentials_for_a_service_account). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#endpoint)`endpoint` An optional endpoint to override the default of `pubsub.googleapis.com:443`. This can be used to connect to a region specific pubsub endpoint. For a list of valid values, see [this document](https://cloud.google.com/pubsub/docs/reference/service_apis_overview#list_of_regional_endpoints). **Type**: `string` **Default**: `""` ```yaml # Examples: endpoint: us-central1-pubsub.googleapis.com:443 # --- endpoint: us-west3-pubsub.googleapis.com:443 ``` ### [](#max_outstanding_bytes)`max_outstanding_bytes` The maximum number of outstanding pending messages to be consumed measured in bytes. **Type**: `int` **Default**: `1000000000` ### [](#max_outstanding_messages)`max_outstanding_messages` The maximum number of outstanding pending messages to be consumed at a given time. **Type**: `int` **Default**: `1000` ### [](#project)`project` The project ID of the target subscription. **Type**: `string` ### [](#subscription)`subscription` The target subscription ID. **Type**: `string` ### [](#sync)`sync` Enable synchronous pull mode. **Type**: `bool` **Default**: `false` --- # Page 55: gcp_spanner_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gcp_spanner_cdc.md --- # gcp_spanner_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_spanner_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/gcp_spanner_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/gcp_spanner_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/gcp_spanner_cdc.adoc description: Creates an input that consumes from a spanner change stream. page-git-created-date: "2025-07-08" page-git-modified-date: "2026-08-11" --- Creates an input that consumes from a spanner change stream. #### Common ```yaml inputs: label: "" gcp_spanner_cdc: credentials_json: "" project_id: "" # No default (required) instance_id: "" # No default (required) database_id: "" # No default (required) stream_id: "" # No default (required) start_timestamp: "" end_timestamp: "" batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) auto_replay_nacks: true ``` #### Advanced ```yaml inputs: label: "" gcp_spanner_cdc: credentials_json: "" project_id: "" # No default (required) instance_id: "" # No default (required) database_id: "" # No default (required) stream_id: "" # No default (required) start_timestamp: "" end_timestamp: "" heartbeat_interval: 10s metadata_table: "" min_watermark_cache_ttl: 5s allowed_mod_types: [] # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) auto_replay_nacks: true ``` Consumes change records from a Google Cloud Spanner change stream. This input allows you to track and process database changes in real-time, making it useful for data replication, event-driven architectures, and maintaining derived data stores. The input reads from a specified change stream within a Spanner database and converts each change record into a message. The message payload contains the change records in JSON format, and metadata is added with details about the Spanner instance, database, and stream. Change streams provide a way to track mutations to your Spanner database tables. For more information about Spanner change streams, refer to the [Google Cloud documentation](https://cloud.google.com/spanner/docs/change-streams). ## [](#fields)Fields ### [](#allowed_mod_types)`allowed_mod_types[]` List of modification types to process. If not specified, all modification types are processed. Allowed values: INSERT, UPDATE, DELETE **Type**: `array` ```yaml # Examples: allowed_mod_types: - INSERT - UPDATE - DELETE ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The maximum total size (in bytes) that a batch can reach before it is passed on for processing or delivery (flushed). When the combined size of all messages in the batch exceeds this limit, the batch is immediately sent to the next stage (such as a processor or output). Set to `0` to disable size-based batching. When disabled, messages are flushed based on other conditions (such as `batching.count` or `batching.period`). **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages at which the batch should be flushed. Set the value to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The length of time after which an incomplete batch should be flushed regardless of its size. Supported time units are `ns`, `us`, `ms`, `s`, `m`, and `h`. For example, `1s` flushes a batch after one second. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, so any attempt to split it into smaller batches with these processors will be ignored. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#credentials_json)`credentials_json` Base64-encoded JSON credentials file for authenticating to GCP with a service account. If not provided, Application Default Credentials (ADC) is used. For more information about how to create a service account and obtain the credentials JSON, see the [Google Cloud documentation](https://cloud.google.com/docs/authentication/getting-started). **Type**: `string` **Default**: `""` ### [](#database_id)`database_id` The ID of the Spanner database to read from. This is the name of the database as it appears in the Spanner console or API. For more information about how to create a Spanner database, see the [Google Cloud documentation](https://cloud.google.com/spanner/docs/create-manage-databases). **Type**: `string` ### [](#end_timestamp)`end_timestamp` The timestamp at which to stop reading change records from the change stream. This is an optional field that allows you to limit the range of change records processed by the input. The timestamp should be in RFC3339 format, such as `2023-10-01T00:00:00Z`. If not provided, the input reads all available change records up to the current time. **Type**: `string` **Default**: `""` ```yaml # Examples: end_timestamp: 2022-01-01T00:00:00Z ``` ### [](#heartbeat_interval)`heartbeat_interval` The interval at which to send heartbeat messages to the output. Heartbeat messages are sent to indicate that the input is still active and processing changes. This can help prevent timeouts in downstream systems. Supported time units are `ns`, `us`, `ms`, `s`, `m`, and `h`. For example, `1s` sends a heartbeat every second. **Type**: `string` **Default**: `10s` ### [](#instance_id)`instance_id` The ID of the Spanner instance to read from. This is the name of the instance as it appears in the Spanner console or API. For more information about how to create a Spanner instance, see the [Google Cloud documentation](https://cloud.google.com/spanner/docs/create-manage-instances). **Type**: `string` ### [](#metadata_table)`metadata_table` The table to store metadata in (default: `cdc_metadata_`). **Type**: `string` **Default**: `""` ### [](#min_watermark_cache_ttl)`min_watermark_cache_ttl` Sets how frequently to query Spanner for the minimum watermark. **Type**: `string` **Default**: `5s` ### [](#project_id)`project_id` The ID of the GCP project that contains the Spanner instance and database. This is the name of the project as it appears in the GCP console or API. For more information about how to create a GCP project, see the [Google Cloud documentation](https://cloud.google.com/resource-manager/docs/creating-managing-projects). **Type**: `string` ### [](#start_timestamp)`start_timestamp` The timestamp at which to start reading change records from the change stream. This is an optional field that allows you to limit the range of change records processed by the input. The timestamp should be in RFC3339 format, such as `2023-10-01T00:00:00Z` (default: current time). **Type**: `string` **Default**: `""` ```yaml # Examples: start_timestamp: 2022-01-01T00:00:00Z ``` ### [](#stream_id)`stream_id` The name of the change stream to track. The stream must exist in the Spanner database. To create a change stream, follow the [Google Cloud documentation](https://cloud.google.com/spanner/docs/change-streams/manage). **Type**: `string` --- # Page 56: generate **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/generate.md --- # generate > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: generate latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/generate page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/generate.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/generate.adoc description: Generates messages at a given interval using a Bloblang mapping executed without a context. This allows you to generate messages for testing your pipeline configs. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates messages at a given interval using a [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) mapping executed without a context. This allows you to generate messages for testing your pipeline configs. #### Common ```yml inputs: label: "" generate: mapping: "" # No default (required) interval: 1s count: 0 batch_size: 1 auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" generate: mapping: "" # No default (required) interval: 1s count: 0 batch_size: 1 auto_replay_nacks: true ``` ## [](#examples)Examples ### [](#cron-scheduled-processing)Cron Scheduled Processing A common use case for the generate input is to trigger processors on a schedule so that the processors themselves can behave similarly to an input. The following configuration reads rows from a PostgreSQL table every 5 minutes. ```yaml input: generate: interval: '@every 5m' mapping: 'root = {}' processors: - sql_select: driver: postgres dsn: postgres://foouser:foopass@localhost:5432/testdb?sslmode=disable table: foo columns: [ "*" ] ``` ### [](#generate-100-rows)Generate 100 Rows The generate input can be used as a convenient way to generate test data. The following example generates 100 rows of structured data by setting an explicit count. The interval field is set to empty, which means data is generated as fast as the downstream components can consume it. ```yaml input: generate: count: 100 interval: "" mapping: | root = if random_int() % 2 == 0 { { "type": "foo", "foo": "is yummy" } } else { { "type": "bar", "bar": "is gross" } } ``` ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batch_size)`batch_size` The number of generated messages that should be accumulated into each batch flushed at the specified interval. **Type**: `int` **Default**: `1` ### [](#count)`count` An optional number of messages to generate, if set above 0 the specified number of messages is generated and then the input will shut down. **Type**: `int` **Default**: `0` ### [](#interval)`interval` The time interval at which messages should be generated, expressed either as a duration string or as a cron expression. If set to an empty string messages will be generated as fast as downstream services can process them. Cron expressions can specify a timezone by prefixing the expression with `TZ=`, where the location name corresponds to a file within the IANA Time Zone database. **Type**: `string` **Default**: `1s` ```yaml # Examples: interval: 5s # --- interval: 1m # --- interval: 1h # --- interval: @every 1s # --- interval: 0,30 */2 * * * * # --- interval: TZ=Europe/London 30 3-6,20-23 * * * ``` ### [](#mapping)`mapping` A [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) mapping to use for generating messages. **Type**: `string` ```yaml # Examples: mapping: root = "hello world" # --- mapping: root = {"test":"message","id":uuid_v4()} ``` --- # Page 57: git **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/git.md --- # git > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: git latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/git page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/git.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/git.adoc page-git-created-date: "2025-05-02" page-git-modified-date: "2026-05-26" --- Clones a Git repository, reads its contents, then polls for new commits at a configurable interval. Any updates are emitted as new messages. ```yml inputs: label: "" git: repository_url: "" # No default (required) branch: main poll_interval: 10s include_patterns: [] exclude_patterns: [] max_file_size: 10485760 checkpoint_cache: "" # No default (optional) checkpoint_key: git_last_commit auth: basic: username: "" password: "" ssh_key: private_key_path: "" private_key: "" passphrase: "" token: value: "" auto_replay_nacks: true ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `git_file_path` - `git_file_size` - `git_file_mode` - `git_file_modified` - `git_commit` - `git_mime_type` - `git_is_binary` - `git_encoding` (present if the file was base64 encoded) - `git_deleted` (only present if the file was deleted) You can access these metadata fields using function interpolation. ## [](#fields)Fields ### [](#auth)`auth` Options for authenticating with your Git repository. **Type**: `object` ### [](#auth-basic)`auth.basic` Allows you to specify basic authentication. **Type**: `object` ### [](#auth-basic-password)`auth.basic.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#auth-basic-username)`auth.basic.username` The username to use for authentication. **Type**: `string` **Default**: `""` ### [](#auth-ssh_key)`auth.ssh_key` Allows you to specify SSH key authentication. **Type**: `object` ### [](#auth-ssh_key-passphrase)`auth.ssh_key.passphrase` The passphrase for your SSH private key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#auth-ssh_key-private_key)`auth.ssh_key.private_key` Your private SSH key. When using encrypted keys, you must also set a value for [`private_key_passphrase`](#auth-ssh_key-passphrase). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#auth-ssh_key-private_key_path)`auth.ssh_key.private_key_path` The path to your private SSH key file. When using encrypted keys, you must also set a value for [`private_key_passphrase`](#auth-ssh_key-passphrase). **Type**: `string` **Default**: `""` ### [](#auth-token)`auth.token` Allows you to specify token-based authentication. **Type**: `object` ### [](#auth-token-value)`auth.token.value` The token value to use for token-based authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#branch)`branch` The repository branch to check out. **Type**: `string` **Default**: `main` ### [](#checkpoint_cache)`checkpoint_cache` Specify a [`cache`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) resource to store the last processed commit hash. After a restart, Redpanda Connect can then continue processing changes from where it left off, avoiding the need to reprocess all detected updates. **Type**: `string` ### [](#checkpoint_key)`checkpoint_key` The key to use when storing the last processed commit hash in the cache. **Type**: `string` **Default**: `git_last_commit` ### [](#exclude_patterns)`exclude_patterns[]` A list of file patterns to exclude. For example, you could choose not to read content from certain Git directories or image files: `'.git/**', '**/*.png'`. These patterns take precedence over `include_patterns`. The following patterns are supported: - Glob patterns: **, `/`**`*/`, `?` - Character ranges: `[a-z]`. Escape any character with a special meaning using a backslash. **Type**: `array` **Default**: `[]` ### [](#include_patterns)`include_patterns[]` A list of file patterns to read from. For example, you could read content from only Markdown and YAML files: `'***/**.md', 'configs/*.yaml'`. The following patterns are supported: - Glob patterns: **, `/`**`*/`, `?` - Character ranges: `[a-z]`. Escape any character with a special meaning using a backslash. If this field is left empty, all files are read from. **Type**: `array` **Default**: `[]` ### [](#max_file_size)`max_file_size` The maximum size of files to read from (in bytes). Files that exceed this limit are skipped. Set to `0` for unlimited file sizes. **Type**: `int` **Default**: `10485760` ### [](#poll_interval)`poll_interval` How frequently this input polls the Git repository for changes. **Type**: `string` **Default**: `10s` ```yaml # Examples: poll_interval: 10s ``` ### [](#repository_url)`repository_url` The URL of the Git repository to clone. **Type**: `string` ```yaml # Examples: repository_url: https://github.com/username/repo.git ``` --- # Page 58: http_client **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/http_client.md --- # http_client > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: http_client latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/http_client page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/http_client.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/http_client.adoc page-git-created-date: "2025-03-04" page-git-modified-date: "2026-05-26" --- Connects to a server and continuously requests single messages. #### Common ```yml inputs: label: "" http_client: url: "" # No default (required) verb: GET headers: {} rate_limit: "" # No default (optional) timeout: 5s payload: "" # No default (optional) stream: enabled: false reconnect: true scanner: lines: {} auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" http_client: url: "" # No default (required) verb: GET headers: {} metadata: include_prefixes: [] include_patterns: [] dump_request_log_level: "" oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] extract_headers: include_prefixes: [] include_patterns: [] rate_limit: "" # No default (optional) timeout: 5s retry_period: 1s max_retry_backoff: 300s retries: 3 follow_redirects: true backoff_on: - 429 drop_on: [] successful_on: [] proxy_url: "" # No default (optional) disable_http2: false payload: "" # No default (optional) drop_empty_bodies: true stream: enabled: false reconnect: true scanner: lines: {} auto_replay_nacks: true ``` ## [](#dynamic-url-and-header-settings)Dynamic URL and header settings You can set the [`url`](#url) and [`headers`](#headers) values dynamically using [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). You can also add [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) to the [`url`](#url) and [`headers`](#headers) fields to implement basic pagination, such as page numbers or tokens, where subsequent requests need to include data from previously-consumed responses. Example: ```yaml input: http_client: url: >- https://api.example.com/search?query=allmyfoos&start_time=${! ( (timestamp_unix()-300).ts_format("2006-01-02T15:04:05Z","UTC").escape_url_query() ) }${! ("&next_token="+this.meta.next_token.not_null()) | "" } verb: GET rate_limit: schedule_searches oauth2: enabled: true token_url: https://api.example.com/oauth2/token client_key: "${EXAMPLE_KEY}" client_secret: "${EXAMPLE_SECRET}" rate_limit_resources: - label: schedule_searches local: count: 1 interval: 30s ``` > 💡 **TIP** > > If pagination requires more complex logic, consider using the [`http` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/http/) combined with a [`generate` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/generate/), which allows you to schedule the processor. ## [](#streaming-messages)Streaming messages If you [enable streaming](#stream-enabled), Redpanda Connect consumes the body of the server response as a continuous stream of data, and breaks the stream down into smaller, logical messages using the [specified scanner](#stream-scanner). This functionality allows you to consume APIs that provide long-lived streamed data feeds, such as stock market feeds. ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay rejected messages (negative acknowledgements) at the output level. If the cause of rejections persists, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#backoff_on)`backoff_on[]` A list of status codes that indicate a request failure, and trigger retries with an increasing backoff period between attempts. **Type**: `array` **Default**: ```yaml - 429 ``` ### [](#basic_auth)`basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#disable_http2)`disable_http2` Whether to disable HTTP/2. By default, HTTP/2 is enabled. **Type**: `bool` **Default**: `false` ### [](#drop_empty_bodies)`drop_empty_bodies` Whether to drop empty payloads received from the target server. **Type**: `bool` **Default**: `true` ### [](#drop_on)`drop_on[]` A list of status codes that indicate a request failure, where the input should not attempt retries. This helps avoid unnecessary retries for requests that are unlikely to succeed. > 📝 **NOTE** > > In these cases, the _request_ is dropped, but the _message_ that triggered the request is retained. **Type**: `array` **Default**: `[]` ### [](#dump_request_log_level)`dump_request_log_level` EXPERIMENTAL: Set the logging level for the request and response payloads of each HTTP request. **Type**: `string` **Default**: `""` **Options**: `TRACE`, `DEBUG`, `INFO`, `WARN`, `ERROR`, `FATAL`, \`\` ### [](#extract_headers)`extract_headers` Specify which response headers to add to the resulting messages as metadata. Header keys are automatically converted to lowercase before matching, so make sure that your patterns target the lowercase versions of the expected header keys. **Type**: `object` ### [](#extract_headers-include_patterns)`extract_headers.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#extract_headers-include_prefixes)`extract_headers.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#follow_redirects)`follow_redirects` Whether or not to transparently follow redirects, i.e. responses with 300-399 status codes. If disabled, the response message will contain the body, status, and headers from the redirect response and the processor will not make a request to the URL set in the Location header of the response. **Type**: `bool` **Default**: `true` ### [](#headers)`headers` A map of headers to add to the request. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: Content-Type: application/octet-stream traceparent: ${! tracing_span().traceparent } ``` ### [](#jwt)`jwt` (beta) Configure JSON Web Token (JWT) authentication. This feature is in beta and may change in future releases. JWT tokens provide secure, stateless authentication between services. **Type**: `object` ### [](#jwt-claims)`jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` Additional key-value pairs to include in the JWT header (optional). These headers provide extra metadata for JWT processing. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` Path to a file containing the PEM-encoded private key using PKCS#1 or PKCS#8 format. The private key must be compatible with the algorithm specified in the `signing_method` field. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` The cryptographic algorithm used to sign the JWT token. Supported algorithms include RS256, RS384, RS512, and EdDSA. This algorithm must be compatible with the private key specified in the `private_key_file` field. **Type**: `string` **Default**: `""` ### [](#max_retry_backoff)`max_retry_backoff` The maximum period to wait between failed requests. **Type**: `string` **Default**: `300s` ### [](#metadata)`metadata` Specify matching rules that determine which metadata keys to add to the HTTP request as headers (optional). **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#oauth)`oauth` Configure OAuth version 1.0 authentication for secure API access. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` The value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` The secret that establishes ownership of the `oauth.access_token` in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` The secret that establishes ownership of the consumer key in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2)`oauth2` Allows you to specify open authentication using OAuth version 2 and the client credentials token flow. **Type**: `object` ### [](#oauth2-client_key)`oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#oauth2-client_secret)`oauth2.client_secret` The secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth2-enabled)`oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2-endpoint_params)`oauth2.endpoint_params` A list of endpoint parameters specified as arrays of strings (optional). **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: bar: - woof foo: - meow - quack ``` ### [](#oauth2-scopes)`oauth2.scopes[]` A list of requested permissions (optional). **Type**: `array` **Default**: `[]` ### [](#oauth2-token_url)`oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#payload)`payload` A payload to deliver for each request (optional). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#proxy_url)`proxy_url` A HTTP proxy URL (optional). **Type**: `string` ### [](#rate_limit)`rate_limit` A [rate limit](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) to throttle requests by (optional). **Type**: `string` ### [](#retries)`retries` The maximum number of retry attempts to make. **Type**: `int` **Default**: `3` ### [](#retry_period)`retry_period` The initial period to wait between failed requests before retrying. **Type**: `string` **Default**: `1s` ### [](#stream)`stream` Enables streaming mode, where the HTTP connection remains open and messages are processed line-by-line. **Type**: `object` ### [](#stream-enabled)`stream.enabled` Enables streaming mode. **Type**: `bool` **Default**: `false` ### [](#stream-reconnect)`stream.reconnect` Whether to automatically reestablish the HTTP connection if it is lost. **Type**: `bool` **Default**: `true` ### [](#stream-scanner)`stream.scanner` The [scanner](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/about/) used to split the stream of bytes into individual messages. Scanners are useful for processing large data sources efficiently without holding the entire data set in memory. For example, the `csv` scanner processes individual rows in a CSV file without loading the entire file in memory. **Type**: `scanner` **Default**: ```yaml lines: {} ``` ### [](#successful_on)`successful_on[]` A list of HTTP status codes that should be considered as successful, even if they are not 2XX codes. This is useful for handling cases where non-2XX codes indicate that the request was processed successfully, such as `303 See Other` or `409 Conflict`. By default, all 2XX codes are considered successful unless they are specified in `backoff_on` or `drop_on` fields. **Type**: `array` **Default**: `[]` ### [](#timeout)`timeout` A static timeout to apply to requests. **Type**: `string` **Default**: `5s` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL to connect to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#verb)`verb` A verb to connect with. **Type**: `string` **Default**: `GET` ```yaml # Examples: verb: POST # --- verb: GET # --- verb: DELETE ``` --- # Page 59: http_server **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/http_server.md --- # http_server > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: http_server latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/http_server page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/http_server.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/http_server.adoc description: Receive messages POSTed over HTTP(S). HTTP 2.0 is supported when using TLS, which is enabled when key and cert files are specified. page-git-created-date: "2026-02-18" page-git-modified-date: "2026-05-26" --- Receive messages sent over HTTP using POST requests. HTTP 2.0 is supported when using TLS, which is enabled when key and cert files are specified. #### Common ```yml inputs: label: "" http_server: address: "" path: /post ws_path: /post/ws allowed_verbs: - "POST" timeout: 5s rate_limit: "" ``` #### Advanced ```yml inputs: label: "" http_server: address: "" path: /post ws_path: /post/ws ws_welcome_message: "" ws_rate_limit_message: "" allowed_verbs: - "POST" timeout: 5s rate_limit: "" cert_file: "" key_file: "" cors: enabled: false allowed_origins: [] sync_response: status: 200 headers: Content-Type: "application/octet-stream" metadata_headers: include_prefixes: [] include_patterns: [] tcp: reuse_addr: false reuse_port: false ``` The field `rate_limit` allows you to specify an optional [`rate_limit` resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/), which will be applied to each HTTP request made and each websocket payload received. When the rate limit is breached HTTP requests will have a 429 response returned with a Retry-After header. Websocket payloads will be dropped and an optional response payload will be sent as per `ws_rate_limit_message`. ## [](#responses)Responses It’s possible to return a response for each message received using [synchronous responses](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/sync_responses/). When doing so you can customize headers with the `sync_response` field `headers`, which can also use [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) in the value based on the response message contents. ## [](#endpoints)Endpoints The following fields specify endpoints that are registered for sending messages, and support path parameters of the form `/{foo}`, which are added to ingested messages as metadata. A path ending in `/` will match against all extensions of that path: ### [](#path-defaults-to-post)`path` (defaults to `/post`) This endpoint expects POST requests where the entire request body is consumed as a single message. If the request contains a multipart `content-type` header as per [RFC1341](https://www.w3.org/Protocols/rfc1341/7_2_Multipart.html) then the multiple parts are consumed as a batch of messages, where each body part is a message of the batch. ### [](#ws_path-defaults-to-postws)`ws_path` (defaults to `/post/ws`) Creates a websocket connection, where payloads received on the socket are passed through the pipeline as a batch of one message. > ⚠️ **CAUTION: Endpoint caveats** > > Endpoint caveats > > Components within a Redpanda Connect config will register their respective endpoints in a non-deterministic order. This means that establishing precedence of endpoints that are registered via multiple `http_server` inputs or outputs (either within brokers or from cohabiting streams) is not possible in a predictable way. > > This ambiguity makes it difficult to ensure that paths which are both a subset of a path registered by a separate component, and end in a slash (`/`) and will therefore match against all extensions of that path, do not prevent the more specific path from matching against requests. > > It is therefore recommended that you ensure paths of separate components do not collide unless they are explicitly non-competing. > > For example, if you were to deploy two separate `http_server` inputs, one with a path `/foo/` and the other with a path `/foo/bar`, it would not be possible to ensure that the path `/foo/` does not swallow requests made to `/foo/bar`. You may specify an optional `ws_welcome_message`, which is a static payload to be sent to all clients once a websocket connection is first established. It’s also possible to specify a `ws_rate_limit_message`, which is a static payload to be sent to clients that have triggered the servers rate limit. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `http_server_user_agent` - `http_server_request_path` - `http_server_verb` - `http_server_remote_ip` - All headers (only first values are taken) - All query parameters - All path parameters - All cookies If HTTPS is enabled, the following fields are added as well: - `http_server_tls_version` - `http_server_tls_subject` - `http_server_tls_cipher_suite` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#examples)Examples ### [](#path-switching)Path Switching This example shows an `http_server` input that captures all requests and processes them by switching on that path: ```yaml input: http_server: path: / allowed_verbs: [ GET, POST ] sync_response: headers: Content-Type: application/json processors: - switch: - check: '@http_server_request_path == "/foo"' processors: - mapping: | root.title = "You Got Fooed!" root.result = content().string().uppercase() - check: '@http_server_request_path == "/bar"' processors: - mapping: 'root.title = "Bar Is Slow"' - sleep: # Simulate a slow endpoint duration: 1s ``` ### [](#mock-oauth-2-0-server)Mock OAuth 2.0 Server This example shows an `http_server` input that mocks an OAuth 2.0 Client Credentials flow server at the endpoint `/oauth2_test`: ```yaml input: http_server: path: /oauth2_test allowed_verbs: [ GET, POST ] sync_response: headers: Content-Type: application/json processors: - log: message: "Received request" level: INFO fields_mapping: | root = @ root.body = content().string() - mapping: | root.access_token = "MTQ0NjJkZmQ5OTM2NDE1ZTZjNGZmZjI3" root.token_type = "Bearer" root.expires_in = 3600 - sync_response: {} - mapping: 'root = deleted()' ``` ## [](#fields)Fields ### [](#address)`address` An alternative address to host from. If left empty the service wide address is used. **Type**: `string` **Default**: `""` ### [](#allowed_verbs)`allowed_verbs[]` An array of verbs that are allowed for the `path` endpoint. **Type**: `array` **Default**: ```yaml - "POST" ``` ### [](#cert_file)`cert_file` Enable TLS by specifying a certificate and key file. Only valid with a custom `address`. **Type**: `string` **Default**: `""` ### [](#cors)`cors` Adds Cross-Origin Resource Sharing headers. Only valid with a custom `address`. **Type**: `object` ### [](#cors-allowed_origins)`cors.allowed_origins[]` An explicit list of origins that are allowed for CORS requests. **Type**: `array` **Default**: `[]` ### [](#cors-enabled)`cors.enabled` Whether to allow CORS requests. **Type**: `bool` **Default**: `false` ### [](#key_file)`key_file` Enable TLS by specifying a certificate and key file. Only valid with a custom `address`. **Type**: `string` **Default**: `""` ### [](#path)`path` The endpoint path to listen for POST requests. **Type**: `string` **Default**: `/post` ### [](#rate_limit)`rate_limit` An optional [rate limit](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) to throttle requests by. **Type**: `string` **Default**: `""` ### [](#sync_response)`sync_response` Customize messages returned via [synchronous responses](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/sync_responses/). **Type**: `object` ### [](#sync_response-headers)`sync_response.headers` Specify headers to return with synchronous responses. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: ```yaml Content-Type: "application/octet-stream" ``` ### [](#sync_response-metadata_headers)`sync_response.metadata_headers` Specify criteria for which metadata values are added to the response as headers. **Type**: `object` ### [](#sync_response-metadata_headers-include_patterns)`sync_response.metadata_headers.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#sync_response-metadata_headers-include_prefixes)`sync_response.metadata_headers.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#sync_response-status)`sync_response.status` Specify the status code to return with synchronous responses. This is a string value, which allows you to customize it based on resulting payloads and their metadata. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `200` ```yaml # Examples: status: ${! json("status") } # --- status: ${! meta("status") } ``` ### [](#tcp)`tcp` TCP listener configuration for the HTTP server. Only valid with a custom `address`. **Type**: `object` ### [](#tcp-reuse_addr)`tcp.reuse_addr` Enable SO\_REUSEADDR, allowing binding to ports in TIME\_WAIT state. Useful for graceful restarts and config reloads where the server needs to rebind to the same port immediately after shutdown. **Type**: `bool` **Default**: `false` ### [](#tcp-reuse_port)`tcp.reuse_port` Enable SO\_REUSEPORT, allowing multiple sockets to bind to the same port for load balancing across multiple processes/threads. **Type**: `bool` **Default**: `false` ### [](#timeout)`timeout` Timeout for requests. If a consumed messages takes longer than this to be delivered the connection is closed, but the message may still be delivered. **Type**: `string` **Default**: `5s` ### [](#ws_path)`ws_path` The endpoint path to create websocket connections from. **Type**: `string` **Default**: `/post/ws` ### [](#ws_rate_limit_message)`ws_rate_limit_message` An optional message to delivery to websocket connections that are rate limited. **Type**: `string` **Default**: `""` ### [](#ws_welcome_message)`ws_welcome_message` An optional message to deliver to fresh websocket connections. **Type**: `string` **Default**: `""` --- # Page 60: inproc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/inproc.md --- # inproc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: inproc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/inproc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/inproc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/inproc.adoc description: Directly connects to an output within the same Redpanda Connect process by a chosen ID, to link isolated streams when running in streams mode. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- ```yml inputs: label: "" inproc: "" ``` Directly connect to an output within a Redpanda Connect process by referencing it by a chosen ID. It is possible to connect multiple inputs to the same inproc ID, resulting in messages dispatching in a round-robin fashion to connected inputs. However, only one output can assume an inproc ID, and will replace existing outputs if a collision occurs. --- # Page 61: jira **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/jira.md --- # jira > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: jira latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/jira page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/jira.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/jira.adoc description: Streams Jira issues, comments, or changelog entries via JQL with incremental polling. page-git-created-date: "2026-07-23" page-git-modified-date: "2026-08-11" --- Streams Jira issues, comments, or changelog entries via JQL with incremental polling. Periodically queries Jira’s REST API using a JQL filter and emits one message per resource. The cursor (max issue `updated` timestamp, plus the set of issue versions already emitted at the boundary) is persisted via the configured cache resource after every fully-acknowledged page, so progress survives restarts — including mid-backfill — and boundary issues are not re-emitted on every poll. Authentication uses API token (email + token) basic auth. The `backoff` settings govern the adaptive backoff applied to 429 responses; retries of 502/503/504 responses use a fixed three-attempt policy. Each message body is the raw JSON of the resource. Metadata fields: - `jira_id` - issue key (issues) / comment ID / changelog history ID - `jira_issue_key` - parent issue key (omitted for resource=issues) - `jira_project` - project key - `jira_updated` - RFC 3339 timestamp of the resource - `jira_event_type` - "issue" / "comment" / "changelog" - `jira_self` - Jira API URL of the resource Limitations (v1): OAuth and the worklogs resource are not yet supported. For resource=comments and resource=changelog, only the first page of child resources (up to ~50 comments or ~100 changelog entries per issue update) is emitted; a WARN is logged when truncation is detected. To fetch the full child set for issues that exceed this limit, query the Jira REST API directly. #### Common ```yml inputs: label: "" jira: auth: email: "" # No default (required) api_token: "" # No default (required) resource: issues jql: "" fields: - "*all" expand: [] page_size: 50 poll_interval: 60s cursor: cache: "" # No default (required) key: "" overlap: 60s auto_replay_nacks: true base_url: "" # No default (required) timeout: 5s ``` #### Advanced ```yml inputs: label: "" jira: auth: email: "" # No default (required) api_token: "" # No default (required) resource: issues jql: "" fields: - "*all" expand: [] page_size: 50 poll_interval: 60s cursor: cache: "" # No default (required) key: "" overlap: 60s auto_replay_nacks: true base_url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] proxy_url: "" disable_http2: false tps_limit: 0 tps_burst: 1 backoff: initial_interval: 1s max_interval: 30s max_retries: 3 tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s http: max_idle_conns: 100 max_idle_conns_per_host: 0 max_conns_per_host: 64 idle_conn_timeout: 1m30s tls_handshake_timeout: 10s expect_continue_timeout: 1s response_header_timeout: 0s disable_keep_alives: false disable_compression: false max_response_header_bytes: 1048576 max_response_body_bytes: 10485760 write_buffer_size: 4096 read_buffer_size: 4096 h2: strict_max_concurrent_requests: false max_decoder_header_table_size: 4096 max_encoder_header_table_size: 4096 max_read_frame_size: 16384 max_receive_buffer_per_connection: 1048576 max_receive_buffer_per_stream: 1048576 send_ping_timeout: 0s ping_timeout: 15s write_byte_timeout: 0s access_log_level: "" access_log_body_limit: 0 ``` ## [](#fields)Fields ### [](#access_log_body_limit)`access_log_body_limit` Maximum bytes of request/response body to include in logs. 0 to skip body logging. **Type**: `int` **Default**: `0` ### [](#access_log_level)`access_log_level` Log level for HTTP request/response logging. Empty disables logging. **Type**: `string` **Default**: `""` **Options**: `` `, `TRACE ``, `DEBUG`, `INFO`, `WARN`, `ERROR` ### [](#auth)`auth` API token authentication. **Type**: `object` ### [](#auth-api_token)`auth.api_token` Jira API token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-email)`auth.email` Email or username of the Jira account. **Type**: `string` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#backoff)`backoff` Adaptive backoff configuration for 429 (Too Many Requests) responses. Always active. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` Initial interval between retries on 429 responses. **Type**: `string` **Default**: `1s` ### [](#backoff-max_interval)`backoff.max_interval` Maximum interval between retries on 429 responses. **Type**: `string` **Default**: `30s` ### [](#backoff-max_retries)`backoff.max_retries` Maximum number of retries on 429 responses. **Type**: `int` **Default**: `3` ### [](#base_url)`base_url` Base URL of the target service (e.g., [https://api.example.com](https://api.example.com)). TLS is enabled automatically for https URLs. **Type**: `string` ### [](#cursor)`cursor` Cursor checkpoint storage. **Type**: `object` ### [](#cursor-cache)`cursor.cache` Name of a cache resource used to persist the cursor. **Type**: `string` ### [](#cursor-key)`cursor.key` Cache key. Defaults to `redpanda_connect_jira_input_`. **Type**: `string` **Default**: `""` ### [](#cursor-overlap)`cursor.overlap` Widens `updated >= cursor - overlap` to absorb minute-boundary effects. Jira JQL’s `updated` operator has minute precision, so this should be set to at least 1m to have an effect. **Type**: `string` **Default**: `60s` ### [](#disable_http2)`disable_http2` Disable HTTP/2 and force HTTP/1.1. **Type**: `bool` **Default**: `false` ### [](#expand)`expand[]` Jira `expand` query parameter. The input automatically adds `changelog` when resource=changelog. **Type**: `array` **Default**: `[]` ### [](#fields-2)`fields[]` Jira `fields` query parameter - narrow this for throughput. **Type**: `array` **Default**: ```yaml - "*all" ``` ### [](#http)`http` HTTP transport settings controlling connection pooling, timeouts, and HTTP/2. **Type**: `object` ### [](#http-disable_compression)`http.disable_compression` Disable automatic decompression of gzip responses. **Type**: `bool` **Default**: `false` ### [](#http-disable_keep_alives)`http.disable_keep_alives` Disable HTTP keep-alive connections; each request uses a new connection. **Type**: `bool` **Default**: `false` ### [](#http-expect_continue_timeout)`http.expect_continue_timeout` Maximum time to wait for a server’s 100-continue response before sending the body. 0 means the body is sent immediately. **Type**: `string` **Default**: `1s` ### [](#http-h2)`http.h2` HTTP/2-specific transport settings. Only applied when HTTP/2 is enabled. **Type**: `object` ### [](#http-h2-max_decoder_header_table_size)`http.h2.max_decoder_header_table_size` Upper limit in bytes for the HPACK header table used to decode headers from the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-h2-max_encoder_header_table_size)`http.h2.max_encoder_header_table_size` Upper limit in bytes for the HPACK header table used to encode headers sent to the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-h2-max_read_frame_size)`http.h2.max_read_frame_size` Largest HTTP/2 frame this endpoint will read. Valid range: 16 KiB to 16 MiB. **Type**: `int` **Default**: `16384` ### [](#http-h2-max_receive_buffer_per_connection)`http.h2.max_receive_buffer_per_connection` Maximum flow-control window size in bytes for data received on a connection. Must be at least 64 KiB and less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-h2-max_receive_buffer_per_stream)`http.h2.max_receive_buffer_per_stream` Maximum flow-control window size in bytes for data received on a single stream. Must be less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-h2-ping_timeout)`http.h2.ping_timeout` Timeout waiting for a PING response before closing the connection. **Type**: `string` **Default**: `15s` ### [](#http-h2-send_ping_timeout)`http.h2.send_ping_timeout` Idle timeout after which a PING frame is sent to verify connection health. 0 disables health checks. **Type**: `string` **Default**: `0s` ### [](#http-h2-strict_max_concurrent_requests)`http.h2.strict_max_concurrent_requests` When true, new requests block when a connection’s concurrency limit is reached instead of opening a new connection. **Type**: `bool` **Default**: `false` ### [](#http-h2-write_byte_timeout)`http.h2.write_byte_timeout` Timeout for writing data to a connection. The timer resets whenever bytes are written. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-idle_conn_timeout)`http.idle_conn_timeout` How long an idle connection remains in the pool before being closed. 0 disables the timeout. **Type**: `string` **Default**: `1m30s` ### [](#http-max_conns_per_host)`http.max_conns_per_host` Maximum total connections (active + idle) per host. 0 means unlimited. **Type**: `int` **Default**: `64` ### [](#http-max_idle_conns)`http.max_idle_conns` Maximum total number of idle (keep-alive) connections across all hosts. 0 means unlimited. **Type**: `int` **Default**: `100` ### [](#http-max_idle_conns_per_host)`http.max_idle_conns_per_host` Maximum idle connections to keep per host. 0 (the default) uses GOMAXPROCS+1. **Type**: `int` **Default**: `0` ### [](#http-max_response_body_bytes)`http.max_response_body_bytes` Maximum bytes of response body the client will read. The response body is wrapped with a limit reader; reads beyond this cap return EOF. 0 disables the limit. **Type**: `int` **Default**: `10485760` ### [](#http-max_response_header_bytes)`http.max_response_header_bytes` Maximum bytes of response headers to allow. **Type**: `int` **Default**: `1048576` ### [](#http-read_buffer_size)`http.read_buffer_size` Size in bytes of the per-connection read buffer. **Type**: `int` **Default**: `4096` ### [](#http-response_header_timeout)`http.response_header_timeout` Maximum time to wait for response headers after writing the full request. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-tls_handshake_timeout)`http.tls_handshake_timeout` Maximum time to wait for a TLS handshake to complete. 0 disables the timeout. **Type**: `string` **Default**: `10s` ### [](#http-write_buffer_size)`http.write_buffer_size` Size in bytes of the per-connection write buffer. **Type**: `int` **Default**: `4096` ### [](#jql)`jql` Jira JQL filter. The input appends an `updated >= cursor` predicate and `ORDER BY updated ASC, key ASC`. Empty means all issues visible to the principal. **Type**: `string` **Default**: `""` ### [](#page_size)`page_size` Issues per Jira page (Jira max 100). **Type**: `int` **Default**: `50` ### [](#poll_interval)`poll_interval` Time to wait between polls once the input has caught up. Minimum 10s. **Type**: `string` **Default**: `60s` ### [](#proxy_url)`proxy_url` HTTP proxy URL. Empty string disables proxying. **Type**: `string` **Default**: `""` ### [](#resource)`resource` Which Jira resource to emit. **Type**: `string` **Default**: `issues` **Options**: `issues`, `comments`, `changelog` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` HTTP request timeout. **Type**: `string` **Default**: `5s` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tps_burst)`tps_burst` Maximum burst size for rate limiting. **Type**: `int` **Default**: `1` ### [](#tps_limit)`tps_limit` Rate limit in requests per second. 0 disables rate limiting. **Type**: `float` **Default**: `0` --- # Page 62: kafka_franz **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/kafka_franz.md --- # kafka_franz > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: kafka_franz latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/kafka_franz page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/kafka_franz.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/kafka_franz.adoc description: A Kafka input using the Franz Kafka client library. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > ⚠️ **WARNING: Deprecated in 4.68.0** > > Deprecated in 4.68.0 > > This component is deprecated and will be removed in the next major version release. Please consider moving onto the unified [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [`redpanda` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) components. A Kafka input using the [Franz Kafka client library](https://github.com/twmb/franz-go). #### Common ```yml inputs: label: "" kafka_franz: seed_brokers: [] # No default (required) topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" kafka_franz: seed_brokers: [] # No default (required) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) rack_id: "" instance_id: "" rebalance_timeout: 45s session_timeout: 1m heartbeat_interval: 3s start_offset: earliest fetch_max_bytes: 50MiB fetch_max_wait: 5s fetch_min_bytes: 1B fetch_max_partition_bytes: 1MiB transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) checkpoint_limit: 1024 commit_period: 5s multi_header: false batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) topic_lag_refresh_period: 5s auto_replay_nacks: true timely_nacks_maximum_wait: "" # No default (optional) ``` When you specify a consumer group in your configuration, this input consumes one or more topics and automatically balances the topic partitions across any other connected clients with the same consumer group. Otherwise, topics are consumed in their entirety or with explicit partitions. This input often out-performs the traditional `kafka` input and provides more useful logs and error messages. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `kafka_key` - `kafka_topic` - `kafka_partition` - `kafka_offset` - `kafka_lag` - `kafka_timestamp_ms` - `kafka_timestamp_unix` - `kafka_tombstone_message` - All record headers ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay rejected messages (negative acknowledgements) at the output level. If the cause of rejections persists, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/) that applies to individual topic partitions in order to batch messages together before flushing them for processing. Batching can be beneficial for performance as well as useful for windowed processing, and doing so this way preserves the ordering of topic partitions. **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_limit)`checkpoint_limit` The maximum number of messages that are processed in parallel inside the same partition before back pressure is applied. When a message with a specific offset is delivered to the output, the offset is only committed when all messages of previous offsets have also been delivered. This behavior ensures at-least-once delivery guarantees. However, in the event of crashes or server faults, it also increases the likelihood of duplicates. To decrease this risk, reduce the `checkpoint_limit` value. **Type**: `int` **Default**: `1024` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#commit_period)`commit_period` The period of time between each commit of the current partition offsets. Offsets are always committed during shutdown. **Type**: `string` **Default**: `5s` ### [](#conn_idle_timeout)`conn_idle_timeout` The maximum duration that connections can remain idle before they are automatically closed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `20s` ### [](#consumer_group)`consumer_group` An optional consumer group. When you specify this value: - The partitions of any topics, specified in the `topics` field, are automatically distributed across consumers sharing a consumer group - Partition offsets are automatically committed and resumed under this name Consumer groups are not supported when you specify explicit partitions to consume from in the `topics` field. **Type**: `string` ### [](#fetch_max_bytes)`fetch_max_bytes` The maximum size of a message batch (in bytes) that a broker tries to send during a client fetch. If individual records exceed the `fetch_max_bytes` value, brokers will still send them. **Type**: `string` **Default**: `50MiB` ### [](#fetch_max_partition_bytes)`fetch_max_partition_bytes` The maximum number of bytes that are consumed from a single partition in a fetch request. This field is equivalent to the Java setting `fetch.max.partition.bytes`. If a single batch is larger than the `fetch_max_partition_bytes` value, the batch is still sent so that the client can make progress. **Type**: `string` **Default**: `1MiB` ### [](#fetch_max_wait)`fetch_max_wait` The maximum period of time a broker can wait for a fetch response to reach the required minimum number of bytes (`fetch_min_bytes`). **Type**: `string` **Default**: `5s` ### [](#fetch_min_bytes)`fetch_min_bytes` The minimum number of bytes that a broker tries to send during a fetch. This field is equivalent to the Java setting `fetch.min.bytes`. **Type**: `string` **Default**: `1B` ### [](#heartbeat_interval)`heartbeat_interval` When you specify a `consumer_group`, `heartbeat_interval` sets how frequently a consumer group member should send heartbeats to Apache Kafka. Apache Kafka uses heartbeats to make sure that a group member’s session is active. You must set `heartbeat_interval` to less than one-third of `session_timeout`. This field is equivalent to the Java `heartbeat.interval.ms` setting and accepts Go duration format strings such as `10s` or `2m`. **Type**: `string` **Default**: `3s` ### [](#instance_id)`instance_id` When you specify a [`consumer_group`](#consumer_group), assign a unique value to `instance_id` to define the group’s static membership, which can prevent unnecessary rebalances during reconnections. When you assign an instance ID, the client does not automatically leave the consumer group when it disconnects. To remove the client, you must use an external admin command on behalf of the instance ID. **Type**: `string` **Default**: `""` ### [](#metadata_max_age)`metadata_max_age` The maximum period of time after which metadata is refreshed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. Lower values provide more responsive topic and partition discovery but may increase broker load. Higher values reduce broker queries but can delay detection of topology changes. **Type**: `string` **Default**: `1m` ### [](#multi_header)`multi_header` Decode headers into lists to allow the handling of multiple values with the same key. **Type**: `bool` **Default**: `false` ### [](#rack_id)`rack_id` A rack specifies where the client is physically located, and changes fetch requests to consume from the closest replica as opposed to the leader replica. **Type**: `string` **Default**: `""` ### [](#rebalance_timeout)`rebalance_timeout` When you specify a [`consumer_group`](#consumer_group), `rebalance_timeout` sets a time limit for all consumer group members to complete their work and commit offsets after a rebalance has begun. The timeout excludes the time taken to detect a failed or late heartbeat, which indicates a rebalance is required. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `45s` ### [](#regexp_topics_exclude)`regexp_topics_exclude[]` A list of regular expression patterns for excluding topics when regex mode is enabled (using `regexp_topics_include` or the deprecated `regexp_topics` boolean). Topics matching any of these patterns will be excluded from consumption, even if they match include patterns. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so use `^` and `$` for exact matching. Exclude patterns are applied after include patterns, providing fine-grained control over topic selection. Example: `regexp_topics_exclude: ["^_", ".**-temp$", ".**-test.*"]` excludes topics starting with underscore, ending with `-temp`, or containing `-test`. **Type**: `array` ### [](#regexp_topics_include)`regexp_topics_include[]` A list of regular expression patterns for matching topics to consume from. When specified, the client will periodically refresh the list of matching topics based on the `metadata_max_age` interval. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so `logs_.` **matches `my-logs_events` and `logs_errors`. Use `^logs_.`**`$` to match only topics starting with `logs_`. This field enables regex mode (replacing the deprecated `regexp_topics` boolean) and cannot be used together with explicit `topics` lists. Use `regexp_topics_exclude` to filter out specific patterns from the matched topics. Example: `regexp_topics_include: ["events_.**", "logs_.**"]` consumes from all topics starting with `events_` or `logs_`. **Type**: `array` ```yaml # Examples: regexp_topics_include: - logs_.* - metrics_.* # --- regexp_topics_include: - "events_[0-9]+" ``` ### [](#request_timeout_overhead)`request_timeout_overhead` Grants an additional buffer or overhead to requests that have timeout fields defined. This field is based on the behavior of Apache Kafka’s `request.timeout.ms` parameter. **Type**: `string` **Default**: `10s` ### [](#sasl)`sasl[]` Specify one or more methods or mechanisms of SASL authentication, which are attempted in order. If the broker supports the first SASL mechanism, all connections use it. If the first mechanism fails, the client picks the first supported mechanism. If the broker does not support any client mechanisms, all connections fail. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM based authentication as specified in RFC5802. | | none | Disable sasl authentication | ### [](#sasl-password)`sasl[].password` A password to provide for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` A username to provide for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to in order. Use commas to separate multiple addresses in a single list item. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#session_timeout)`session_timeout` When you specify a `consumer_group`, `session_timeout` sets the maximum interval between heartbeats sent by a consumer group member to the broker. If a broker doesn’t receive a heartbeat from a group member before the timeout expires, it removes the member from the consumer group and initiates a rebalance. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1m` ### [](#start_offset)`start_offset` Specify the offset from which this input starts or restarts consuming messages. Restarts occur when the `OffsetOutOfRange` error is seen during a fetch. **Type**: `string` **Default**: `earliest` | Option | Summary | | --- | --- | | committed | Prevents consuming a partition in a group if the partition has no prior commits. Corresponds to Kafka’s auto.offset.reset=none option | | earliest | Start from the earliest offset. Corresponds to Kafka’s auto.offset.reset=earliest option. | | latest | Start from the latest offset. Corresponds to Kafka’s auto.offset.reset=latest option. | ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timely_nacks_maximum_wait)`timely_nacks_maximum_wait` EXPERIMENTAL: Specify a maximum period of time in which each message can be consumed and awaiting either acknowledgement or rejection before rejection is instead forced. This can be useful for avoiding situations where certain downstream components can result in blocked confirmation of delivery that exceeds SLAs. Accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic_lag_refresh_period)`topic_lag_refresh_period` The interval between refresh cycles. During each cycle, this input queries the Redpanda Connect server to calculate the topic lag minus the number of produced messages that remain to be read from each topic/partition pair by the specified consumer group. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `5s` ### [](#topics)`topics[]` A list of topics to consume from. Use commas to separate multiple topics in a single element. When a `consumer_group` is specified, partitions are automatically distributed across consumers of a topic. Otherwise, all partitions are consumed. Alternatively, you can specify explicit partitions to consume by using a colon after the topic name. For example, `foo:0` would consume the partition `0` of the topic foo. This syntax supports ranges. For example, `foo:0-10` would consume partitions `0` through to `10` inclusive. It is also possible to specify an explicit offset to consume from by adding another colon after the partition. For example, `foo:0:10` would consume the partition `0` of the topic `foo` starting from the offset `10`. If the offset is not present (or remains unspecified) then the field `start_offset` determines which offset to start from. **Type**: `array` ```yaml # Examples: topics: - foo - bar # --- topics: - things.* # --- topics: - "foo,bar" # --- topics: - "foo:0" - "bar:1" - "bar:3" # --- topics: - "foo:0,bar:1,bar:3" # --- topics: - "foo:0-5" ``` ### [](#transaction_isolation_level)`transaction_isolation_level` The isolation level for handling transactional messages. This setting determines how transactions are processed and affects data consistency guarantees. **Type**: `string` **Default**: `read_uncommitted` | Option | Summary | | --- | --- | | read_committed | If set, only committed transactional records are processed. | | read_uncommitted | If set, then uncommitted records are processed. | --- # Page 63: kafka **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/kafka.md --- # kafka > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: kafka latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/kafka page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/kafka.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/kafka.adoc description: Connects to Kafka brokers and consumes one or more topics. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > ⚠️ **WARNING: Deprecated in 4.68.0** > > Deprecated in 4.68.0 > > This component is deprecated and will be removed in the next major version release. Please consider moving onto the unified [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [`redpanda` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) components. Connects to Kafka brokers and consumes one or more topics. #### Common ```yml inputs: label: "" kafka: addresses: [] # No default (required) topics: [] # No default (required) target_version: "" # No default (optional) consumer_group: "" checkpoint_limit: 1024 auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" kafka: addresses: [] # No default (required) topics: [] # No default (required) target_version: "" # No default (optional) tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: mechanism: none user: "" password: "" access_token: "" token_cache: "" token_key: "" consumer_group: "" client_id: benthos instance_id: "" # No default (optional) rack_id: "" start_from_oldest: true checkpoint_limit: 1024 auto_replay_nacks: true timely_nacks_maximum_wait: "" # No default (optional) commit_period: 1s max_processing_period: 100ms extract_tracing_map: "" # No default (optional) group: session_timeout: 10s heartbeat_interval: 3s rebalance_timeout: 60s fetch_buffer_cap: 256 multi_header: false batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` Offsets are managed within Kafka under the specified consumer group, and partitions for each topic are automatically balanced across members of the consumer group. The Kafka input allows parallel processing of messages from different topic partitions, and messages of the same topic partition are processed with a maximum parallelism determined by the field [`checkpoint_limit`](#checkpoint_limit). To enforce ordered processing of partition messages, set the [`checkpoint_limit`](#checkpoint_limit) to `1`, which makes sure that a message is only processed after the previous message is delivered. Batching messages before processing can be enabled using the [`batching`](#batching) field, and this batching is performed per-partition such that messages of a batch will always originate from the same partition. This batching mechanism is capable of creating batches of greater size than the [`checkpoint_limit`](#checkpoint_limit), in which case the next batch will only be created upon delivery of the current one. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `kafka_key` - `kafka_topic` - `kafka_partition` - `kafka_offset` - `kafka_lag` - `kafka_timestamp_ms` - `kafka_timestamp_unix` - `kafka_tombstone_message` - All existing message headers (version 0.11+) The field `kafka_lag` is the calculated difference between the high water mark offset of the partition at the time of ingestion and the current message offset. You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#ordering)Ordering By default messages of a topic partition can be processed in parallel, up to a limit determined by the field `checkpoint_limit`. However, if strict ordered processing is required then this value must be set to 1 in order to process shard messages in lock-step. When doing so it is recommended that you perform batching at this component for performance as it will not be possible to batch lock-stepped messages at the output level. ## [](#troubleshooting)Troubleshooting If you’re seeing issues writing to or reading from Kafka with this component then it’s worth trying out the newer [`kafka_franz` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/kafka_franz/). - I’m seeing logs that report `Failed to connect to kafka: kafka: client has run out of available brokers to talk to (Is your cluster reachable?)`, but the brokers are definitely reachable. Unfortunately this error message will appear for a wide range of connection problems even when the broker endpoint can be reached. Double check your authentication configuration and also ensure that you have [enabled TLS](#tlsenabled) if applicable. ## [](#fields)Fields ### [](#addresses)`addresses[]` A list of broker addresses to connect to. If an item of the list contains commas it will be expanded into multiple addresses. **Type**: `array` ```yaml # Examples: addresses: - "localhost:9092" # --- addresses: - "localhost:9041,localhost:9042" # --- addresses: - "localhost:9041" - "localhost:9042" ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_limit)`checkpoint_limit` The maximum number of messages of the same topic and partition that can be processed at a given time. Increasing this limit enables parallel processing and batching at the output level to work on individual partitions. Any given offset will not be committed unless all messages under that offset are delivered in order to preserve at least once delivery guarantees. **Type**: `int` **Default**: `1024` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `benthos` ### [](#commit_period)`commit_period` The period of time between each commit of the current partition offsets. Offsets are always committed during shutdown. **Type**: `string` **Default**: `1s` ### [](#consumer_group)`consumer_group` An identifier for the consumer group of the connection. This field can be explicitly made empty in order to disable stored offsets for the consumed topic partitions. **Type**: `string` **Default**: `""` ### [](#extract_tracing_map)`extract_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that attempts to extract an object containing tracing propagation information, which will then be used as the root tracing span for the message. The specification of the extracted fields must match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: extract_tracing_map: root = @ # --- extract_tracing_map: root = this.meta.span ``` ### [](#fetch_buffer_cap)`fetch_buffer_cap` The maximum number of unprocessed messages to fetch at a given time. **Type**: `int` **Default**: `256` ### [](#group)`group` Tuning parameters for consumer group synchronization. **Type**: `object` ### [](#group-heartbeat_interval)`group.heartbeat_interval` A period in which heartbeats should be sent out. **Type**: `string` **Default**: `3s` ### [](#group-rebalance_timeout)`group.rebalance_timeout` A period after which rebalancing is abandoned if unresolved. **Type**: `string` **Default**: `60s` ### [](#group-session_timeout)`group.session_timeout` A period after which a consumer of the group is kicked after no heartbeats. **Type**: `string` **Default**: `10s` ### [](#instance_id)`instance_id` When you specify a [`consumer_group`](#consumer_group), assign a unique value to `instance_id` to help brokers identify each input after restarts and prevent unnecessary rebalances. **Type**: `string` ### [](#max_processing_period)`max_processing_period` A maximum estimate for the time taken to process a message, this is used for tuning consumer group synchronization. **Type**: `string` **Default**: `100ms` ### [](#multi_header)`multi_header` Decode headers into lists to allow handling of multiple values with the same key **Type**: `bool` **Default**: `false` ### [](#rack_id)`rack_id` A rack identifier for this client. **Type**: `string` **Default**: `""` ### [](#sasl)`sasl` Enables SASL authentication. **Type**: `object` ### [](#sasl-access_token)`sasl.access_token` A static OAUTHBEARER access token **Type**: `string` **Default**: `""` ### [](#sasl-mechanism)`sasl.mechanism` The SASL authentication mechanism, if left empty SASL authentication is not used. **Type**: `string` **Default**: `none` | Option | Summary | | --- | --- | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. NOTE: When using plain text auth it is extremely likely that you’ll also need to enable TLS. | | SCRAM-SHA-256 | Authentication using the SCRAM-SHA-256 mechanism. | | SCRAM-SHA-512 | Authentication using the SCRAM-SHA-512 mechanism. | | none | Default, no SASL authentication. | ### [](#sasl-password)`sasl.password` A PLAIN password. It is recommended that you use environment variables to populate this field. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: ${PASSWORD} ``` ### [](#sasl-token_cache)`sasl.token_cache` Instead of using a static `access_token` allows you to query a [`cache`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) resource to fetch OAUTHBEARER tokens from **Type**: `string` **Default**: `""` ### [](#sasl-token_key)`sasl.token_key` Required when using a `token_cache`, the key to query the cache with for tokens. **Type**: `string` **Default**: `""` ### [](#sasl-user)`sasl.user` A PLAIN username. It is recommended that you use environment variables to populate this field. **Type**: `string` **Default**: `""` ```yaml # Examples: user: ${USER} ``` ### [](#start_from_oldest)`start_from_oldest` Determines whether to consume from the oldest available offset, otherwise messages are consumed from the latest offset. The setting is applied when creating a new consumer group or the saved offset no longer exists. **Type**: `bool` **Default**: `true` ### [](#target_version)`target_version` The version of the Kafka protocol to use. This limits the capabilities used by the client and should ideally match the version of your brokers. Defaults to the oldest supported stable version. **Type**: `string` ```yaml # Examples: target_version: 2.1.0 # --- target_version: 3.1.0 ``` ### [](#timely_nacks_maximum_wait)`timely_nacks_maximum_wait` EXPERIMENTAL: Specify a maximum period of time in which each message can be consumed and awaiting either acknowledgement or rejection before rejection is instead forced. This can be useful for avoiding situations where certain downstream components can result in blocked confirmation of delivery that exceeds SLAs. Accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#topics)`topics[]` A list of topics to consume from. Multiple comma separated topics can be listed in a single element. Partitions are automatically distributed across consumers of a topic. Alternatively, it’s possible to specify explicit partitions to consume from with a colon after the topic name, e.g. `foo:0` would consume the partition 0 of the topic foo. This syntax supports ranges, e.g. `foo:0-10` would consume partitions 0 through to 10 inclusive. **Type**: `array` ```yaml # Examples: topics: - foo - bar # --- topics: - "foo,bar" # --- topics: - "foo:0" - "bar:1" - "bar:3" # --- topics: - "foo:0,bar:1,bar:3" # --- topics: - "foo:0-5" ``` --- # Page 64: microsoft_sql_server_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/microsoft_sql_server_cdc.md --- # microsoft_sql_server_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: microsoft_sql_server_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/microsoft_sql_server_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/microsoft_sql_server_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/microsoft_sql_server_cdc.adoc description: Enables Change Data Capture by consuming from Microsoft SQL Server's change tables. page-git-created-date: "2025-10-24" page-git-modified-date: "2026-08-11" --- Enables Change Data Capture by consuming from Microsoft SQL Server’s change tables. #### Common ```yaml inputs: label: "" microsoft_sql_server_cdc: connection_string: "" # No default (required) stream_snapshot: false max_parallel_snapshot_tables: 1 snapshot_max_batch_size: 1000 include: [] # No default (required) exclude: [] # No default (optional) checkpoint_cache: "" # No default (optional) checkpoint_cache_table_name: rpcn.CdcCheckpointCache checkpoint_cache_connection_string: "" # No default (optional) checkpoint_cache_key: microsoft_sql_server_cdc checkpoint_limit: 1024 stream_backoff_interval: 5s auto_replay_nacks: true batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yaml inputs: label: "" microsoft_sql_server_cdc: connection_string: "" # No default (required) stream_snapshot: false max_parallel_snapshot_tables: 1 snapshot_max_batch_size: 1000 include: [] # No default (required) exclude: [] # No default (optional) checkpoint_cache: "" # No default (optional) checkpoint_cache_table_name: rpcn.CdcCheckpointCache checkpoint_cache_connection_string: "" # No default (optional) checkpoint_cache_key: microsoft_sql_server_cdc checkpoint_limit: 1024 stream_backoff_interval: 5s auto_replay_nacks: true batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` Streams changes from a Microsoft SQL Server database for Change Data Capture (CDC). Additionally, if `stream_snapshot` is set to true, then the existing data in the database is also streamed too. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `database_schema` (The database schema for the table where the message originates from) - `schema` (The table schema in benthos common schema format, compatible with processors like parquet\_encode) - `table` (Name of the table that the message originated from) - `operation` (Type of operation that generated the message: "read", "delete", "insert", or "update\_before" and "update\_after". "read" is from messages that are read in the initial snapshot phase.) - `lsn` (the Log Sequence Number in Microsoft SQL Server) ## [](#permissions)Permissions To use the default Microsoft SQL Server cache, the user must have permissions to create tables and stored procedures. Refer to [`checkpoint_cache_table_name`](#checkpoint_cache_table_name) for additional details. ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_cache)`checkpoint_cache` A [cache resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) to store the current Log Sequence Number (LSN) position. This enables the connector to resume from the last processed position after restarts, preventing data loss and duplicate processing. The cache stores the highest LSN that has been successfully delivered downstream. **Type**: `string` ### [](#checkpoint_cache_connection_string)`checkpoint_cache_connection_string` An optional connection string for a remote Microsoft SQL Server to use for the checkpoint cache. When set, this creates the checkpoint cache table on the remote server instead of the source database. If `checkpoint_cache` is also set, that takes precedence. **Type**: `string` ```yaml # Examples: checkpoint_cache_connection_string: sqlserver://username:password@remotehost/instance?param1=value¶m2=value ``` ### [](#checkpoint_cache_key)`checkpoint_cache_key` The key to use to store the snapshot position in `checkpoint_cache`. An alternative key can be provided if multiple CDC inputs share the same cache. **Type**: `string` **Default**: `microsoft_sql_server_cdc` ### [](#checkpoint_cache_table_name)`checkpoint_cache_table_name` The multipart identifier for the checkpoint cache table name. If no `checkpoint_cache` field is specified, this input will automatically create a table and stored procedure under the `rpcn` schema to act as a checkpoint cache. This table stores the latest processed Log Sequence Number (LSN) that has been successfully delivered, allowing Redpanda Connect to resume from that point upon restart rather than reconsume the entire change table. **Type**: `string` **Default**: `rpcn.CdcCheckpointCache` ```yaml # Examples: checkpoint_cache_table_name: dbo.checkpoint_cache ``` ### [](#checkpoint_limit)`checkpoint_limit` The maximum number of messages that can be processed concurrently before applying back pressure. Higher values enable better parallelization and batching but increase memory usage. Messages are processed in LSN order, and a given LSN is only acknowledged after all previous LSNs have been successfully delivered, ensuring at-least-once guarantees. **Type**: `int` **Default**: `1024` ### [](#connection_string)`connection_string` The connection string for the Microsoft SQL Server database. Use the format `sqlserver://username:password@host/instance?param1=value¶m2=value`. For Windows Authentication, use `sqlserver://host/instance?trusted_connection=yes`. Include additional parameters like `TrustServerCertificate=true` for self-signed certificates or `encrypt=disable` to disable encryption. **Type**: `string` ```yaml # Examples: connection_string: sqlserver://username:password@host/instance?param1=value¶m2=value ``` ### [](#exclude)`exclude[]` Regular expressions for tables to exclude from CDC streaming. Use this to filter out specific tables from the include patterns. Table names should follow the `schema.table` format. Exclude patterns are applied after include patterns, allowing you to include broad patterns while excluding specific tables. **Type**: `array` ```yaml # Examples: exclude: dbo.privatetable ``` ### [](#include)`include[]` Regular expressions for tables to include in CDC streaming. Specify table names using the format `schema.table` (such as `dbo.orders`, `sales.customers`). Each pattern is treated as a regular expression, allowing wildcards and pattern matching. All specified tables must have CDC enabled in SQL Server. **Type**: `array` ```yaml # Examples: include: dbo.products ``` ### [](#max_parallel_snapshot_tables)`max_parallel_snapshot_tables` Specifies a number of tables that will be processed in parallel during the snapshot processing stage. **Type**: `int` **Default**: `1` ### [](#snapshot_max_batch_size)`snapshot_max_batch_size` The maximum number of rows to stream in a single batch during the initial snapshot phase. Larger batch sizes can improve throughput for initial data loads but may increase memory usage. This setting only applies when `stream_snapshot` is enabled. **Type**: `int` **Default**: `1000` ### [](#stream_backoff_interval)`stream_backoff_interval` The time interval to wait between polling attempts when no new CDC data is available. For low-traffic tables, increasing this value reduces database load and network traffic. Use Go duration format like `5s`, `30s`, or `1m`. Shorter intervals provide lower latency for new changes but increase server load. **Type**: `string` **Default**: `5s` ```yaml # Examples: stream_backoff_interval: 500ms # --- stream_backoff_interval: 5s # --- stream_backoff_interval: 1m ``` ### [](#stream_snapshot)`stream_snapshot` Whether to stream a snapshot of all existing data before streaming CDC changes. When enabled, the connector first queries all existing table data, then switches to streaming incremental changes from the transaction log. Set to `false` to start streaming only new changes from the current LSN position. **Type**: `bool` **Default**: `false` ```yaml # Examples: stream_snapshot: true ``` --- # Page 65: mongodb_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/mongodb_cdc.md --- # mongodb_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mongodb_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/mongodb_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/mongodb_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/mongodb_cdc.adoc page-git-created-date: "2025-03-11" page-git-modified-date: "2026-05-26" --- Streams data changes from a MongoDB replica set, using MongoDB’s [change streams](https://www.mongodb.com/docs/manual/changeStreams/) to capture data updates. #### Common ```yml inputs: label: "" mongodb_cdc: url: "" # No default (required) database: "" # No default (required) username: "" password: "" collections: [] # No default (required) checkpoint_key: mongodb_cdc_checkpoint checkpoint_cache: "" # No default (required) checkpoint_interval: 5s checkpoint_limit: 1000 read_batch_size: 1000 read_max_wait: 1s stream_snapshot: false snapshot_parallelism: 1 auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" mongodb_cdc: url: "" # No default (required) database: "" # No default (required) username: "" password: "" aws: enabled: false region: "" # No default (optional) session_duration: 1h id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) roles: [] # No default (optional) collections: [] # No default (required) checkpoint_key: mongodb_cdc_checkpoint checkpoint_cache: "" # No default (required) checkpoint_interval: 5s checkpoint_limit: 1000 checkpoint_write_timeout: 10s on_unresumable_position: fail read_batch_size: 1000 read_max_wait: 1s stream_snapshot: false snapshot_parallelism: 1 snapshot_auto_bucket_sharding: false document_mode: update_lookup json_marshal_mode: canonical app_name: benthos auto_replay_nacks: true ``` ## [](#prerequisites)Prerequisites - MongoDB version 6 or later - Network access from the cluster where your Redpanda Connect pipeline is running to the source database environment. For detailed networking information, including how to set up a VPC peering connection, see [Redpanda Cloud Networking](https://docs.redpanda.com/cloud-data-platform/networking/). - A MongoDB database running as a [replica set](https://www.mongodb.com/docs/manual/replication/#replication-in-mongodb) or in a [sharded cluster](https://www.mongodb.com/docs/manual/sharding/) using replica set [protocol version 1](https://www.mongodb.com/docs/manual/reference/replica-configuration/#rsconf.protocolVersion). - A MongoDB database using the [WiredTiger](https://www.mongodb.com/docs/manual/core/wiredtiger/#storage-wiredtiger) storage engine. ## [](#enable-connectivity-from-cloud-based-data-sources-byoc)Enable connectivity from cloud-based data sources (BYOC) To establish a secure connection between a cloud-based data source and Redpanda Connect, you must add the NAT Gateway IP address of your Redpanda cluster to the allowlist of your data source. ## [](#data-capture-method)Data capture method The `mongodb_cdc` input uses [change streams](https://www.mongodb.com/docs/manual/changeStreams/) to capture data changes, which does not propagate _all_ changes to Redpanda Connect. To capture all changes in a MongoDB cluster, including deletions, enable pre- and post-image saving for the cluster and [required collections](#collections). For more information, see [`document_mode` options](#document_mode) and the [MongoDB documentation](https://www.mongodb.com/docs/manual/changeStreams/#change-streams-with-document-pre—​and-post-images). ## [](#data-replication)Data replication Redpanda Connect allows you to specify which [database collections](#collections) in your source database to receive changes from. You can also run the `mongodb_cdc` input in one of two modes, depending on whether you need a snapshot of existing data before streaming updates. - Snapshot mode: Redpanda Connect first captures a snapshot of all data in the selected collections and streams the contents before processing changes from the last recorded [operations log (oplog)](https://www.mongodb.com/docs/manual/core/replica-set-oplog/) position. - Streaming mode: Redpanda Connect skips the snapshot and processes only the most recent data changes, starting from the latest oplog position. ### [](#snapshot-mode)Snapshot mode If you set the [`stream_snapshot` field](#stream_snapshot) to `true`, Redpanda Connect connects to your MongoDB database and does the following to capture a snapshot of all data in the selected collections: 1. Records the latest oplog position. 2. Determines the strategy for splitting the snapshot data down into shards or chunks for more efficient processing: 1. If [`snapshot_auto_bucket_sharding`](#snapshot_auto_bucket_sharding) is set to `false`, the internal `$splitVector` command is used to compute shards. 2. If [`snapshot_auto_bucket_sharding`](#snapshot_auto_bucket_sharding) is set to `true`, the [`$bucketAuto`](https://www.mongodb.com/docs/manual/reference/operator/aggregation/bucketAuto/) command is used instead. This setting is for environments, such as MongoDB Atlas, where the `$splitVector` command is not available. 3. This input then uses the number of connections specified in [`snapshot-parallelism`](#snapshot_parallelism) to read the selected collections. > 📝 **NOTE** > > If the pipeline restarts during this process, Redpanda Connect must start the snapshot capture from scratch to store the current oplog position in the [`checkpoint_cache`](#checkpoint_cache). 4. Finally, the input uses the stored oplog position to catch up with changes that occurred during snapshot processing. ### [](#streaming-mode)Streaming mode If you set the [`stream_snapshot` field](#stream_snapshot) to `false`, Redpanda Connect connects to your MongoDB database and starts processing data changes from the latest oplog position. If the pipeline restarts, Redpanda Connect resumes processing updates from the last oplog position written to the [`checkpoint_cache`](#checkpoint_cache). ## [](#metadata)Metadata Each message emitted by this plugin has the following metadata: - `operation`: either "insert", "replace", "delete" or "update" for changes streamed. Documents from the initial snapshot have the operation set to "read". - `collection`: the collection the document was written to. - `operation_time`: the oplog time for when this operation occurred. - `schema`: the collection schema in benthos common schema format (set as immutable metadata). Extracted from the collection’s `$jsonSchema` validator if available, otherwise inferred from the first document seen. Not present on messages where no schema could be determined (e.g. deletes without pre-images when no prior schema is cached). ## [](#fields)Fields ### [](#app_name)`app_name` The client application name. **Type**: `string` **Default**: `benthos` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay rejected messages (negative acknowledgements) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#aws)`aws` AWS IAM authentication using the `MONGODB-AWS` mechanism, for example against MongoDB Atlas. When enabled, IAM credentials are used instead of a static username and password. Role-derived session credentials are resolved when the component connects and are re-resolved whenever it reconnects. The `mongodb` processor and cache establish their client once at creation and cannot refresh expiring session credentials, so `role`, `roles` and session tokens are rejected for those components; use the ambient credential chain or long-lived access keys with them. For long-running pipelines, prefer the ambient credential chain (leave keys and roles unset), which the driver refreshes automatically. **Type**: `object` ### [](#aws-enabled)`aws.enabled` Enable AWS IAM authentication using the driver-native `MONGODB-AWS` mechanism. The MongoDB Atlas database user must be created with the AWS IAM authentication type, and connections require TLS. When no static credentials or roles are configured, the ambient AWS credential chain (environment variables, EC2 instance profile, EKS pod role) is used and expiring credentials are refreshed automatically. **Type**: `bool` **Default**: `false` ### [](#aws-id)`aws.id` The ID of credentials to use. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region used when assuming roles (for STS calls). Only used when `role` or `roles` are configured; the ambient and static-key paths ignore it. If no region is specified then the environment default is used. **Type**: `string` ### [](#aws-role)`aws.role` Optional AWS IAM role ARN to assume for authentication. Cannot be combined with `roles`; use the `roles` array instead when chaining multiple roles. **Type**: `string` ### [](#aws-role_external_id)`aws.role_external_id` Optional external ID for the role assumption. Only used with the `role` field, which cannot be combined with `roles`. **Type**: `string` ### [](#aws-roles)`aws.roles[]` Optional array of AWS IAM roles to assume for authentication. Roles can be assumed in sequence, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. Cannot be combined with `role`. **Type**: `array` ### [](#aws-roles-role)`aws.roles[].role` AWS IAM role ARN to assume. **Type**: `string` **Default**: `""` ### [](#aws-roles-role_external_id)`aws.roles[].role_external_id` Optional external ID for the role assumption. **Type**: `string` **Default**: `""` ### [](#aws-secret)`aws.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-session_duration)`aws.session_duration` The duration of the STS session requested when assuming roles. AWS requires at least 15 minutes and caps sessions created through role chaining at one hour. Only used when `role` or `roles` are configured. When using `mongodb_cdc` with role assumption, credentials are freshly resolved after the initial snapshot completes, so the streaming phase starts with a full session. The snapshot itself must still complete within a single session duration: snapshot progress is not checkpointed, so a credential expiry mid-snapshot restarts the snapshot from scratch after reconnecting. Once the snapshot completes and is fully acknowledged, its position is checkpointed, so later restarts resume the stream without re-running the snapshot. For very large snapshots prefer the ambient credential chain. **Type**: `string` **Default**: `1h` ### [](#aws-token)`aws.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#checkpoint_cache)`checkpoint_cache` Specify a [`cache` resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) to store the oplog position for the most recent data update streamed to Redpanda Connect. After a restart, Redpanda Connect can continue processing changes from this position, avoiding the need to reprocess all collection updates. **Type**: `string` ### [](#checkpoint_interval)`checkpoint_interval` The interval between writing checkpoints to the cache. **Type**: `string` **Default**: `5s` ### [](#checkpoint_key)`checkpoint_key` The key identifier used to store the oplog position in [`checkpoint_cache`](#checkpoint_cache). If you have multiple `mongodb_cdc` inputs sharing the same cache, you can provide an alternative key. **Type**: `string` **Default**: `mongodb_cdc_checkpoint` ### [](#checkpoint_limit)`checkpoint_limit` The maximum number of in-flight messages emitted from this input. Increasing this limit enables parallel processing, and batching at the output level. To preserve at-least-once guarantees, any given oplog position is not acknowledged until all messages under that offset are delivered. **Type**: `int` **Default**: `1000` ### [](#checkpoint_write_timeout)`checkpoint_write_timeout` Bounds the checkpoint writes that run outside the normal read loop - storing the position a completed snapshot reached, and clearing a position that can no longer be resumed from - so that a shutdown racing either of them is not extended indefinitely by a slow cache. Raise this for slow remote caches (for example `redis` or `dynamodb`), where losing the post-snapshot write costs a full re-snapshot on the next start. **Type**: `string` **Default**: `10s` ### [](#collections)`collections[]` A list of collections to stream changes from. Specify each collection name as a separate item. **Type**: `array` ### [](#database)`database` The name of the MongoDB database to stream changes from. **Type**: `string` ### [](#document_mode)`document_mode` The mode in which MongoDB emits document changes to Redpanda Connect, specifically updates and deletes. **Type**: `string` **Default**: `update_lookup` | Option | Summary | | --- | --- | | partial_update | In this mode update operations only have a description of the update operation, which follows the following schema: { "_id": , "operations": [ # type == set means that the value was updated like so: # root.foo."bar.baz" = "world" {"path": ["foo", "bar.baz"], "type": "set", "value":"world"}, # type == unset means that the value was deleted like so: # root.qux = deleted() {"path": ["qux"], "type": "unset", "value": null}, # type == truncatedArray means that the array at that path was truncated to value number of elements # root.array = this.array.slice(2) {"path": ["array"], "type": "truncatedArray", "value": 2} ] } | | pre_and_post_images | Uses pre and post image collection to emit the full documents for update and delete operations. To use and configure this mode see the setup steps in the ^MongoDB documentation. | | update_lookup | In this mode insert, replace and update operations have the full document emitted and deletes only have the _id field populated. Documents updates lookup the full document. This corresponds to the updateLookup option, see the ^MongoDB documentation for more information. | ### [](#json_marshal_mode)`json_marshal_mode` Controls the format used to convert a message from BSON to JSON when it is received by Redpanda Connect. **Type**: `string` **Default**: `canonical` | Option | Summary | | --- | --- | | canonical | A string format that emphasizes type preservation at the expense of readability and interoperability. That is, conversion from canonical to BSON will generally preserve type information except in certain specific cases. | | relaxed | A string format that emphasizes readability and interoperability at the expense of type preservation.That is, conversion from relaxed format to BSON can lose type information. | ### [](#on_unresumable_position)`on_unresumable_position` What to do when the stored stream position can no longer be resumed from (for example it has aged out of the oplog) and `stream_snapshot` is disabled, so there is no snapshot to recover with: `fail` stops the input with an error, preserving the checkpoint for inspection; `reset` clears the checkpoint and restarts streaming from the current oplog position, skipping the changes between the lost position and now. When `stream_snapshot` is enabled this field has no effect: recovery re-runs the snapshot, which loses nothing. **Type**: `string` **Default**: `fail` **Options**: `fail`, `reset` ### [](#password)`password` The password to connect to the database. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#read_batch_size)`read_batch_size` The number of documents to fetch in each message batch from MongoDB. **Type**: `int` **Default**: `1000` ### [](#read_max_wait)`read_max_wait` The maximum duration MongoDB waits to accumulate the [`read_batch_size`](#read_batch_size) documents on a change stream before returning the batch to Redpanda Connect. **Type**: `string` **Default**: `1s` ### [](#snapshot_auto_bucket_sharding)`snapshot_auto_bucket_sharding` Uses the [`$bucketAuto`](https://www.mongodb.com/docs/manual/reference/operator/aggregation/bucketAuto/) command instead of the default, `$splitVector`, to split the snapshot data into chunks for processing. This is required for environments, such as MongoDB Atlas, where the `$splitVector` command is not available. To enable parallel processing in these environments: - Set this field to to `true`. - Set `stream_snapshot` to `true`. - Increase `snapshot_parallelism` to a value greater than `1`. **Type**: `bool` **Default**: `false` ### [](#snapshot_parallelism)`snapshot_parallelism` Specifies the number of connections to use when reading the initial snapshot from one or more collections. Increase this number to enable parallel processing of the snapshot. This feature uses the `$splitVector` command to split snapshot data into chunks for more efficient processing. This field is only applicable when `stream_snapshot` is set to `true`. **Type**: `int` **Default**: `1` ### [](#stream_snapshot)`stream_snapshot` When set to `true`, this input streams a snapshot of all existing data in the source collections before streaming data changes. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target MongoDB server. **Type**: `string` ```yaml # Examples: url: mongodb://localhost:27017 ``` ### [](#username)`username` The username to connect to the database. **Type**: `string` **Default**: `""` --- # Page 66: mongodb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/mongodb.md --- # mongodb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mongodb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/mongodb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/mongodb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/mongodb.adoc description: Executes a query and creates a message for each document received. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- Executes a query and creates a message for each document received. #### Common ```yml inputs: label: "" mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" collection: "" # No default (required) query: "" # No default (required) auto_replay_nacks: true batch_size: "" # No default (optional) sort: "" # No default (optional) limit: "" # No default (optional) ``` #### Advanced ```yml inputs: label: "" mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" app_name: benthos aws: enabled: false region: "" # No default (optional) session_duration: 1h id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) roles: [] # No default (optional) collection: "" # No default (required) operation: find json_marshal_mode: canonical query: "" # No default (required) auto_replay_nacks: true batch_size: "" # No default (optional) sort: "" # No default (optional) limit: "" # No default (optional) ``` Once the documents from the query are exhausted, this input shuts down, allowing the pipeline to gracefully terminate (or the next input in a [sequence](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/) to execute). ## [](#fields)Fields ### [](#app_name)`app_name` The client application name. **Type**: `string` **Default**: `benthos` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#aws)`aws` AWS IAM authentication using the `MONGODB-AWS` mechanism, for example against MongoDB Atlas. When enabled, IAM credentials are used instead of a static username and password. Role-derived session credentials are resolved when the component connects and are re-resolved whenever it reconnects. The `mongodb` processor and cache establish their client once at creation and cannot refresh expiring session credentials, so `role`, `roles` and session tokens are rejected for those components; use the ambient credential chain or long-lived access keys with them. For long-running pipelines, prefer the ambient credential chain (leave keys and roles unset), which the driver refreshes automatically. **Type**: `object` ### [](#aws-enabled)`aws.enabled` Enable AWS IAM authentication using the driver-native `MONGODB-AWS` mechanism. The MongoDB Atlas database user must be created with the AWS IAM authentication type, and connections require TLS. When no static credentials or roles are configured, the ambient AWS credential chain (environment variables, EC2 instance profile, EKS pod role) is used and expiring credentials are refreshed automatically. **Type**: `bool` **Default**: `false` ### [](#aws-id)`aws.id` The ID of credentials to use. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region used when assuming roles (for STS calls). Only used when `role` or `roles` are configured; the ambient and static-key paths ignore it. If no region is specified then the environment default is used. **Type**: `string` ### [](#aws-role)`aws.role` Optional AWS IAM role ARN to assume for authentication. Cannot be combined with `roles`; use the `roles` array instead when chaining multiple roles. **Type**: `string` ### [](#aws-role_external_id)`aws.role_external_id` Optional external ID for the role assumption. Only used with the `role` field, which cannot be combined with `roles`. **Type**: `string` ### [](#aws-roles)`aws.roles[]` Optional array of AWS IAM roles to assume for authentication. Roles can be assumed in sequence, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. Cannot be combined with `role`. **Type**: `array` ### [](#aws-roles-role)`aws.roles[].role` AWS IAM role ARN to assume. **Type**: `string` **Default**: `""` ### [](#aws-roles-role_external_id)`aws.roles[].role_external_id` Optional external ID for the role assumption. **Type**: `string` **Default**: `""` ### [](#aws-secret)`aws.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-session_duration)`aws.session_duration` The duration of the STS session requested when assuming roles. AWS requires at least 15 minutes and caps sessions created through role chaining at one hour. Only used when `role` or `roles` are configured. For long-running pipelines, prefer the ambient credential chain over a fixed session duration, since the driver refreshes ambient credentials automatically as they near expiry. **Type**: `string` **Default**: `1h` ### [](#aws-token)`aws.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#batch_size)`batch_size` A explicit number of documents to batch up before flushing them for processing. Must be greater than `0`. Operations: `find`, `aggregate` **Type**: `int` ```yaml # Examples: batch_size: 1000 ``` ### [](#collection)`collection` The collection to select from. **Type**: `string` ### [](#database)`database` The name of the target MongoDB database. **Type**: `string` ### [](#json_marshal_mode)`json_marshal_mode` The json\_marshal\_mode setting is optional and controls the format of the output message. **Type**: `string` **Default**: `canonical` | Option | Summary | | --- | --- | | canonical | A string format that emphasizes type preservation at the expense of readability and interoperability. That is, conversion from canonical to BSON will generally preserve type information except in certain specific cases. | | relaxed | A string format that emphasizes readability and interoperability at the expense of type preservation.That is, conversion from relaxed format to BSON can lose type information. | ### [](#limit)`limit` An explicit maximum number of documents to return. Operations: `find` **Type**: `int` ### [](#operation)`operation` The mongodb operation to perform. **Type**: `string` **Default**: `find` **Options**: `find`, `aggregate` ### [](#password)`password` The password to connect to the database. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#query)`query` Bloblang expression describing MongoDB query. **Type**: `string` ```yaml # Examples: query: |- root.from = {"$lte": timestamp_unix()} root.to = {"$gte": timestamp_unix()} ``` ### [](#sort)`sort` An object specifying fields to sort by, and the respective sort order (`1` ascending, `-1` descending). Note: The driver currently appears to support only one sorting key. Operations: `find` **Type**: `object` ```yaml # Examples: sort: name: 1 # --- sort: age: -1 ``` ### [](#url)`url` The URL of the target MongoDB server. **Type**: `string` ```yaml # Examples: url: mongodb://localhost:27017 ``` ### [](#username)`username` The username to connect to the database. **Type**: `string` **Default**: `""` --- # Page 67: mqtt **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/mqtt.md --- # mqtt > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mqtt latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/mqtt page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/mqtt.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/mqtt.adoc description: Subscribe to topics on MQTT brokers. page-git-created-date: "2024-11-07" page-git-modified-date: "2026-05-26" --- Subscribe to topics on MQTT brokers. #### Common ```yml inputs: label: "" mqtt: urls: [] # No default (required) client_id: "" connect_timeout: 30s topics: [] # No default (required) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" mqtt: urls: [] # No default (required) client_id: "" dynamic_client_id_suffix: "" # No default (optional) connect_timeout: 30s will: enabled: false qos: 0 retained: false topic: "" payload: "" user: "" password: "" keepalive: 30 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] topics: [] # No default (required) qos: 1 clean_session: true auto_replay_nacks: true ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `mqtt_duplicate` - `mqtt_qos` - `mqtt_retained` - `mqtt_topic` - `mqtt_message_id` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#clean_session)`clean_session` Set whether the connection is non-persistent. **Type**: `bool` **Default**: `true` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `""` ### [](#connect_timeout)`connect_timeout` The maximum amount of time to wait in order to establish a connection before the attempt is abandoned. **Type**: `string` **Default**: `30s` ```yaml # Examples: connect_timeout: 1s # --- connect_timeout: 500ms ``` ### [](#dynamic_client_id_suffix)`dynamic_client_id_suffix` Append a dynamically generated suffix to the specified `client_id` on each run of the pipeline. This can be useful when clustering Redpanda Connect producers. **Type**: `string` | Option | Summary | | --- | --- | | nanoid | append a nanoid of length 21 characters | ### [](#keepalive)`keepalive` Max seconds of inactivity before a keepalive message is sent. **Type**: `int` **Default**: `30` ### [](#password)`password` A password to connect with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#qos)`qos` The level of delivery guarantee to enforce. Has options 0, 1, 2. **Type**: `int` **Default**: `1` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#topics)`topics[]` A list of topics to consume from. **Type**: `array` ### [](#urls)`urls[]` A list of URLs to connect to. Use the format `scheme://host:port`, where: - `scheme` is one of the following: `tcp`, `ssl`, `ws` - `host` is the IP address or hostname - `port` is the port on which the MQTT broker accepts connections If an item in the list contains commas, it is expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "tcp://localhost:1883" ``` ### [](#user)`user` A username to connect with. **Type**: `string` **Default**: `""` ### [](#will)`will` Set last will message in case of Redpanda Connect failure **Type**: `object` ### [](#will-enabled)`will.enabled` Whether to enable last will messages. **Type**: `bool` **Default**: `false` ### [](#will-payload)`will.payload` Set payload for last will message. **Type**: `string` **Default**: `""` ### [](#will-qos)`will.qos` Set QoS for last will message. Valid values are: 0, 1, 2. **Type**: `int` **Default**: `0` ### [](#will-retained)`will.retained` Set retained for last will message. **Type**: `bool` **Default**: `false` ### [](#will-topic)`will.topic` Set topic for last will message. **Type**: `string` **Default**: `""` --- # Page 68: mysql_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/mysql_cdc.md --- # mysql_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mysql_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/mysql_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/mysql_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/mysql_cdc.adoc page-git-created-date: "2025-02-20" page-git-modified-date: "2026-05-26" --- Streams data changes from a MySQL database, using MySQL’s binary log to capture data updates. This input is built on the [`mysql-canal` library](https://github.com/go-mysql-org/go-mysql?tab=readme-ov-file#replication) but uses a custom approach for streaming historical data. #### Common ```yml inputs: label: "" mysql_cdc: flavor: mysql dsn: "" # No default (required) tables: [] # No default (required) checkpoint_cache: "" # No default (required) checkpoint_key: mysql_binlog_position snapshot_max_batch_size: 1000 stream_snapshot: "" # No default (required) max_parallel_snapshot_tables: 1 auto_replay_nacks: true checkpoint_limit: 1024 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml inputs: label: "" mysql_cdc: flavor: mysql dsn: "" # No default (required) tables: [] # No default (required) checkpoint_cache: "" # No default (required) checkpoint_key: mysql_binlog_position snapshot_max_batch_size: 1000 max_reconnect_attempts: 10 stream_snapshot: "" # No default (required) max_parallel_snapshot_tables: 1 auto_replay_nacks: true checkpoint_limit: 1024 tls: skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] aws: enabled: false region: "" # No default (optional) endpoint: "" # No default (required) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) roles: [] # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` The `mysql_cdc` input uses MySQL’s [binary log (`binlog`)](https://dev.mysql.com/doc/refman/8.0/en/binary-log.html) to capture changes made to a MySQL database in real time and streams them to Redpanda Connect. Redpanda Connect allows you to specify which [database tables](#tables) in your source database to receive changes from. There are also [two replication modes](#choose-a-replication-mode) to choose from. ## [](#prerequisites)Prerequisites - MySQL version 8 or later - Network access from the cluster where your Redpanda Connect pipeline is running to the source database environment. For detailed networking information, including how to set up a VPC peering connection, see [Redpanda Cloud Networking](https://docs.redpanda.com/cloud-data-platform/networking/). - A MySQL instance with binary logging enabled ### [](#configuration-resources)Configuration resources #### Cloud platforms - [Change data capture on Amazon RDS for MySQL](https://aws.amazon.com/blogs/database/enable-change-data-capture-on-amazon-rds-for-mysql-applications-that-are-using-xa-transactions/) - [Azure MySQL Database (CDC)](https://learn.microsoft.com/en-us/fabric/real-time-hub/add-source-mysql-database-cdc) - [Google Cloud SQL for MySQL](https://cloud.google.com/datastream/docs/configure-cloudsql-mysql) #### Self-hosted MySQL - [Binary Logging Options and Variables](https://dev.mysql.com/doc/refman/8.4/en/replication-options-binary-log.html) ## [](#choose-a-replication-mode)Choose a replication mode You can run the `mysql_cdc` input in one of two modes, depending on whether you need a snapshot of existing data. - Snapshot mode: Redpanda Connect first captures a snapshot of all data in the selected tables and streams the contents before processing changes from the last recorded binlog position. - Streaming mode: Redpanda Connect skips the snapshot and processes only the most recent data changes, starting from the latest binlog position. ### [](#snapshot-mode)Snapshot mode If you set the [`stream_snapshot` field](#stream_snapshot) to `true`, Redpanda Connect connects to your MySQL database and does the following to capture a snapshot of all data in the selected tables: 1. Executes the `FLUSH TABLES WITH READ LOCK` query to write any outstanding table updates to disk, and locks the tables. 2. Runs the `START TRANSACTION WITH CONSISTENT SNAPSHOT` statement to create a new transaction with a consistent view of all data, capturing the state of the database at the moment the transaction started. 3. Reads the current binlog position. 4. Runs the `UNLOCK TABLES` statement to release the database. 5. Preserves the initial transaction for data integrity. > 📝 **NOTE** > > If the pipeline restarts during this process, Redpanda Connect must start the snapshot capture from scratch to store the current binlog position in the [`checkpoint_cache`](#checkpoint_cache). After the snapshot is taken, the input executes SELECT statements to extract data from the selected tables in two stages: 1. The input finds the primary keys of a table. 2. It selects the data ordered by primary key. Finally, the input uses the stored binlog position to catch up with changes that occurred during snapshot processing. ### [](#streaming-mode)Streaming mode If you set the [`stream_snapshot` field](#stream_snapshot) to `false`, Redpanda Connect connects to your MySQL database and starts processing data changes from the latest binlog position. If the pipeline restarts, Redpanda Connect resumes processing updates from the last binlog position written to the [`checkpoint_cache`](#checkpoint_cache). ## [](#binlog-rotation)Binlog rotation While the `mysql_cdc` input is streaming changes to Redpanda Connect, your MySQL server may rotate the binlog file. When this occurs, Redpanda Connect flushes the existing message batch and stores the new binlog position so that it can resume processing using the latest offset. ## [](#data-mappings)Data mappings The following table shows how selected MySQL data types are mapped to data types supported in Redpanda Connect. All other data types are mapped to string values. | MySQL data type | Bloblang value | | --- | --- | | TEXT, VARCHAR | A string value, for example: "this data" | | BINARY, VARBINARY, TINYBLOB, BLOB, MEDIUMBLOB, LONGBLOB | An array of byte values, for example: [byte1,byte2,byte3] | | DECIMAL, NUMERIC, TINYINT, SMALLINT, MEDIUMINT, INT, BIGINT, YEAR | A standard numeric type, for example: 123 | | FLOAT, DOUBLE | A 64-bit decimal (float64), for example: 123.1234 | | DATETIME, TIMESTAMP | A Bloblang timestamp, for example:1257894000000 2009-11-10 23:00:00 +0000 UTC | | SET | An array of strings, for example: ["apple", "banana", "orange"] | | JSON | A map object of the JSON, for example: {"red": 1, "blue": 2, "green": 3} | ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `operation`: The type of operation (insert, update, delete, or read for snapshot messages) - `table`: The name of the table - `binlog_position`: The binlog position (for CDC messages only, not set for snapshot messages) - `schema`: The table schema in benthos common schema format, compatible with processors like parquet\_encode ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay rejected messages (negative acknowledgements) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#aws)`aws` AWS IAM authentication configuration for MySQL instances. When enabled, IAM credentials are used to generate temporary authentication tokens instead of a static password. **Type**: `object` ### [](#aws-enabled)`aws.enabled` Enable AWS IAM authentication for MySQL. When enabled, an IAM authentication token is generated and used as the password. When using IAM authentication ensure `max_reconnect_attempts` is set to a low value to ensure it can refresh credentials. **Type**: `bool` **Default**: `false` ### [](#aws-endpoint)`aws.endpoint` The MySQL endpoint hostname (e.g., mydb.abc123.us-east-1.rds.amazonaws.com). **Type**: `string` ### [](#aws-id)`aws.id` The ID of credentials to use. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region where the MySQL instance is located. If no region is specified then the environment default will be used. **Type**: `string` ### [](#aws-role)`aws.role` Optional AWS IAM role ARN to assume for authentication. Alternatively, use `roles` array for role chaining instead. **Type**: `string` ### [](#aws-role_external_id)`aws.role_external_id` Optional external ID for the role assumption. Only used with the `role` field. Alternatively, use `roles` array for role chaining instead. **Type**: `string` ### [](#aws-roles)`aws.roles[]` Optional array of AWS IAM roles to assume for authentication. Roles can be assumed in sequence, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. **Type**: `array` ### [](#aws-roles-role)`aws.roles[].role` AWS IAM role ARN to assume. **Type**: `string` **Default**: `""` ### [](#aws-roles-role_external_id)`aws.roles[].role_external_id` Optional external ID for the role assumption. **Type**: `string` **Default**: `""` ### [](#aws-secret)`aws.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-token)`aws.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_cache)`checkpoint_cache` Specify a `cache` resource to store the binlog position of the most recent data update delivered to Redpanda Connect. After a restart, Redpanda Connect can continue processing changes from this last known position, avoiding the need to reprocess all table updates. **Type**: `string` ### [](#checkpoint_key)`checkpoint_key` The key identifier used to store the binlog position in [`checkpoint_cache`](#checkpoint_cache). If you have multiple `mysql_cdc` inputs sharing the same cache, you can provide an alternative key. **Type**: `string` **Default**: `mysql_binlog_position` ### [](#checkpoint_limit)`checkpoint_limit` The maximum number of messages that this input can process at a given time. Increasing this limit enables parallel processing, and batching at the output level. To preserve at-least-once guarantees, any given binlog position is not acknowledged until all messages under that offset are delivered. **Type**: `int` **Default**: `1024` ### [](#dsn)`dsn` The data source name (DSN) of the MySQL database from which you want to stream updates. Use the format `user:password@tcp(localhost:3306)/database`. **Type**: `string` ```yaml # Examples: dsn: user:password@tcp(localhost:3306)/database ``` ### [](#flavor)`flavor` The type of MySQL database to connect to. **Type**: `string` **Default**: `mysql` | Option | Summary | | --- | --- | | mariadb | MariaDB flavored databases. | | mysql | MySQL flavored databases. | ### [](#max_parallel_snapshot_tables)`max_parallel_snapshot_tables` Specifies the number of tables that will be snapshotted in parallel. **Type**: `int` **Default**: `1` ### [](#max_reconnect_attempts)`max_reconnect_attempts` The maximum number of attempts the MySQL driver will try to re-establish a broken connection before Connect attempts reconnection. A zero or negative number means infinite retry attempts. **Type**: `int` **Default**: `10` ### [](#snapshot_max_batch_size)`snapshot_max_batch_size` The maximum number of table rows to fetch in each batch when taking a snapshot. This option is only available when `stream_snapshot` is set to `true`. **Type**: `int` **Default**: `1000` ### [](#stream_snapshot)`stream_snapshot` When set to `true`, this input streams a snapshot of all existing data in the source database before streaming data changes. To use this setting, all database tables that you want to replicate _must_ have a primary key. **Type**: `bool` ### [](#tables)`tables[]` A list of the database table names to stream changes from. Specify each table name as a separate item. **Type**: `array` ```yaml # Examples: tables: - table1 - table2 ``` ### [](#tls)`tls` Using this field overrides the SSL/TLS settings in the environment and DSN. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` --- # Page 69: nats_jetstream **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/nats_jetstream.md --- # nats_jetstream > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats_jetstream latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/nats_jetstream page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/nats_jetstream.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/nats_jetstream.adoc description: Reads messages from NATS JetStream subjects. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Reads messages from NATS JetStream subjects. #### Common ```yml inputs: label: "" nats_jetstream: urls: [] # No default (required) queue: "" # No default (optional) subject: "" # No default (optional) durable: "" # No default (optional) stream: "" # No default (optional) bind: "" # No default (optional) deliver: all ``` #### Advanced ```yml inputs: label: "" nats_jetstream: urls: [] # No default (required) max_reconnects: "" # No default (optional) queue: "" # No default (optional) subject: "" # No default (optional) durable: "" # No default (optional) stream: "" # No default (optional) bind: "" # No default (optional) create_stream: false deliver: all ack_wait: 30s max_ack_pending: 1024 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) extract_tracing_map: "" # No default (optional) ``` ## [](#consume-mirrored-streams)Consume mirrored streams When a stream being consumed is mirrored in a different JetStream domain, the stream cannot be resolved from the subject name alone. You must specify the stream name as well as the subject (if applicable). ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `nats_subject` - `nats_sequence_stream` - `nats_sequence_consumer` - `nats_num_delivered` - `nats_num_pending` - `nats_domain` - `nats_timestamp_unix_nano` - `nats_consumer` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#ack_wait)`ack_wait` The maximum amount of time NATS server should wait for an ack from consumer. **Type**: `string` **Default**: `30s` ```yaml # Examples: ack_wait: 100ms # --- ack_wait: 5m ``` ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#bind)`bind` Indicates that the subscription should use an existing consumer. **Type**: `bool` ### [](#create_stream)`create_stream` Whether to automatically create the stream if it doesn’t exist (requires the stream field to be set). **Type**: `bool` **Default**: `false` ### [](#deliver)`deliver` Determines which messages to deliver when consuming without a durable subscriber. **Type**: `string` **Default**: `all` | Option | Summary | | --- | --- | | all | Deliver all available messages. | | last | Deliver starting with the last published messages. | | last_per_subject | Deliver starting with the last published message per subject. | | new | Deliver starting from now, not taking into account any previous messages. | ### [](#durable)`durable` Preserve the state of your consumer under a durable name. **Type**: `string` ### [](#extract_tracing_map)`extract_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that attempts to extract an object containing tracing propagation information, which will then be used as the root tracing span for the message. The specification of the extracted fields must match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: extract_tracing_map: root = @ # --- extract_tracing_map: root = this.meta.span ``` ### [](#max_ack_pending)`max_ack_pending` The maximum number of outstanding acks to be allowed before consuming is halted. **Type**: `int` **Default**: `1024` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#queue)`queue` An optional queue group to consume as. **Type**: `string` ### [](#stream)`stream` A stream to consume from. Either a subject or stream must be specified. **Type**: `string` ### [](#subject)`subject` A subject to consume from. Supports wildcards for consuming multiple subjects. Either a subject or stream must be specified. **Type**: `string` ```yaml # Examples: subject: foo.bar.baz # --- subject: foo.*.baz # --- subject: foo.bar.* # --- subject: foo.> ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 70: nats_kv **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/nats_kv.md --- # nats_kv > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats_kv latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/nats_kv page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/nats_kv.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/nats_kv.adoc description: Watches for updates in a NATS key-value bucket. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Watches for updates in a NATS key-value bucket. #### Common ```yml inputs: label: "" nats_kv: urls: [] # No default (required) bucket: "" # No default (required) key: > auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" nats_kv: urls: [] # No default (required) max_reconnects: "" # No default (optional) bucket: "" # No default (required) key: > auto_replay_nacks: true ignore_deletes: false include_history: false meta_only: false tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `nats_kv_key` - `nats_kv_bucket` - `nats_kv_revision` - `nats_kv_delta` - `nats_kv_operation` - `nats_kv_created` ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the corresponding user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#bucket)`bucket` The name of the KV bucket. **Type**: `string` ```yaml # Examples: bucket: my_kv_bucket ``` ### [](#ignore_deletes)`ignore_deletes` Do not send delete markers as messages. **Type**: `bool` **Default**: `false` ### [](#include_history)`include_history` Include all the history per key, not just the last one. **Type**: `bool` **Default**: `false` ### [](#key)`key` Key to watch for updates, can include wildcards. **Type**: `string` **Default**: `>` ```yaml # Examples: key: foo.bar.baz # --- key: foo.*.baz # --- key: foo.bar.* # --- key: foo.> ``` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#meta_only)`meta_only` Retrieve only the metadata of the entry **Type**: `bool` **Default**: `false` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 71: nats **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/nats.md --- # nats > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/nats page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/nats.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/nats.adoc description: Subscribe to a NATS subject. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Subscribe to a NATS subject. #### Common ```yml inputs: label: "" nats: urls: [] # No default (required) subject: "" # No default (required) queue: "" # No default (optional) auto_replay_nacks: true send_ack: true ``` #### Advanced ```yml inputs: label: "" nats: urls: [] # No default (required) max_reconnects: "" # No default (optional) subject: "" # No default (required) queue: "" # No default (optional) auto_replay_nacks: true send_ack: true nak_delay: "" # No default (optional) prefetch_count: 500000 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) extract_tracing_map: "" # No default (optional) ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `nats_subject` - `nats_reply_subject` - All message headers (when supported by the connection) You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the corresponding user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#extract_tracing_map)`extract_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that attempts to extract an object containing tracing propagation information, which will then be used as the root tracing span for the message. The specification of the extracted fields must match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: extract_tracing_map: root = @ # --- extract_tracing_map: root = this.meta.span ``` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#nak_delay)`nak_delay` An optional delay duration on redelivering a message when negatively acknowledged. **Type**: `string` ```yaml # Examples: nak_delay: 1m ``` ### [](#prefetch_count)`prefetch_count` The maximum number of messages to pull at a time. **Type**: `int` **Default**: `500000` ### [](#queue)`queue` An optional queue group to consume as. **Type**: `string` ### [](#send_ack)`send_ack` Whether an automatic acknowledgment is sent as a reply to each message. When enabled, these replies are sent only when data has been delivered to all outputs. **Type**: `bool` **Default**: `true` ### [](#subject)`subject` A subject to consume from. Supports wildcards for consuming multiple subjects. Either a subject or stream must be specified. **Type**: `string` ```yaml # Examples: subject: foo.bar.baz # --- subject: foo.*.baz # --- subject: foo.bar.* # --- subject: foo.> ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 72: oracledb_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/oracledb_cdc.md --- # oracledb_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: oracledb_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/oracledb_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/oracledb_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/oracledb_cdc.adoc description: Enables Change Data Capture by consuming from OracleDB. page-git-created-date: "2026-03-31" page-git-modified-date: "2026-08-11" --- Enables Change Data Capture by consuming from OracleDB. Streams changes from an Oracle database for Change Data Capture (CDC). Use the `snapshot_mode` field to control whether existing data is captured in an initial snapshot before streaming changes. #### Common ```yml inputs: label: "" oracledb_cdc: connection_string: "" # No default (required) wallet_path: "" # No default (optional) wallet_password: "" # No default (optional) snapshot_mode: "" # No default (optional) max_parallel_snapshot_tables: 1 snapshot_max_batch_size: 1000 logminer: scn_window_size: 20000 min_scn_window_size: 1000 max_scn_window_size: 100000 backoff_interval: 5s mining_interval: 300ms strategy: online_catalog max_transaction_events: 0 lob_enabled: true transaction_cache: "" # No default (optional) transaction_cache_key: oracledb_cdc max_session_age: 0s snapshot_filters: "" # No default (optional) include: [] # No default (required) exclude: [] # No default (optional) checkpoint_cache: "" # No default (optional) checkpoint_cache_table_name: RPCN.CDC_CHECKPOINT_CACHE checkpoint_cache_key: oracledb_cdc checkpoint_limit: 1024 pdb_name: "" # No default (optional) auto_replay_nacks: true batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml inputs: label: "" oracledb_cdc: connection_string: "" # No default (required) wallet_path: "" # No default (optional) wallet_password: "" # No default (optional) snapshot_mode: "" # No default (optional) max_parallel_snapshot_tables: 1 snapshot_max_batch_size: 1000 logminer: scn_window_size: 20000 min_scn_window_size: 1000 max_scn_window_size: 100000 backoff_interval: 5s mining_interval: 300ms strategy: online_catalog max_transaction_events: 0 lob_enabled: true transaction_cache: "" # No default (optional) transaction_cache_key: oracledb_cdc max_session_age: 0s snapshot_filters: "" # No default (optional) include: [] # No default (required) exclude: [] # No default (optional) checkpoint_cache: "" # No default (optional) checkpoint_cache_table_name: RPCN.CDC_CHECKPOINT_CACHE checkpoint_cache_key: oracledb_cdc checkpoint_limit: 1024 pdb_name: "" # No default (optional) auto_replay_nacks: true batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `database_schema`: The database schema for the table where the message originates from. - `table_name`: Name of the table that the message originated from. - `operation`: Type of operation that generated the message: "read", "delete", "insert", or "update". "read" is from messages that are read in the initial snapshot phase. - `scn`: The System Change Number in Oracle. Messages published as part of a snapshot will contain Oracle’s current SCN captured at time of snapshot. - `transaction_id`: The Oracle transaction ID in `USN.SLOT.SEQ` format, identifying the transaction that produced the change. Not present on snapshot (`read`) messages. - `source_ts_ms`: The timestamp of when Oracle wrote the change record into the redo log, expressed as milliseconds since the Unix epoch. This reflects the database server’s wall-clock time at the moment the DML executed, not the transaction commit time. - `commit_ts_ms`: The timestamp of the transaction commit, expressed as milliseconds since the Unix epoch. Sourced from `V$LOGMNR_CONTENTS.TIMESTAMP` on the COMMIT redo record — this is Oracle’s wall-clock time when the commit was written to the redo log, not a dedicated commit-timestamp column. For snapshot (`read`) messages, this reflects Oracle’s `SYSTIMESTAMP` at the moment the snapshot SCN was captured, so all snapshot messages share the same value. - `schema`: The table schema, for use with schema-aware downstream processors such as `schema_registry_encode`. When new columns are detected in CDC events, the schema is automatically refreshed from the Oracle catalog. Dropped columns are reflected after a connector restart. > 📝 **NOTE** > > `source_ts_ms` is not present on snapshot (`read`) messages. ## [](#permissions)Permissions When using the default Oracle-based cache, the Connect user requires permission to create tables and stored procedures, and the rpcn schema must already exist. See `checkpoint_cache_table_name` for more information. ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_cache)`checkpoint_cache` A [cache resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) to use for storing the current System Change Number (SCN) that has been successfully delivered. This allows Redpanda Connect to continue from that SCN upon restart, rather than consume the entire state of OracleDB redo logs. If not set, the default Oracle-based cache is used. See `checkpoint_cache_table_name` for more information. **Type**: `string` ### [](#checkpoint_cache_key)`checkpoint_cache_key` The key to use to store the snapshot position in `checkpoint_cache`. An alternative key can be provided if multiple CDC inputs share the same cache. **Type**: `string` **Default**: `oracledb_cdc` ### [](#checkpoint_cache_table_name)`checkpoint_cache_table_name` The identifier for the checkpoint cache table name. If no `checkpoint_cache` field is specified, this input will automatically create a table and stored procedure under the `rpcn` schema to act as a checkpoint cache. This table stores the latest processed System Change Number (SCN) that has been successfully delivered, allowing Redpanda Connect to resume from that point upon restart rather than reconsume the entire redo log. When `pdb_name` is set and this field is left at its default value, the table name is automatically derived per PDB (e.g. `RPCN.CDC_CHECKPOINT_MYPDB`) to avoid SCN collisions between pipelines monitoring different PDBs. Set this field explicitly to opt out of that auto-derivation. **Type**: `string` **Default**: `RPCN.CDC_CHECKPOINT_CACHE` ```yaml # Examples: checkpoint_cache_table_name: RPCN.CHECKPOINT_CACHE ``` ### [](#checkpoint_limit)`checkpoint_limit` The maximum number of messages that can be processed at a given time. Increasing this limit enables parallel processing and batching at the output level. Any given System Change Number (SCN) will not be acknowledged unless all messages under that offset are delivered in order to preserve at least once delivery guarantees. **Type**: `int` **Default**: `1024` ### [](#connection_string)`connection_string` The connection string of the Oracle database to connect to. You can supply additional connection options as URL query parameters, for example: `oracle://user:password@host:1522/service?WALLET=/opt/oracle/wallet&SSL=true`. **Type**: `string` ```yaml # Examples: connection_string: oracle://username:password@host:port/service_name # --- connection_string: oracle://user:password@host:1522/service?WALLET=/opt/oracle/wallet&SSL=true ``` ### [](#exclude)`exclude[]` Regular expressions for tables to exclude. **Type**: `array` ```yaml # Examples: exclude: SCHEMA.PRIVATETABLE ``` ### [](#include)`include[]` Regular expressions for tables to include. **Type**: `array` ```yaml # Examples: include: SCHEMA.PRODUCTS ``` ### [](#logminer)`logminer` LogMiner configuration settings. **Type**: `object` ### [](#logminer-backoff_interval)`logminer.backoff_interval` The interval between attempts to check for new changes once all data is processed. For low traffic tables increasing this value can reduce network traffic to the server. **Type**: `string` **Default**: `5s` ```yaml # Examples: backoff_interval: 5s # --- backoff_interval: 1m ``` ### [](#logminer-lob_enabled)`logminer.lob_enabled` When enabled, large object (CLOB, BLOB) columns are included in both snapshot and streaming change events. When disabled, these columns are still present but contain no values. Enabling this option introduces additional performance overhead and increases memory requirements. **Type**: `bool` **Default**: `true` ### [](#logminer-max_scn_window_size)`logminer.max_scn_window_size` The maximum SCN range that can be mined in a single cycle. The window starts at scn\_window\_size and grows by scn\_window\_size each cycle that ends at the cap (backlog present), up to this limit. It shrinks by the same step each cycle that catches up to the database. This allows the connector to automatically mine larger windows during heavy backlog and smaller windows during steady state. **Type**: `int` **Default**: `100000` ### [](#logminer-max_session_age)`logminer.max_session_age` The maximum duration a single LogMiner session may stay open before being forcibly ended and restarted, even if the underlying redo log files haven’t changed. By default, a LogMiner session is only restarted when a redo log switch is detected. On databases where switches are infrequent, a session can stay open for a long time, and LogMiner has been observed to accumulate server-side PGA memory (particularly around online catalog dictionary lookups) until Oracle terminates the session with ORA-04036. Setting this forces a periodic restart independent of log switches. Set to 0 (default) to disable and restart only on log switches. **Type**: `string` **Default**: `0s` ```yaml # Examples: max_session_age: 20m ``` ### [](#logminer-max_transaction_events)`logminer.max_transaction_events` The maximum number of events that can be buffered for a single transaction. If a transaction exceeds this limit it is discarded and its events will not be emitted. Set to 0 to disable the limit. **Type**: `int` **Default**: `0` ### [](#logminer-min_scn_window_size)`logminer.min_scn_window_size` The minimum SCN gap required before starting a new LogMiner session. When the gap between the connector’s current position and the database’s current SCN is smaller than this value, the mining cycle is skipped and the connector backs off instead. This prevents excessive LogMiner start/stop cycles on low-traffic databases where Oracle background activity advances the SCN without producing relevant events. Set to 0 to disable. **Type**: `int` **Default**: `1000` ### [](#logminer-mining_interval)`logminer.mining_interval` The interval between mining cycles during normal operation. Controls how frequently LogMiner polls for new changes when not caught up. **Type**: `string` **Default**: `300ms` ```yaml # Examples: mining_interval: 100ms # --- mining_interval: 1s ``` ### [](#logminer-scn_window_size)`logminer.scn_window_size` The SCN range to mine per cycle. Each cycle reads changes between the current SCN and current SCN + scn\_window\_size. Smaller values mean more frequent queries with lower memory usage but higher overhead; larger values reduce query frequency and improve throughput at the cost of higher memory usage per cycle. **Type**: `int` **Default**: `20000` ### [](#logminer-strategy)`logminer.strategy` Controls how LogMiner retrieves data dictionary information. `online_catalog` uses the current data dictionary for best performance but cannot capture DDL changes. Currently, only `online_catalog` is supported. **Type**: `string` **Default**: `online_catalog` ### [](#logminer-transaction_cache)`logminer.transaction_cache` A [cache resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) to use for buffering in-flight transactions. When set, DML events are serialized and stored in the named cache rather than held in memory, reducing connector memory usage for workloads with large or long-running transactions. If not set, an in-memory buffer is used. Each in-flight transaction is stored as N+1 cache entries: one metadata key holding the transaction ID, start SCN, and event count; and one event key per DML event. A transaction with 1000 events occupies 1001 cache entries. Each AddEvent call writes exactly two keys regardless of how many events the transaction has already accumulated. This cache is designed for low-latency stores with cheap per-operation cost. Redis and Memcached are the recommended backends. The built-in `memory:{}` cache works but provides no durability across restarts. High-latency or per-request-cost stores such as S3 or DynamoDB are not recommended - a transaction with 1000 events generates approximately 3000 cache operations across its lifetime, and because LogMiner processes events on a single goroutine, per-call latency directly reduces throughput. A backend that causes timeouts or errors will also cause the mining cycle to restart from an earlier checkpoint SCN, which can result in duplicate event delivery. **Type**: `string` ### [](#logminer-transaction_cache_key)`logminer.transaction_cache_key` The key prefix used when storing transactions in `transaction_cache`. An alternative prefix must be set if multiple `oracledb_cdc` inputs share the same cache resource, since Oracle transaction IDs (USN.SLOT.SEQ) are only unique within a single Oracle instance and would otherwise collide. **Type**: `string` **Default**: `oracledb_cdc` ### [](#max_parallel_snapshot_tables)`max_parallel_snapshot_tables` Specifies a number of tables that will be processed in parallel during the snapshot processing stage. **Type**: `int` **Default**: `1` ### [](#pdb_name)`pdb_name` The name of the pluggable database (PDB) to monitor. When connecting to a CDB root, LogMiner output is scoped to this PDB via SRC\_CON\_NAME filtering and catalog queries use ALTER SESSION SET CONTAINER to switch context. Requires GRANT SET CONTAINER TO CONTAINER=ALL. **Type**: `string` ### [](#snapshot_filters)`snapshot_filters` A map of fully-qualified table names (for example `SCHEMA.TABLE`) to SQL `SELECT` queries that override the default snapshot query for each table. Use this to filter or shape the rows captured during the initial snapshot. Each query must project every column of the table’s primary key (all columns of a composite key), even if it otherwise selects only a subset of columns. During a snapshot, Redpanda Connect pages through a table’s rows by filtering and sorting on the full primary key against the query’s own result set. If a primary key column isn’t projected, the snapshot fails part-way through, after the first batch of rows is read. **Type**: `object` ```yaml # Examples: snapshot_filters: TESTDB.PRODUCTS: SELECT * FROM TESTDB.PRODUCTS WHERE ID > 1000 TESTDB.USERS: SELECT * FROM TESTDB.USERS ``` ### [](#snapshot_max_batch_size)`snapshot_max_batch_size` The maximum number of rows to be streamed in a single batch when taking a snapshot. **Type**: `int` **Default**: `1000` ### [](#snapshot_mode)`snapshot_mode` Controls snapshot behavior. `none` (default) skips snapshotting and starts streaming from the current SCN. `snapshot_only` performs a full snapshot, persists the SCN checkpoint, then stops without streaming. `snapshot_and_stream` performs a full snapshot then transitions to streaming. **Type**: `string` **Options**: `none`, `snapshot_only`, `snapshot_and_stream` ### [](#wallet_password)`wallet_password` Password for the `ewallet.p12` PKCS#12 wallet file. Only use this when the wallet directory contains `ewallet.p12` rather than `cwallet.sso`. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#wallet_path)`wallet_path` Path to the Oracle Wallet directory. When set, this automatically enables SSL. The directory must contain either `cwallet.sso` (auto-login, does not require a password) or `ewallet.p12` (requires `wallet_password`). **Type**: `string` ```yaml # Examples: wallet_path: /opt/oracle/wallet ``` --- # Page 73: otlp_grpc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/otlp_grpc.md --- # otlp_grpc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: otlp_grpc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/otlp_grpc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/otlp_grpc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/otlp_grpc.adoc description: Receive OpenTelemetry traces, logs, and metrics via OTLP/gRPC protocol. page-git-created-date: "2026-01-23" page-git-modified-date: "2026-08-11" --- Receive OpenTelemetry traces, logs, and metrics via OTLP/gRPC protocol. Exposes an OpenTelemetry Collector gRPC receiver that accepts traces, logs, and metrics via gRPC. Telemetry data is received in OTLP protobuf format and converted to individual Redpanda OTEL v1 protobuf messages. Each signal (span, log record, or metric) becomes a separate message with embedded Resource and Scope metadata, optimized for Kafka partitioning. #### Common ```yml inputs: label: "" otlp_grpc: encoding: json address: 0.0.0.0:4317 rate_limit: "" ``` #### Advanced ```yml inputs: label: "" otlp_grpc: encoding: json address: 0.0.0.0:4317 tls: enabled: false cert_file: "" key_file: "" auth_token: "" max_recv_msg_size: 4194304 rate_limit: "" tcp: reuse_addr: false reuse_port: false schema_registry: url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} common_subject: "" trace_subject: "" log_subject: "" metric_subject: "" ``` ## [](#protocols)Protocols This input supports OTLP/gRPC on the default port 4317 using the standard OTLP protobuf format for all signal types (traces, logs, metrics). ## [](#output-format)Output format Each OTLP export request is unbatched into individual messages: - **Traces**: One message per span - **Logs**: One message per log record - **Metrics**: One message per metric Messages are encoded in Redpanda OTEL v1 protobuf format. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `signal_type` - The signal type: "trace", "log", or "metric" You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#authentication)Authentication When `auth_token` is configured, clients must include the token in the gRPC metadata. ### [](#go-client-example)Go client example ```go import ( "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc" ) exporter, err := otlptracegrpc.New(ctx, otlptracegrpc.WithEndpoint("localhost:4317"), otlptracegrpc.WithInsecure(), // or WithTLSCredentials() for TLS otlptracegrpc.WithHeaders(map[string]string{ "authorization": "Bearer your-token-here", }), ) ``` ### [](#environment-variable)Environment variable ```bash export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer your-token-here" ``` ## [](#rate-limiting)Rate limiting An optional rate limit resource can be specified to throttle incoming requests. When the rate limit is breached, requests will receive a ResourceExhausted gRPC status code. ## [](#fields)Fields ### [](#address)`address` The address to listen on for gRPC connections. **Type**: `string` **Default**: `0.0.0.0:4317` ### [](#auth_token)`auth_token` Optional bearer token for authentication. When set, requests must include 'authorization: Bearer ' metadata. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#encoding)`encoding` Encoding format for messages in the batch. Options: 'protobuf' or 'json'. **Type**: `string` **Default**: `json` **Options**: `protobuf`, `json` ### [](#max_recv_msg_size)`max_recv_msg_size` Maximum size of gRPC messages to receive in bytes. **Type**: `int` **Default**: `4194304` ### [](#rate_limit)`rate_limit` An optional rate limit resource to throttle requests. **Type**: `string` **Default**: `""` ### [](#schema_registry)`schema_registry` Optional Schema Registry configuration for adding Schema Registry wire format headers to messages. **Type**: `object` ### [](#schema_registry-basic_auth)`schema_registry.basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#schema_registry-basic_auth-enabled)`schema_registry.basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-basic_auth-password)`schema_registry.basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-basic_auth-username)`schema_registry.basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#schema_registry-common_subject)`schema_registry.common_subject` Schema subject name for the common protobuf schema. Only used when encoding is 'protobuf'. Defaults to 'redpanda-otel-common' for protobuf encoding or 'redpanda-otel-common-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt)`schema_registry.jwt` (beta) Allows you to specify JWT authentication. **Type**: `object` ### [](#schema_registry-jwt-claims)`schema_registry.jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-enabled)`schema_registry.jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-jwt-headers)`schema_registry.jwt.headers` Add optional key/value headers to the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-private_key_file)`schema_registry.jwt.private_key_file` A file with the PEM encoded via PKCS1 or PKCS8 as private key. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt-signing_method)`schema_registry.jwt.signing_method` A method used to sign the token such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#schema_registry-log_subject)`schema_registry.log_subject` Schema subject name for log data. Defaults to 'redpanda-otel-logs' for protobuf encoding or 'redpanda-otel-logs-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-metric_subject)`schema_registry.metric_subject` Schema subject name for metric data. Defaults to 'redpanda-otel-metrics' for protobuf encoding or 'redpanda-otel-metrics-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth)`schema_registry.oauth` Allows you to specify open authentication via OAuth version 1. **Type**: `object` ### [](#schema_registry-oauth-access_token)`schema_registry.oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-access_token_secret)`schema_registry.oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_key)`schema_registry.oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_secret)`schema_registry.oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-enabled)`schema_registry.oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-oauth2)`schema_registry.oauth2` Allows you to specify open authentication via OAuth version 2 using the client credentials token flow. **Type**: `object` ### [](#schema_registry-oauth2-client_key)`schema_registry.oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth2-client_secret)`schema_registry.oauth2.client_secret` A secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth2-enabled)`schema_registry.oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-oauth2-endpoint_params)`schema_registry.oauth2.endpoint_params` A list of optional endpoint parameters, values should be arrays of strings. **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: audience: - https://example.com resource: - https://api.example.com ``` ### [](#schema_registry-oauth2-scopes)`schema_registry.oauth2.scopes[]` A list of optional requested permissions. **Type**: `array` **Default**: `[]` ### [](#schema_registry-oauth2-token_url)`schema_registry.oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-timeout)`schema_registry.timeout` HTTP client timeout for Schema Registry requests. **Type**: `string` **Default**: `5s` ### [](#schema_registry-tls)`schema_registry.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#schema_registry-tls-client_certs)`schema_registry.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#schema_registry-tls-client_certs-cert)`schema_registry.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-cert_file)`schema_registry.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key)`schema_registry.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key_file)`schema_registry.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-password)`schema_registry.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#schema_registry-tls-enable_renegotiation)`schema_registry.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-enabled)`schema_registry.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-root_cas)`schema_registry.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#schema_registry-tls-root_cas_file)`schema_registry.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#schema_registry-tls-skip_cert_verify)`schema_registry.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#schema_registry-trace_subject)`schema_registry.trace_subject` Schema subject name for trace data. Defaults to 'redpanda-otel-traces' for protobuf encoding or 'redpanda-otel-traces-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-url)`schema_registry.url` Schema Registry URL for schema operations. **Type**: `string` ```yaml # Examples: url: http://localhost:8081 ``` ### [](#tcp)`tcp` TCP listener socket configuration. **Type**: `object` ### [](#tcp-reuse_addr)`tcp.reuse_addr` Enable SO\_REUSEADDR, allowing binding to ports in TIME\_WAIT state. Useful for graceful restarts and config reloads where the server needs to rebind to the same port immediately after shutdown. **Type**: `bool` **Default**: `false` ### [](#tcp-reuse_port)`tcp.reuse_port` Enable SO\_REUSEPORT, allowing multiple sockets to bind to the same port for load balancing across multiple processes/threads. **Type**: `bool` **Default**: `false` ### [](#tls)`tls` TLS configuration for gRPC. **Type**: `object` ### [](#tls-cert_file)`tls.cert_file` Path to the TLS certificate file. **Type**: `string` **Default**: `""` ### [](#tls-enabled)`tls.enabled` Enable TLS connections. **Type**: `bool` **Default**: `false` ### [](#tls-key_file)`tls.key_file` Path to the TLS key file. **Type**: `string` **Default**: `""` --- # Page 74: otlp_http **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/otlp_http.md --- # otlp_http > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: otlp_http latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/otlp_http page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/otlp_http.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/otlp_http.adoc description: Receive OpenTelemetry traces, logs, and metrics via OTLP/HTTP protocol. page-git-created-date: "2026-01-23" page-git-modified-date: "2026-08-11" --- Receive OpenTelemetry traces, logs, and metrics via OTLP/HTTP protocol. Exposes an OpenTelemetry Collector HTTP receiver that accepts traces, logs, and metrics via HTTP. Telemetry data is received in OTLP format (both protobuf and JSON) at standard OTLP endpoints and converted to individual Redpanda OTEL v1 protobuf messages. Each signal (span, log record, or metric) becomes a separate message with embedded Resource and Scope metadata, optimized for Kafka partitioning. #### Common ```yml inputs: label: "" otlp_http: encoding: json address: 0.0.0.0:4318 rate_limit: "" ``` #### Advanced ```yml inputs: label: "" otlp_http: encoding: json address: 0.0.0.0:4318 tls: enabled: false cert_file: "" key_file: "" auth_token: "" read_timeout: 10s write_timeout: 10s max_body_size: 4194304 rate_limit: "" tcp: reuse_addr: false reuse_port: false schema_registry: url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} common_subject: "" trace_subject: "" log_subject: "" metric_subject: "" ``` ## [](#endpoints)Endpoints This input exposes the following standard OTLP HTTP endpoints: - `/v1/traces` - OpenTelemetry traces - `/v1/logs` - OpenTelemetry logs - `/v1/metrics` - OpenTelemetry metrics ## [](#protocols)Protocols This input supports OTLP/HTTP on the default port 4318. It accepts both: - `application/x-protobuf` - OTLP protobuf format - `application/json` - OTLP JSON format ## [](#output-format)Output format Each OTLP export request is unbatched into individual messages: - **Traces**: One message per span - **Logs**: One message per log record - **Metrics**: One message per metric Messages are encoded in Redpanda OTEL v1 protobuf format. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `signal_type` - The signal type: "trace", "log", or "metric" You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#authentication)Authentication When `auth_token` is configured, clients must include the token in the HTTP Authorization header. ### [](#go-client-example)Go client example ```go import ( "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp" ) exporter, err := otlptracehttp.New(ctx, otlptracehttp.WithEndpoint("localhost:4318"), otlptracehttp.WithInsecure(), // or WithTLSClientConfig() for TLS otlptracehttp.WithHeaders(map[string]string{ "Authorization": "Bearer your-token-here", }), ) ``` ### [](#curl-example)cURL example ```bash curl -X POST http://localhost:4318/v1/traces \ -H "Content-Type: application/x-protobuf" \ -H "Authorization: Bearer your-token-here" \ --data-binary @traces.pb ``` ### [](#environment-variable)Environment variable ```bash export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer your-token-here" ``` ## [](#rate-limiting)Rate limiting An optional rate limit resource can be specified to throttle incoming requests. When the rate limit is breached, requests will receive a 429 (Too Many Requests) response. ## [](#fields)Fields ### [](#address)`address` The address to listen on for HTTP connections. **Type**: `string` **Default**: `0.0.0.0:4318` ### [](#auth_token)`auth_token` Optional bearer token for authentication. When set, requests must include 'Authorization: Bearer ' header. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#encoding)`encoding` Encoding format for messages in the batch. Options: 'protobuf' or 'json'. **Type**: `string` **Default**: `json` **Options**: `protobuf`, `json` ### [](#max_body_size)`max_body_size` Maximum size of HTTP request body in bytes. **Type**: `int` **Default**: `4194304` ### [](#rate_limit)`rate_limit` An optional rate limit resource to throttle requests. **Type**: `string` **Default**: `""` ### [](#read_timeout)`read_timeout` Maximum duration for reading the entire request. **Type**: `string` **Default**: `10s` ### [](#schema_registry)`schema_registry` Optional Schema Registry configuration for adding Schema Registry wire format headers to messages. **Type**: `object` ### [](#schema_registry-basic_auth)`schema_registry.basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#schema_registry-basic_auth-enabled)`schema_registry.basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-basic_auth-password)`schema_registry.basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-basic_auth-username)`schema_registry.basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#schema_registry-common_subject)`schema_registry.common_subject` Schema subject name for the common protobuf schema. Only used when encoding is 'protobuf'. Defaults to 'redpanda-otel-common' for protobuf encoding or 'redpanda-otel-common-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt)`schema_registry.jwt` (beta) Allows you to specify JWT authentication. **Type**: `object` ### [](#schema_registry-jwt-claims)`schema_registry.jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-enabled)`schema_registry.jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-jwt-headers)`schema_registry.jwt.headers` Add optional key/value headers to the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-private_key_file)`schema_registry.jwt.private_key_file` A file with the PEM encoded via PKCS1 or PKCS8 as private key. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt-signing_method)`schema_registry.jwt.signing_method` A method used to sign the token such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#schema_registry-log_subject)`schema_registry.log_subject` Schema subject name for log data. Defaults to 'redpanda-otel-logs' for protobuf encoding or 'redpanda-otel-logs-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-metric_subject)`schema_registry.metric_subject` Schema subject name for metric data. Defaults to 'redpanda-otel-metrics' for protobuf encoding or 'redpanda-otel-metrics-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth)`schema_registry.oauth` Allows you to specify open authentication via OAuth version 1. **Type**: `object` ### [](#schema_registry-oauth-access_token)`schema_registry.oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-access_token_secret)`schema_registry.oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_key)`schema_registry.oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_secret)`schema_registry.oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-enabled)`schema_registry.oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-oauth2)`schema_registry.oauth2` Allows you to specify open authentication via OAuth version 2 using the client credentials token flow. **Type**: `object` ### [](#schema_registry-oauth2-client_key)`schema_registry.oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth2-client_secret)`schema_registry.oauth2.client_secret` A secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth2-enabled)`schema_registry.oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-oauth2-endpoint_params)`schema_registry.oauth2.endpoint_params` A list of optional endpoint parameters, values should be arrays of strings. **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: audience: - https://example.com resource: - https://api.example.com ``` ### [](#schema_registry-oauth2-scopes)`schema_registry.oauth2.scopes[]` A list of optional requested permissions. **Type**: `array` **Default**: `[]` ### [](#schema_registry-oauth2-token_url)`schema_registry.oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-timeout)`schema_registry.timeout` HTTP client timeout for Schema Registry requests. **Type**: `string` **Default**: `5s` ### [](#schema_registry-tls)`schema_registry.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#schema_registry-tls-client_certs)`schema_registry.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#schema_registry-tls-client_certs-cert)`schema_registry.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-cert_file)`schema_registry.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key)`schema_registry.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key_file)`schema_registry.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-password)`schema_registry.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#schema_registry-tls-enable_renegotiation)`schema_registry.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-enabled)`schema_registry.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-root_cas)`schema_registry.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#schema_registry-tls-root_cas_file)`schema_registry.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#schema_registry-tls-skip_cert_verify)`schema_registry.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#schema_registry-trace_subject)`schema_registry.trace_subject` Schema subject name for trace data. Defaults to 'redpanda-otel-traces' for protobuf encoding or 'redpanda-otel-traces-json' for JSON encoding. **Type**: `string` **Default**: `""` ### [](#schema_registry-url)`schema_registry.url` Schema Registry URL for schema operations. **Type**: `string` ```yaml # Examples: url: http://localhost:8081 ``` ### [](#tcp)`tcp` TCP listener socket configuration. **Type**: `object` ### [](#tcp-reuse_addr)`tcp.reuse_addr` Enable SO\_REUSEADDR, allowing binding to ports in TIME\_WAIT state. Useful for graceful restarts and config reloads where the server needs to rebind to the same port immediately after shutdown. **Type**: `bool` **Default**: `false` ### [](#tcp-reuse_port)`tcp.reuse_port` Enable SO\_REUSEPORT, allowing multiple sockets to bind to the same port for load balancing across multiple processes/threads. **Type**: `bool` **Default**: `false` ### [](#tls)`tls` TLS configuration for HTTP. **Type**: `object` ### [](#tls-cert_file)`tls.cert_file` Path to the TLS certificate file. **Type**: `string` **Default**: `""` ### [](#tls-enabled)`tls.enabled` Enable TLS connections. **Type**: `bool` **Default**: `false` ### [](#tls-key_file)`tls.key_file` Path to the TLS key file. **Type**: `string` **Default**: `""` ### [](#write_timeout)`write_timeout` Maximum duration for writing the response. **Type**: `string` **Default**: `10s` --- # Page 75: postgres_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/postgres_cdc.md --- # postgres_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: postgres_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/postgres_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/postgres_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/postgres_cdc.adoc page-git-created-date: "2024-12-05" page-git-modified-date: "2026-05-26" --- Streams data changes from a PostgreSQL database using logical replication. There is also a configuration option to [stream all existing data](#stream_snapshot) from the database. ```yml inputs: label: "" postgres_cdc: dsn: "" # No default (required) include_transaction_markers: false stream_snapshot: false snapshot_batch_size: 1000 schema: "" # No default (required) tables: [] # No default (required) checkpoint_limit: 1024 temporary_slot: false slot_name: "" # No default (required) pg_standby_timeout: 10s pg_wal_monitor_interval: 3s max_parallel_snapshot_tables: 1 auto_replay_nacks: true batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` The `postgres_cdc` input uses logical replication to capture changes made to a PostgreSQL database in real time and streams them to Redpanda Connect. Redpanda Connect uses this replication method to allow you to choose which database tables in your source database to receive changes from. There are also [two replication modes](#choose-a-replication-mode) to choose from, and an [option to receive TOAST and deleted values](#receive-toast-and-deleted-values) in your data updates. ## [](#prerequisites)Prerequisites - PostgreSQL version 14 or later - Network access from the cluster where your Redpanda Connect pipeline is running to the source database environment. For detailed networking information, including how to set up a VPC peering connection, see [Redpanda Cloud Networking](https://docs.redpanda.com/cloud-data-platform/networking/). - Logical replication enabled on your PostgreSQL cluster To check whether logical replication is already enabled, run the following query: ```SQL SHOW wal_level; ``` If the `wal_level` value is `logical`, you can start to use this connector. Otherwise, choose from the following sets of instructions to update your replication settings. ### Cloud platforms - [Amazon RDS for PostgreSQL DB](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/PostgreSQL.Concepts.General.FeatureSupport.LogicalReplication.html) - [Azure Database for PostgreSQL](https://learn.microsoft.com/en-us/azure/postgresql/flexible-server/concepts-logical#prerequisites-for-logical-replication-and-logical-decoding) - [Google Cloud SQL for PostgreSQL](https://cloud.google.com/sql/docs/postgres/replication/configure-logical-replication), including creating a user with replication privileges - [Neon](https://neon.tech/docs/guides/logical-replication-guide) ### Self-Hosted PostgreSQL Use an account with sufficient permissions (superuser) to update your replication settings. 1. Open the `postgresql.conf` file. 2. Find the `wal_level` parameter. 3. Update the parameter value to `wal_level = logical`. If you already use replication slots, you may need to increase the limit on replication slots (`max_replication_slots`). The `max_wal_senders` parameter value must also be greater than or equal to `max_replication_slots`. 4. Restart the PostgreSQL server. For this input to make a successful connection to your database, also make sure that it allows replication connections. 1. Open the `pg_hba.conf` file. 2. Update this line. ```yaml host replication /32 md5 ``` Replace the following placeholders with your own values: - ``: The username from an account with superuser privileges. - ``: The IP address of the server where you are running Redpanda Connect. 3. Restart the PostgreSQL server. ## [](#choose-a-replication-mode)Choose a replication mode When you run a pipeline that uses the `postgres_cdc` input, Redpanda Connect connects to your PostgreSQL database and creates a replication slot. The replication slot uses a copy of the Write-Ahead Log (WAL) file to subscribe to changes in your database records as they are applied to the database. There are two replication modes you can choose from: snapshot mode and streaming mode. In snapshot mode, Redpanda Connect first takes a snapshot of the database and streams the contents before processing changes from the WAL. In streaming mode, Redpanda Connect directly processes changes from the WAL starting from the most recent changes without taking a snapshot first. For local testing, you can use the [example pipeline on this page](#example-pipeline), which runs in snapshot mode. ### [](#snapshot-mode)Snapshot mode If you set the [`stream_snapshot` field](#stream_snapshot) to `true`, Redpanda Connect: 1. Creates a snapshot of your database. 2. Streams the contents of the tables specified in the `postgres_cdc` input. 3. Starts processing changes in the WAL that occurred since the snapshot was taken, and streams them to Redpanda Connect. Once the initial replication process is complete, the snapshot is removed and the input keeps a connection open to the database so that it can receive data updates. If the pipeline restarts during the replication process, Redpanda Connect resumes processing data changes from where it left off. If there are other interruptions while the snapshot is taken, you may need to restart the snapshot process. For more information, see [Troubleshoot replication failures](#troubleshoot_replication_failures). ### [](#streaming-mode)Streaming mode If you set the [`stream_snapshot` field](#stream_snapshot) to `false`, Redpanda Connect starts processing data changes from the end of the WAL. If the pipeline restarts, Redpanda Connect resumes processing data changes from the last acknowledged position in the WAL. ## [](#monitor-the-replication-process)Monitor the replication process You can monitor the initial replication of data using the following metrics: | Metric name | Description | | --- | --- | | replication_lag_bytes | Indicates how far the connector is lagging behind the source database when processing the transaction log. | | postgres_snapshot_progress | Shows the progress of snapshot processing for each table. | ## [](#troubleshoot-replication-failures)Troubleshoot replication failures If the database snapshot fails, the replication slot has only an incomplete record of the existing data in your database. To maintain data integrity, you must drop the replication slot manually in your source database and run the Redpanda Connect pipeline again. ```SQL SELECT pg_drop_replication_slot(SLOT_NAME); ``` ## [](#receive-toast-and-deleted-values)Receive TOAST and deleted values For full visibility of all data updates, you can also choose to stream [TOAST](https://www.postgresql.org/docs/current/storage-toast.html) and deleted values. To enable this option, run the following query on your source database: ```SQL ALTER TABLE large_data REPLICA IDENTITY FULL; ``` ## [](#data-mappings)Data mappings The following table shows how selected PostgreSQL data types are mapped to data types supported in Redpanda Connect. All other data types are mapped to string values. | PostgreSQL data type | Bloblang value | | --- | --- | | TEXT, TIMESTAMP, UUID, VARCHAR | JSON strings, for example: this data | | BOOL | Boolean JSON fields, for example: true or false | | Numeric types (INT4) | JSON number types, for example: 1. | | JSONB | JSON objects, for example: { "message": "message text" } | | INTEGER[] | An array of integer values, for example: [1,2,3] | | TEXT[] | An array of string values, for example: ["value1", "value2", "value3"] | | INET | A string that contains an IP address, for example: "192.168.1.1" | | POINT | A string that represents a point in a two-dimensional plane, for example: (x, y) | | TSRANGE | A string that includes range bounds, for example: [2010-01-01 14:30, 2010-01-01 15:30) | | TSVECTOR | A string that includes vector data, for example: "'the':2 'question':3 'is':4" | ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `table`: The name of the database table from which the message originated. - `operation`: The type of database operation that generated the message, such as `read`, `insert`, `update`, `delete`, `begin` and `commit`. A `read` operation occurs when a snapshot of the database is processed. The `begin` and `commit` operations are only included if the `include_transaction_markers` field is set to `true`. - `lsn`: The [Log Sequence Number](https://www.postgresql.org/docs/current/datatype-pg-lsn.html) of each data update from the source PostgreSQL database. The `lsn` values are strings that can be sorted to determine the order in which data updates were written to the WAL. ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay rejected messages (negative acknowledgements) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#aws)`aws` AWS IAM authentication configuration for PostgreSQL instances. When enabled, IAM credentials are used to generate temporary authentication tokens instead of a static password. This is useful for connecting to Amazon RDS or Aurora PostgreSQL instances with IAM database authentication enabled. The generated tokens are valid for 15 minutes and are automatically refreshed. For more information about AWS credentials configuration, see the [credentials for AWS](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/) guide. **Type**: `object` ### [](#aws-enabled)`aws.enabled` Enable AWS IAM authentication for PostgreSQL. When enabled, an IAM authentication token is generated and used as the password. **Type**: `bool` **Default**: `false` ### [](#aws-endpoint)`aws.endpoint` The PostgreSQL endpoint hostname (e.g., mydb.abc123.us-east-1.rds.amazonaws.com). **Type**: `string` ### [](#aws-id)`aws.id` The ID of credentials to use. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region where the PostgreSQL instance is located. If no region is specified then the environment default will be used. **Type**: `string` ### [](#aws-role)`aws.role` Optional AWS IAM role ARN to assume for authentication. Alternatively, use `roles` array for role chaining instead. **Type**: `string` ### [](#aws-role_external_id)`aws.role_external_id` Optional external ID for the role assumption. Only used with the `role` field. Alternatively, use `roles` array for role chaining instead. **Type**: `string` ### [](#aws-roles)`aws.roles[]` Optional array of AWS IAM roles to assume for authentication. Roles can be assumed in sequence, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. **Type**: `array` ### [](#aws-roles-role)`aws.roles[].role` AWS IAM role ARN to assume. **Type**: `string` **Default**: `""` ### [](#aws-roles-role_external_id)`aws.roles[].role_external_id` Optional external ID for the role assumption. **Type**: `string` **Default**: `""` ### [](#aws-secret)`aws.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-token)`aws.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_limit)`checkpoint_limit` The maximum number of messages that this input can process at a given time. Increasing this limit enables parallel processing, and batching at the output level. To preserve at-least-once guarantees, any given log sequence number (LSN) is not acknowledged until all messages under that offset are delivered. **Type**: `int` **Default**: `1024` ### [](#dsn)`dsn` The data source name (DSN) of the PostgreSQL database from which you want to stream updates. Use the format `postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​]`. For example, if you wanted to disable SSL in a secure environment, you would add `sslmode=disable` to the connection string. **Type**: `string` ```yaml # Examples: dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable ``` ### [](#heartbeat_interval)`heartbeat_interval` The interval between heartbeat messages, which Redpanda Connect writes to the WAL using the `pg_logical_emit_message` function. Heartbeat messages are useful when you subscribe to data changes from tables with low activity, while other tables in the database have higher-frequency updates. Heartbeat messages allow Redpanda Connect to periodically acknowledge new messages even when no data updates occur. Each acknowledgement advances the committed point in the WAL, which ensures that PostgreSQL can safely reclaim older log segments, preventing excessive disk space usage. Set `heartbeat_interval` to `0s` to disable heartbeats. **Type**: `string` **Default**: `1h` ```yaml # Examples: heartbeat_interval: 0s # --- heartbeat_interval: 24h ``` ### [](#include_transaction_markers)`include_transaction_markers` When set to `true`, creates empty messages for `BEGIN` and `COMMIT` operations which start and complete each transaction. Messages with the `operation` metadata field set to `BEGIN` or `COMMIT` have null message payloads. **Type**: `bool` **Default**: `false` ### [](#max_parallel_snapshot_tables)`max_parallel_snapshot_tables` Specify the maximum number of tables that are processed in parallel when the initial snapshot of the source database is taken. **Type**: `int` **Default**: `1` ### [](#pg_standby_timeout)`pg_standby_timeout` Specify the standby timeout after which an idle connection is refreshed to keep the connection alive. **Type**: `string` **Default**: `10s` ```yaml # Examples: pg_standby_timeout: 30s ``` ### [](#pg_wal_monitor_interval)`pg_wal_monitor_interval` How often to report changes to the replication lag and write them to Redpanda Connect metrics. **Type**: `string` **Default**: `3s` ```yaml # Examples: pg_wal_monitor_interval: 6s ``` ### [](#schema)`schema` The PostgreSQL schema from which to replicate data. **Type**: `string` ```yaml # Examples: schema: public # --- schema: "MyCaseSensitiveSchemaNeedingQuotes" ``` ### [](#signal_table_name)`signal_table_name` The name of the table used to send control signals to the connector, excluding the schema. The table must exist in the schema configured via the `schema` field, and must not also appear in `tables`, since the signal table is implicitly added to the publication and excluded from snapshot scans, so listing it in both places is rejected at startup. It must have at least these columns; startup validation checks column names only, not types, so a wrong column type (for example `data JSONB` instead of `TEXT`) is only caught at runtime, on the first signal row read: - **id**: any type representable as a string (for example `SERIAL`, `BIGSERIAL`, `UUID`, `VARCHAR`) - **type**: the signal type (see supported signals below). Should be `VARCHAR` or another string type. - **data**: a JSON object containing signal parameters. Should be `TEXT`. Create the table with: ```sql CREATE TABLE . ( id SERIAL PRIMARY KEY, type VARCHAR(32), data TEXT ); ``` Signal rows are published as regular output messages (`operation=insert`, `table=`). To exclude them from downstream processing, filter on the `table` metadata field using a `mapping` processor: ```yaml pipeline: processors: - mapping: | root = if @table == "rpcn_signal_table" { deleted() } else { this } ``` **Supported signals** **`log`**: recognized and logged when received. The `data` column must contain a JSON object with a `message` key, whose value is written to the connector’s log output. ```sql INSERT INTO . (type, data) VALUES ('log', '{"message": "Signal message"}'); ``` **Type**: `string` **Default**: `""` ```yaml # Examples: signal_table_name: rpcn_signal_table ``` ### [](#slot_name)`slot_name` The name of the PostgreSQL logical replication slot to use. If not provided, a random name is generated unless you create a replication slot manually before starting replication. **Type**: `string` ```yaml # Examples: slot_name: my_test_slot ``` ### [](#snapshot_batch_size)`snapshot_batch_size` The number of table rows to fetch in each batch when querying the snapshot. This option is only available when `stream_snapshot` is set to `true`. **Type**: `int` **Default**: `1000` ```yaml # Examples: snapshot_batch_size: 10000 ``` ### [](#stream_snapshot)`stream_snapshot` When set to `true`, this input streams a snapshot of all existing data in the source database before streaming data changes. To use this setting, all database tables that you want to replicate _must_ have a primary key. **Type**: `bool` **Default**: `false` ```yaml # Examples: stream_snapshot: true ``` ### [](#tables)`tables[]` A list of database table names to include in the snapshot and logical replication. Specify each table name as a separate item. **Type**: `array` ```yaml # Examples: tables: - my_table_1 - "MyCaseSensitiveTableNeedingQuotes" ``` ### [](#temporary_slot)`temporary_slot` If set to `true`, the input creates a temporary replication slot that is automatically dropped when the connection to your source database is closed. You might use this option to: - Avoid data accumulating in the replication slot when a pipeline is paused or stopped - Test the connector If the pipeline is restarted, another data snapshot is taken before data updates are streamed. **Type**: `bool` **Default**: `false` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#unchanged_toast_value)`unchanged_toast_value` Specify the value to emit when unchanged [TOAST values](#receive-toast-and-deleted-values) appear in the message stream. Unchanged values occur for data updates and deletes when `REPLICA IDENTITY` is not set to `FULL`. **Type**: `unknown` **Default**: ```yaml null ``` ```yaml # Examples: unchanged_toast_value: __redpanda_connect_unchanged_toast_value__ ``` ## [](#example-pipeline)Example pipeline You can run the following pipeline locally to check that data updates are streamed from your source database to Redpanda Connect. All transactions are written to stdout. ```yml input: label: "postgres_cdc" postgres_cdc: dsn: postgres://user:password@host:port/dbname include_transaction_markers: false slot_name: test_slot_native_decoder snapshot_batch_size: 100000 stream_snapshot: true temporary_slot: true schema: schema_name tables: - table_name cache_resources: - label: data_caching file: directory: /tmp/cache output: label: main stdout: {} ``` --- # Page 76: read_until **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/read_until.md --- # read_until > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: read_until latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/read_until page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/read_until.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/read_until.adoc description: Reads messages from a child input until a consumed message passes a Bloblang query, at which point the input closes. It is also possible to configure a timeout after which the input is closed if no new messages arrive in that period. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Reads messages from a child input until a consumed message passes a [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/), at which point the input closes. It is also possible to configure a timeout after which the input is closed if no new messages arrive in that period. ```yml inputs: label: "" read_until: input: "" # No default (required) check: "" # No default (optional) idle_timeout: "" # No default (optional) restart_input: false ``` Messages are read continuously while the query check returns false, when the query returns true the message that triggered the check is sent out and the input is closed. Use this to define inputs where the stream should end once a certain message appears. If the idle timeout is configured, the input will be closed if no new messages arrive after that period of time. Use this field if you want to empty out and close an input that doesn’t have a logical end. Sometimes inputs close themselves. For example, when the `file` input type reaches the end of a file it will shut down. By default this type will also shut down. If you wish for the input type to be restarted every time it shuts down until the query check is met then set `restart_input` to `true`. ## [](#metadata)Metadata A metadata key `benthos_read_until` containing the value `final` is added to the first part of the message that triggers the input to stop. ## [](#fields)Fields ### [](#check)`check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether the input should now be closed. **Type**: `string` ```yaml # Examples: check: this.type == "foo" # --- check: count("messages") >= 100 ``` ### [](#idle_timeout)`idle_timeout` The maximum amount of time without receiving new messages after which the input is closed. **Type**: `string` ```yaml # Examples: idle_timeout: 5s ``` ### [](#input)`input` The child input to consume from. **Type**: `input` ### [](#restart_input)`restart_input` Whether the input should be reopened if it closes itself before the condition has resolved to true. **Type**: `bool` **Default**: `false` ## [](#examples)Examples ### [](#consume-n-messages)Consume N Messages A common reason to use this input is to consume only N messages from an input and then stop. This can easily be done with the [`count` function](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#count): ```yaml # Only read 100 messages, and then exit. input: read_until: check: count("messages") >= 100 input: kafka: addresses: [ TODO ] topics: [ foo, bar ] consumer_group: foogroup ``` ### [](#read-from-a-kafka-and-close-when-empty)Read from a kafka and close when empty A common reason to use this input is a job that consumes all messages and exits once its empty: ```yaml # Consumes all messages and exit when the last message was consumed 5s ago. input: read_until: idle_timeout: 5s input: kafka: addresses: [ TODO ] topics: [ foo, bar ] consumer_group: foogroup ``` --- # Page 77: redis_list **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redis_list.md --- # redis_list > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_list latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/redis_list page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/redis_list.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/redis_list.adoc description: Pops messages from the beginning of a Redis list using the BLPop command. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Pops messages from the beginning of a Redis list using the BLPop command. #### Common ```yml inputs: label: "" redis_list: url: "" # No default (required) key: "" # No default (required) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" redis_list: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] key: "" # No default (required) auto_replay_nacks: true max_in_flight: 0 timeout: 5s command: blpop ``` ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#command)`command` The command used to pop elements from the Redis list **Type**: `string` **Default**: `blpop` **Options**: `blpop`, `brpop` ### [](#key)`key` The key of a list to read from. **Type**: `string` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#max_in_flight)`max_in_flight` Optionally sets a limit on the number of messages that can be flowing through a Redpanda Connect stream pending acknowledgment from the input at any given time. Once a message has been either acknowledged or rejected (nacked) it is no longer considered pending. If the input produces logical batches then each batch is considered a single count against the maximum. **WARNING**: Batching policies at the output level will stall if this field limits the number of messages below the batching threshold. Zero (default) or lower implies no limit. **Type**: `int` **Default**: `0` ### [](#timeout)`timeout` The length of time to poll for new messages before reattempting. **Type**: `string` **Default**: `5s` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 78: redis_pubsub **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redis_pubsub.md --- # redis_pubsub > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_pubsub latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/redis_pubsub page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/redis_pubsub.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/redis_pubsub.adoc description: Consume from a Redis publish/subscribe channel using either the SUBSCRIBE or PSUBSCRIBE commands. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consume from a Redis publish/subscribe channel using either the SUBSCRIBE or PSUBSCRIBE commands. #### Common ```yml inputs: label: "" redis_pubsub: url: "" # No default (required) channels: [] # No default (required) use_patterns: false auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" redis_pubsub: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] channels: [] # No default (required) use_patterns: false auto_replay_nacks: true ``` In order to subscribe to channels using the `PSUBSCRIBE` command set the field `use_patterns` to `true`, then you can include glob-style patterns in your channel names. For example: - `h?llo` subscribes to hello, hallo and hxllo - `h*llo` subscribes to hllo and heeeello - `h[ae]llo` subscribes to hello and hallo, but not hillo Use `\` to escape special characters if you want to match them verbatim. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `redis_pubsub_channel` - `redis_pubsub_pattern` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#channels)`channels[]` A list of channels to consume from. **Type**: `array` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` ### [](#use_patterns)`use_patterns` Whether to use the PSUBSCRIBE command, allowing for glob-style patterns within target channel names. **Type**: `bool` **Default**: `false` --- # Page 79: redis_scan **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redis_scan.md --- # redis_scan > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_scan latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/redis_scan page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/redis_scan.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/redis_scan.adoc description: Scans the set of keys in the current selected database and gets their values, using the Scan and Get commands. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Scans the set of keys in the current selected database and gets their values, using the Scan and Get commands. #### Common ```yml inputs: label: "" redis_scan: url: "" # No default (required) auto_replay_nacks: true match: "" ``` #### Advanced ```yml inputs: label: "" redis_scan: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] auto_replay_nacks: true match: "" ``` Optionally, iterates only elements matching a blob-style pattern. For example: - `**foo**` iterates only keys which contain `foo` in it. - `foo*` iterates only keys starting with `foo`. This input generates a message for each key value pair in the following format: ```json {"key":"foo","value":"bar"} ``` ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#match)`match` Iterates only elements matching the optional glob-style pattern. By default, it matches all elements. **Type**: `string` **Default**: `""` ```yaml # Examples: match: * # --- match: 1* # --- match: foo* # --- match: foo # --- match: *4* ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 80: redis_streams **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redis_streams.md --- # redis_streams > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_streams latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/redis_streams page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/redis_streams.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/redis_streams.adoc description: Pulls messages from Redis (v5.0+) streams with the XREADGROUP command. The client_id should be unique for each consumer of a group. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Pulls messages from Redis (v5.0+) streams with the XREADGROUP command. The `client_id` should be unique for each consumer of a group. #### Common ```yml inputs: label: "" redis_streams: url: "" # No default (required) body_key: body streams: [] # No default (required) auto_replay_nacks: true limit: 10 client_id: "" consumer_group: "" ``` #### Advanced ```yml inputs: label: "" redis_streams: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] body_key: body streams: [] # No default (required) auto_replay_nacks: true limit: 10 client_id: "" consumer_group: "" create_streams: true start_from_oldest: true commit_period: 1s timeout: 1s ``` Redis stream entries are key/value pairs, as such it is necessary to specify the key that contains the body of the message. All other keys/value pairs are saved as metadata fields. ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#body_key)`body_key` The field key to extract the raw message from. All other keys will be stored in the message as metadata. **Type**: `string` **Default**: `body` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `""` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#commit_period)`commit_period` The period of time between each commit of the current offset. Offsets are always committed during shutdown. **Type**: `string` **Default**: `1s` ### [](#consumer_group)`consumer_group` An identifier for the consumer group of the stream. **Type**: `string` **Default**: `""` ### [](#create_streams)`create_streams` Create subscribed streams if they do not exist (MKSTREAM option). **Type**: `bool` **Default**: `true` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#limit)`limit` The maximum number of messages to consume from a single request. **Type**: `int` **Default**: `10` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#start_from_oldest)`start_from_oldest` If an offset is not found for a stream, determines whether to consume from the oldest available offset, otherwise messages are consumed from the latest offset. **Type**: `bool` **Default**: `true` ### [](#streams)`streams[]` A list of streams to consume from. **Type**: `array` ### [](#timeout)`timeout` The length of time to poll for new messages before reattempting. **Type**: `string` **Default**: `1s` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 81: redpanda_common **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda_common.md --- # redpanda_common > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda_common latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/redpanda_common page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/redpanda_common.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/redpanda_common.adoc description: Consumes data from a Redpanda (Kafka) broker, using credentials defined in a common top-level redpanda config block. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- > ⚠️ **WARNING: Deprecated in 4.68.0** > > Deprecated in 4.68.0 > > This component is deprecated and will be removed in the next major version release. Please consider moving onto the unified [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [`redpanda` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) components. Consumes data from a Redpanda (Kafka) broker, using credentials from a common `redpanda` configuration block. To avoid duplicating Redpanda cluster credentials in your `redpanda_common` input, output, or any other components in your data pipeline, you can use a single [`redpanda` configuration block](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/redpanda/about/). For more details, see the [Pipeline example](#pipeline-example). > 📝 **NOTE** > > If you need to move topic data between Redpanda clusters or other Apache Kafka clusters, consider using the [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) instead. #### Common ```yml inputs: label: "" redpanda_common: topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" redpanda_common: topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) rack_id: "" instance_id: "" rebalance_timeout: 45s session_timeout: 1m heartbeat_interval: 3s start_offset: earliest fetch_max_bytes: 50MiB fetch_max_wait: 5s fetch_min_bytes: 1B fetch_max_partition_bytes: 1MiB transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) commit_period: 5s partition_buffer_bytes: 1MB topic_lag_refresh_period: 5s max_yield_batch_bytes: 32KB auto_replay_nacks: true timely_nacks_maximum_wait: "" # No default (optional) ``` ## [](#pipeline-example)Pipeline example This data pipeline reads data from `topic_A` and `topic_B` on a Redpanda cluster, and then writes the data to `topic_C` on the same cluster. The cluster details are configured within the `redpanda` configuration block, so you only need to configure them once. This is a useful feature when you have multiple inputs and outputs in the same data pipeline that need to connect to the same cluster. ```none input: redpanda_common: topics: [ topic_A, topic_B ] output: redpanda_common: topic: topic_C key: ${! @id } redpanda: seed_brokers: [ "127.0.0.1:9092" ] tls: enabled: true sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ## [](#consumer-groups)Consumer groups When you specify a consumer group in your configuration, this input consumes one or more topics and automatically balances the topic partitions across any other connected clients with the same consumer group. Otherwise, topics are consumed in their entirety or with explicit partitions. ### [](#delivery-guarantees)Delivery guarantees If you choose to use consumer groups, the offsets of records received by Redpanda Connect are committed automatically. In the event of restarts, this input uses the committed offsets to resume data consumption where it left off. Redpanda Connect guarantees at-least-once delivery. Records are only confirmed as delivered when all downstream outputs that a record is routed to have also confirmed delivery. ## [](#ordering)Ordering To preserve the order of topic partitions: - Records consumed from each partition are processed and delivered in the order that they are received - Only one batch of records of a given partition is processed at a time This approach means that although records from different partitions may be processed in parallel, records from the same partition are processed in sequential order. ### [](#delivery-errors)Delivery errors The order in which records are delivered may be disrupted by delivery errors and any error-handling mechanisms that start up. Redpanda Connect uses at-least-once delivery unless instructed otherwise, and this includes reattempting delivery of data when the ordering of that data is no longer guaranteed. For example, a batch of records is sent to an output broker and only a subset of records are delivered. In this scenario, Redpanda Connect (by default) attempts to deliver the records that failed, even though these delivery failures may have been sent before records that were delivered successfully. #### [](#use-a-fallback-output)Use a fallback output To prevent delivery errors from disrupting the order of records, you must specify a [`fallback`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/) output in your pipeline configuration. When adding a `fallback` output, it is good practice to set the `auto_retry_nacks` field to `false`. This also improves the throughput of your pipeline. For example, the following configuration includes a `fallback` output. If Redpanda Connect fails to write delivery errors to the `foo` topic, it then attempts to write them into a dead letter queue topic (`foo_dlq`), which is retried indefinitely as a way to apply back pressure. ```yaml output: fallback: - redpanda_common: topic: foo - retry: output: redpanda_common: topic: foo_dlq ``` ## [](#batching)Batching Records are processed and delivered from each partition in the same batches as they are received from brokers. Batch sizes are dynamically sized in order to optimize throughput, but you can tune them further using the following configuration fields: - `fetch_max_partition_bytes` - `fetch_max_bytes` You can break batches down further using the [`split`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/split/) processor. ## [](#metrics)Metrics This input emits a `redpanda_lag` metric with `topic` and `partition` labels for each consumed topic. The metric records the number of produced messages that remain to be read from each topic/partition pair by the specified consumer group. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `kafka_key` - `kafka_topic` - `kafka_partition` - `kafka_offset` - `kafka_lag` - `kafka_timestamp_ms` - `kafka_timestamp_unix` - `kafka_tombstone_message` - All record headers ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#commit_period)`commit_period` The period of time between each commit of the current partition offsets. Offsets are always committed during shutdown. **Type**: `string` **Default**: `5s` ### [](#consumer_group)`consumer_group` An optional consumer group. When this value is specified: - The partitions of any topics, specified in the `topics` field, are automatically distributed across consumers sharing a consumer group - Partition offsets are automatically committed and resumed under this name Consumer groups are not supported when you specify explicit partitions to consume from in the `topics` field. **Type**: `string` ### [](#fetch_max_bytes)`fetch_max_bytes` The maximum number of bytes that a broker tries to send during a fetch. If individual records are larger than the `fetch_max_bytes` value, brokers will still send them. **Type**: `string` **Default**: `50MiB` ### [](#fetch_max_partition_bytes)`fetch_max_partition_bytes` The maximum number of bytes that are consumed from a single partition in a fetch request. This field is equivalent to the Java setting `fetch.max.partition.bytes`. If a single batch is larger than the `fetch_max_partition_bytes` value, the batch is still sent so that the client can make progress. **Type**: `string` **Default**: `1MiB` ### [](#fetch_max_wait)`fetch_max_wait` The maximum period of time a broker can wait for a fetch response to reach the required minimum number of bytes (`fetch_min_bytes`). **Type**: `string` **Default**: `5s` ### [](#fetch_min_bytes)`fetch_min_bytes` The minimum number of bytes that a broker tries to send during a fetch. This field is equivalent to the Java setting `fetch.min.bytes`. **Type**: `string` **Default**: `1B` ### [](#heartbeat_interval)`heartbeat_interval` When you specify a `consumer_group`, `heartbeat_interval` sets how frequently a consumer group member should send heartbeats to Apache Kafka. Apache Kafka uses heartbeats to make sure that a group member’s session is active. You must set `heartbeat_interval` to less than one-third of `session_timeout`. This field is equivalent to the Java `heartbeat.interval.ms` setting and accepts Go duration format strings such as `10s` or `2m`. **Type**: `string` **Default**: `3s` ### [](#instance_id)`instance_id` When you specify a [`consumer_group`](#consumer_group), assign a unique value to `instance_id` to define the group’s static membership, which can prevent unnecessary rebalances during reconnections. When you assign an instance ID, the client does not automatically leave the consumer group when it disconnects. To remove the client, you must use an external admin command on behalf of the instance ID. **Type**: `string` **Default**: `""` ### [](#max_yield_batch_bytes)`max_yield_batch_bytes` The maximum size (in bytes) for each batch yielded by this input. This value must be less than or equal to the `partition_buffer_bytes`. If using Redpanda output, this value should not be greater than the `max_message_bytes` option value (1MB by default), and for high-throughput scenarios they should be equal. **Type**: `string` **Default**: `32KB` ### [](#partition_buffer_bytes)`partition_buffer_bytes` A buffer size (in bytes) for each consumed partition, which allows the internal queuing of records before they are flushed. Increasing this value may improve throughput but results in higher memory utilization. Each buffer can grow slightly beyond this value. **Type**: `string` **Default**: `1MB` ### [](#rack_id)`rack_id` A rack specifies where the client is physically located, and changes fetch requests to consume from the closest replica as opposed to the leader replica. **Type**: `string` **Default**: `""` ### [](#rebalance_timeout)`rebalance_timeout` When you specify a [`consumer_group`](#consumer_group), `rebalance_timeout` sets a time limit for all consumer group members to complete their work and commit offsets after a rebalance has begun. The timeout excludes the time taken to detect a failed or late heartbeat, which indicates a rebalance is required. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `45s` ### [](#regexp_topics_exclude)`regexp_topics_exclude[]` A list of regular expression patterns for excluding topics when regex mode is enabled (using `regexp_topics_include` or the deprecated `regexp_topics` boolean). Topics matching any of these patterns will be excluded from consumption, even if they match include patterns. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so use `^` and `$` for exact matching. Exclude patterns are applied after include patterns, providing fine-grained control over topic selection. Example: `regexp_topics_exclude: ["^_", ".**-temp$", ".**-test.*"]` excludes topics starting with underscore, ending with `-temp`, or containing `-test`. **Type**: `array` ### [](#regexp_topics_include)`regexp_topics_include[]` A list of regular expression patterns for matching topics to consume from. When specified, the client will periodically refresh the list of matching topics based on the `metadata_max_age` interval. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so `logs_.` **matches `my-logs_events` and `logs_errors`. Use `^logs_.`**`$` to match only topics starting with `logs_`. This field enables regex mode (replacing the deprecated `regexp_topics` boolean) and cannot be used together with explicit `topics` lists. Use `regexp_topics_exclude` to filter out specific patterns from the matched topics. Example: `regexp_topics_include: ["events_.**", "logs_.**"]` consumes from all topics starting with `events_` or `logs_`. **Type**: `array` ```yaml # Examples: regexp_topics_include: - logs_.* - metrics_.* # --- regexp_topics_include: - "events_[0-9]+" ``` ### [](#session_timeout)`session_timeout` When you specify a `consumer_group`, `session_timeout` sets the maximum interval between heartbeats sent by a consumer group member to the broker. If a broker doesn’t receive a heartbeat from a group member before the timeout expires, it removes the member from the consumer group and initiates a rebalance. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1m` ### [](#start_offset)`start_offset` Specify the offset from which this input starts or restarts consuming messages. Restarts occur when the `OffsetOutOfRange` error is seen during a fetch. **Type**: `string` **Default**: `earliest` | Option | Summary | | --- | --- | | committed | Prevents consuming a partition in a group if the partition has no prior commits. Corresponds to Kafka’s auto.offset.reset=none option | | earliest | Start from the earliest offset. Corresponds to Kafka’s auto.offset.reset=earliest option. | | latest | Start from the latest offset. Corresponds to Kafka’s auto.offset.reset=latest option. | ### [](#timely_nacks_maximum_wait)`timely_nacks_maximum_wait` EXPERIMENTAL: Specify a maximum period of time in which each message can be consumed and awaiting either acknowledgement or rejection before rejection is instead forced. This can be useful for avoiding situations where certain downstream components can result in blocked confirmation of delivery that exceeds SLAs. Accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` ### [](#topic_lag_refresh_period)`topic_lag_refresh_period` The interval between refresh cycles. During each cycle, this input queries the Redpanda Connect server to calculate the topic lag minus the number of produced messages that remain to be read from each topic/partition pair by the specified consumer group. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `5s` ### [](#topics)`topics[]` A list of topics to consume from. Use commas to separate multiple topics in a single element. When a `consumer_group` is specified, partitions are automatically distributed across consumers of a topic. Otherwise, all partitions are consumed. Alternatively, you can specify explicit partitions to consume by using a colon after the topic name. For example, `foo:0` would consume the partition `0` of the topic foo. This syntax supports ranges. For example, `foo:0-10` would consume partitions `0` through to `10` inclusive. It is also possible to specify an explicit offset to consume from by adding another colon after the partition. For example, `foo:0:10` would consume the partition `0` of the topic `foo` starting from the offset `10`. If the offset is not present (or remains unspecified) then the field `start_offset` determines which offset to start from. **Type**: `array` ```yaml # Examples: topics: - foo - bar # --- topics: - things.* # --- topics: - "foo,bar" # --- topics: - "foo:0" - "bar:1" - "bar:3" # --- topics: - "foo:0,bar:1,bar:3" # --- topics: - "foo:0-5" ``` ### [](#transaction_isolation_level)`transaction_isolation_level` The isolation level for handling transactional messages. This setting determines how transactions are processed and affects data consistency guarantees. **Type**: `string` **Default**: `read_uncommitted` | Option | Summary | | --- | --- | | read_committed | If set, only committed transactional records are processed. | | read_uncommitted | If set, then uncommitted records are processed. | --- # Page 82: redpanda_migrator **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda_migrator.md --- # redpanda_migrator > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda_migrator page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/redpanda_migrator page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/redpanda_migrator.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/redpanda_migrator.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-10-02" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Unified Kafka consumer for migrating data between Kafka/Redpanda clusters. Use this input with the [`redpanda_migrator` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_migrator/) to safely transfer topic data, ACLs, schemas, and consumer group offsets between clusters. This component is designed for migration scenarios. ### Common ```yml inputs: label: "" redpanda_migrator: seed_brokers: [] # No default (required) topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) schema_registry: url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} auto_replay_nacks: true ``` ### Advanced ```yml inputs: label: "" redpanda_migrator: seed_brokers: [] # No default (required) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) rack_id: "" instance_id: "" rebalance_timeout: 45s session_timeout: 1m heartbeat_interval: 3s start_offset: earliest fetch_max_bytes: 50MiB fetch_max_wait: 5s fetch_min_bytes: 1B fetch_max_partition_bytes: 1MiB transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) commit_period: 5s partition_buffer_bytes: 1MB topic_lag_refresh_period: 5s max_yield_batch_bytes: 32KB schema_registry: url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} auto_replay_nacks: true ``` The `redpanda_migrator` input: - Reads a batch of messages from a broker. - Waits for the `redpanda_migrator` output to acknowledge the writes before updating the Kafka consumer group offset. - Provides the same delivery guarantees and ordering semantics as the [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/). Specify a consumer group to make this input consume one or more topics and automatically balance the topic partitions across any other connected clients with the same consumer group. Otherwise, topics are consumed in their entirety or with explicit partitions. This input requires a corresponding `redpanda_migrator` output in the same pipeline. Each pipeline must have both input and output components configured. For capabilities, guarantees, scheduling, and examples, see the output documentation. ## [](#requirements)Requirements - Must be paired with a `redpanda_migrator` output in the same pipeline. - Requires access to a source Kafka or Redpanda cluster. - Consumer group configuration is recommended for partition balancing. - When the source cluster enforces ACLs, a consumer ACL alone is not enough for the source principal: it needs at minimum topic `READ` **and** `DESCRIBE_CONFIGS`, plus consumer group and cluster permissions. A `READ` ACL grants `DESCRIBE` but not `DESCRIBE_CONFIGS`, so the migrator consumes messages but fails to create topics with `TOPIC_AUTHORIZATION_FAILED`. See [Required permissions](https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/redpanda_migrator/#required-permissions). ## [](#multiple-migrator-pairs)Multiple migrator pairs When using multiple migrator pairs in a single pipeline, coordination is based on the `label` field. The label of the input and output must match exactly for correct pairing. If labels do not match, migration fails for that pair. ## [](#performance-tuning-for-high-throughput)Performance tuning for high throughput For workloads with high message rates or large messages, adjust the following settings to optimize throughput: On this input component: - `partition_buffer_bytes`: Set to 2MB to increase per-partition buffer size - `max_yield_batch_bytes`: Set to 1MB to allow larger batches to be yielded On the paired `redpanda_migrator` output component: - `max_in_flight`: Set to the total number of partitions being copied in parallel (up to all partitions in the cluster) > 📝 **NOTE** > > Setting `max_yield_batch_bytes` over 1MB is counter-productive unless you change the broker settings to allow bigger messages or batches. The `partition_buffer_bytes` setting allows for partition readahead. ## [](#metrics)Metrics This input emits an `input_redpanda_migrator_lag` metric with `topic` and `partition` labels for each consumed topic. This metric records the number of produced messages that remain to be read from each topic/partition pair by the specified consumer group. Monitor this metric to track migration progress and detect bottlenecks. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - kafka\_key - kafka\_topic - kafka\_partition - kafka\_offset - kafka\_lag - kafka\_timestamp\_ms - kafka\_timestamp\_unix - All record headers ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#commit_period)`commit_period` The period of time between each commit of the current partition offsets. Offsets are always committed during shutdown. **Type**: `string` **Default**: `5s` ### [](#conn_idle_timeout)`conn_idle_timeout` The maximum duration that connections can remain idle before they are automatically closed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `20s` ### [](#consumer_group)`consumer_group` An optional consumer group. When specified, the partitions of specified topics are automatically distributed across consumers sharing a consumer group, and partition offsets are automatically committed and resumed under this name. Consumer groups are not supported when explicit partitions are specified to consume from in the `topics` field. **Type**: `string` ### [](#fetch_max_bytes)`fetch_max_bytes` The maximum number of bytes that a broker tries to send during a fetch. If individual records are larger than the `fetch_max_bytes` value, brokers still send them. **Type**: `string` **Default**: `50MiB` ### [](#fetch_max_partition_bytes)`fetch_max_partition_bytes` The maximum number of bytes that are consumed from a single partition in a fetch request. This field is equivalent to the Java setting `fetch.max.partition.bytes`. If a single batch is larger than the `fetch_max_partition_bytes` value, the batch is still sent so that the client can make progress. **Type**: `string` **Default**: `1MiB` ### [](#fetch_max_wait)`fetch_max_wait` The maximum period of time a broker can wait for a fetch response to reach the required minimum number of bytes (`fetch_min_bytes`). **Type**: `string` **Default**: `5s` ### [](#fetch_min_bytes)`fetch_min_bytes` The minimum number of bytes that a broker tries to send during a fetch. This field is equivalent to the Java setting `fetch.min.bytes`. **Type**: `string` **Default**: `1B` ### [](#heartbeat_interval)`heartbeat_interval` When you specify a `consumer_group`, `heartbeat_interval` sets how frequently a consumer group member should send heartbeats to Apache Kafka. Apache Kafka uses heartbeats to make sure that a group member’s session is active. You must set `heartbeat_interval` to less than one-third of `session_timeout`. This field is equivalent to the Java `heartbeat.interval.ms` setting and accepts Go duration format strings such as `10s` or `2m`. **Type**: `string` **Default**: `3s` ### [](#instance_id)`instance_id` When you specify a [`consumer_group`](#consumer_group), assign a unique value to `instance_id` to define the group’s static membership, which can prevent unnecessary rebalances during reconnections. When you assign an instance ID, the client does not automatically leave the consumer group when it disconnects. To remove the client, you must use an external admin command on behalf of the instance ID. **Type**: `string` **Default**: `""` ### [](#max_yield_batch_bytes)`max_yield_batch_bytes` The maximum size (in bytes) for each batch yielded by this input. This value must be less than or equal to the `partition_buffer_bytes`. If using Redpanda output, this value should not be greater than the `max_message_bytes` option value (1MB by default), and for high-throughput scenarios they should be equal. **Type**: `string` **Default**: `32KB` ### [](#metadata_max_age)`metadata_max_age` The maximum period of time after which metadata is refreshed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. Lower values provide more responsive topic and partition discovery but may increase broker load. Higher values reduce broker queries but can delay detection of topology changes. **Type**: `string` **Default**: `1m` ### [](#partition_buffer_bytes)`partition_buffer_bytes` A buffer size (in bytes) for each consumed partition, which allows the internal queuing of records before they are flushed. Increasing this value may improve throughput but results in higher memory utilization. Each buffer can grow slightly beyond this value. **Type**: `string` **Default**: `1MB` ### [](#rack_id)`rack_id` A rack identifier for this client. **Type**: `string` **Default**: `""` ### [](#rebalance_timeout)`rebalance_timeout` When you specify a [`consumer_group`](#consumer_group), `rebalance_timeout` sets a time limit for all consumer group members to complete their work and commit offsets after a rebalance has begun. The timeout excludes the time taken to detect a failed or late heartbeat, which indicates a rebalance is required. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `45s` ### [](#regexp_topics_exclude)`regexp_topics_exclude[]` A list of regular expression patterns for excluding topics when regex mode is enabled (using `regexp_topics_include` or the deprecated `regexp_topics` boolean). Topics matching any of these patterns will be excluded from consumption, even if they match include patterns. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so use `^` and `$` for exact matching. Exclude patterns are applied after include patterns, providing fine-grained control over topic selection. Example: `regexp_topics_exclude: ["^_", ".**-temp$", ".**-test.*"]` excludes topics starting with underscore, ending with `-temp`, or containing `-test`. **Type**: `array` ### [](#regexp_topics_include)`regexp_topics_include[]` A list of regular expression patterns for matching topics to consume from. When specified, the client will periodically refresh the list of matching topics based on the `metadata_max_age` interval. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so `logs_.` **matches `my-logs_events` and `logs_errors`. Use `^logs_.`**`$` to match only topics starting with `logs_`. This field enables regex mode (replacing the deprecated `regexp_topics` boolean) and cannot be used together with explicit `topics` lists. Use `regexp_topics_exclude` to filter out specific patterns from the matched topics. Example: `regexp_topics_include: ["events_.**", "logs_.**"]` consumes from all topics starting with `events_` or `logs_`. **Type**: `array` ```yaml # Examples: regexp_topics_include: - logs_.* - metrics_.* # --- regexp_topics_include: - "events_[0-9]+" ``` ### [](#request_timeout_overhead)`request_timeout_overhead` Grants an additional buffer or overhead to requests that have timeout fields defined. This field is based on the behavior of Apache Kafka’s `request.timeout.ms` parameter. **Type**: `string` **Default**: `10s` ### [](#sasl)`sasl[]` Specify one or more methods of SASL authentication, which are tried in order. If the broker supports the first mechanism, all connections use that mechanism. If the first mechanism fails, the client picks the first supported mechanism. Connections fail if the broker does not support any client mechanisms. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM based authentication as specified in RFC5802. | | none | Disable sasl authentication | ### [](#sasl-password)`sasl[].password` A password to provide for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` A username to provide for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#schema_registry)`schema_registry` Configuration for schema registry integration. Enables migration of schema subjects, versions, and compatibility settings between clusters. **Type**: `object` ### [](#schema_registry-basic_auth)`schema_registry.basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#schema_registry-basic_auth-enabled)`schema_registry.basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-basic_auth-password)`schema_registry.basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-basic_auth-username)`schema_registry.basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt)`schema_registry.jwt` (beta) Allows you to specify JWT authentication. **Type**: `object` ### [](#schema_registry-jwt-claims)`schema_registry.jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-enabled)`schema_registry.jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-jwt-headers)`schema_registry.jwt.headers` Add optional key/value headers to the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-private_key_file)`schema_registry.jwt.private_key_file` A file with the PEM encoded via PKCS1 or PKCS8 as private key. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt-signing_method)`schema_registry.jwt.signing_method` A method used to sign the token such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth)`schema_registry.oauth` Allows you to specify open authentication via OAuth version 1. **Type**: `object` ### [](#schema_registry-oauth-access_token)`schema_registry.oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-access_token_secret)`schema_registry.oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_key)`schema_registry.oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_secret)`schema_registry.oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-enabled)`schema_registry.oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-timeout)`schema_registry.timeout` HTTP client timeout for schema registry requests. **Type**: `string` **Default**: `5s` ### [](#schema_registry-tls)`schema_registry.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#schema_registry-tls-client_certs)`schema_registry.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#schema_registry-tls-client_certs-cert)`schema_registry.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-cert_file)`schema_registry.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key)`schema_registry.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key_file)`schema_registry.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-password)`schema_registry.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#schema_registry-tls-enable_renegotiation)`schema_registry.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-enabled)`schema_registry.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-root_cas)`schema_registry.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#schema_registry-tls-root_cas_file)`schema_registry.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#schema_registry-tls-skip_cert_verify)`schema_registry.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#schema_registry-url)`schema_registry.url` The base URL of the schema registry service. Required for schema migration functionality. **Type**: `string` ```yaml # Examples: url: http://localhost:8081 # --- url: https://schema-registry.example.com:8081 ``` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to in order. Use commas to separate multiple addresses in a single list item. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#session_timeout)`session_timeout` When you specify a `consumer_group`, `session_timeout` sets the maximum interval between heartbeats sent by a consumer group member to the broker. If a broker doesn’t receive a heartbeat from a group member before the timeout expires, it removes the member from the consumer group and initiates a rebalance. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1m` ### [](#start_offset)`start_offset` Specify the offset from which this input starts or restarts consuming messages. Restarts occur when the `OffsetOutOfRange` error is seen during a fetch. **Type**: `string` **Default**: `earliest` | Option | Summary | | --- | --- | | committed | Prevents consuming a partition in a group if the partition has no prior commits. Corresponds to Kafka’s auto.offset.reset=none option | | earliest | Start from the earliest offset. Corresponds to Kafka’s auto.offset.reset=earliest option. | | latest | Start from the latest offset. Corresponds to Kafka’s auto.offset.reset=latest option. | ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic_lag_refresh_period)`topic_lag_refresh_period` The interval between refresh cycles. During each cycle, this input queries the Redpanda Connect server to calculate the topic lag minus the number of produced messages that remain to be read from each topic/partition pair by the specified consumer group. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `5s` ### [](#topics)`topics[]` A list of topics to consume from. Use commas to separate multiple topics in a single element. When a `consumer_group` is specified, partitions are automatically distributed across consumers of a topic. Otherwise, all partitions are consumed. Alternatively, you can specify explicit partitions to consume by using a colon after the topic name. For example, `foo:0` would consume the partition `0` of the topic foo. This syntax supports ranges. For example, `foo:0-10` would consume partitions `0` through to `10` inclusive. It is also possible to specify an explicit offset to consume from by adding another colon after the partition. For example, `foo:0:10` would consume the partition `0` of the topic `foo` starting from the offset `10`. If the offset is not present (or remains unspecified) then the field `start_offset` determines which offset to start from. **Type**: `array` ```yaml # Examples: topics: - foo - bar # --- topics: - things.* # --- topics: - "foo,bar" # --- topics: - "foo:0" - "bar:1" - "bar:3" # --- topics: - "foo:0,bar:1,bar:3" # --- topics: - "foo:0-5" ``` ### [](#transaction_isolation_level)`transaction_isolation_level` The isolation level for handling transactional messages. This setting determines how transactions are processed and affects data consistency guarantees. **Type**: `string` **Default**: `read_uncommitted` | Option | Summary | | --- | --- | | read_committed | If set, only committed transactional records are processed. | | read_uncommitted | If set, then uncommitted records are processed. | ## [](#troubleshooting)Troubleshooting - Ensure the input and output `label` fields match exactly. - Both input and output must be present in the pipeline. - Verify consumer group configuration for partition balancing. - Monitor the lag metric for stalled migration. ## [](#suggested-reading)Suggested reading - [`redpanda_migrator` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_migrator/) - [Migrating from legacy components](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/migrate-unified-redpanda-migrator/) --- # Page 83: redpanda **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda.md --- # redpanda > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/redpanda page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/redpanda.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/redpanda.adoc page-git-created-date: "2024-11-19" page-git-modified-date: "2026-05-26" --- Consumes topic data from one or more Kafka brokers. #### Common ```yml inputs: label: "" redpanda: seed_brokers: [] # No default (optional) topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" redpanda: seed_brokers: [] # No default (optional) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s topics: [] # No default (optional) regexp_topics_include: [] # No default (optional) regexp_topics_exclude: [] # No default (optional) rack_id: "" instance_id: "" rebalance_timeout: 45s session_timeout: 1m heartbeat_interval: 3s start_offset: earliest fetch_max_bytes: 50MiB fetch_max_wait: 5s fetch_min_bytes: 1B fetch_max_partition_bytes: 1MiB transaction_isolation_level: read_uncommitted consumer_group: "" # No default (optional) commit_period: 5s partition_buffer_bytes: 1MB topic_lag_refresh_period: 5s max_yield_batch_bytes: 32KB unordered_processing: enabled: false checkpoint_limit: 1024 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) auto_replay_nacks: true timely_nacks_maximum_wait: "" # No default (optional) extract_tracing_map: "" # No default (optional) ``` ## [](#consumer-groups)Consumer groups When you specify a consumer group in your configuration, this input consumes one or more topics and automatically balances the topic partitions across any other connected clients with the same consumer group. Otherwise, topics are consumed in their entirety or with explicit partitions. ## [](#delivery-guarantees)Delivery guarantees If you choose to use consumer groups, the offsets of records received by Redpanda Connect are committed automatically. In the event of restarts, this input uses the committed offsets to resume data consumption where it left off. Redpanda Connect guarantees at-least-once delivery. Records are only confirmed as delivered when all downstream outputs that a record is routed to have also confirmed delivery. ## [](#ordering)Ordering To preserve the order of topic partitions: - Records consumed from each partition are processed and delivered in the order that they are received - Only one batch of records of a given partition is processed at a time This approach means that although records from different partitions may be processed in parallel, records from the same partition are processed in sequential order. ### [](#delivery-errors)Delivery errors The order in which records are delivered may be disrupted by delivery errors and any error-handling mechanisms that start up. Redpanda Connect leans towards at-least-once delivery unless instructed otherwise, and this includes reattempting delivery of data when the ordering of that data is no longer guaranteed. For example, a batch of records is sent to an output broker and only a subset of records are delivered. In this scenario, Redpanda Connect (by default) attempts to deliver the records that failed, even though these delivery failures may have been sent before records that were delivered successfully. #### [](#use-a-fallback-output)Use a fallback output To prevent delivery errors from disrupting the order of records, you must specify a [`fallback`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/) output in your pipeline configuration. When adding a `fallback` output, it is good practice to set the `auto_retry_nacks` field to `false`. This also improves the throughput of your pipeline. For example, the following configuration includes a `fallback` output. If Redpanda Connect fails to write delivery errors to the `foo` topic, it then attempts to write them into a dead letter queue topic (`foo_dlq`), which is retried indefinitely as a way to apply back pressure. ```yaml output: fallback: - redpanda_common: topic: foo - retry: output: redpanda_common: topic: foo_dlq ``` ## [](#batching)Batching Records are processed and delivered from each partition in the same batches as they are received from brokers. Batch sizes are dynamically sized in order to optimize throughput, but you can tune them further using the following configuration fields: - `fetch_max_partition_bytes` - `fetch_max_bytes` You can break batches down further using the [`split`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/split/) processor. ## [](#metrics)Metrics This input emits a `redpanda_lag` metric with `topic` and `partition` labels for each consumed topic. The metric records the number of produced messages that remain to be read from each topic/partition pair by the specified consumer group. ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `kafka_key` - `kafka_topic` - `kafka_partition` - `kafka_offset` - `kafka_lag` - `kafka_timestamp_ms` - `kafka_timestamp_unix` - `kafka_tombstone_message` - All record headers ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#commit_period)`commit_period` The period of time between each commit of the current partition offsets. Offsets are always committed during shutdown. **Type**: `string` **Default**: `5s` ### [](#conn_idle_timeout)`conn_idle_timeout` The maximum duration that connections can remain idle before they are automatically closed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `20s` ### [](#consumer_group)`consumer_group` An optional consumer group. When this value is specified: - The partitions of any topics, specified in the `topics` field, are automatically distributed across consumers sharing a consumer group - Partition offsets are automatically committed and resumed under this name Consumer groups are not supported when you specify explicit partitions to consume from in the `topics` field. **Type**: `string` ### [](#extract_tracing_map)`extract_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that attempts to extract an object containing tracing propagation information, which will then be used as the root tracing span for the message. The specification of the extracted fields must match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: extract_tracing_map: root = @ # --- extract_tracing_map: root = this.meta.span ``` ### [](#fetch_max_bytes)`fetch_max_bytes` The maximum number of bytes that a broker tries to send during a fetch. If individual records are larger than the `fetch_max_bytes` value, brokers will still send them. **Type**: `string` **Default**: `50MiB` ### [](#fetch_max_partition_bytes)`fetch_max_partition_bytes` The maximum number of bytes that are consumed from a single partition in a fetch request. This field is equivalent to the Java setting `fetch.max.partition.bytes`. If a single batch is larger than the `fetch_max_partition_bytes` value, the batch is still sent so that the client can make progress. **Type**: `string` **Default**: `1MiB` ### [](#fetch_max_wait)`fetch_max_wait` The maximum period of time a broker can wait for a fetch response to reach the required minimum number of bytes (`fetch_min_bytes`). **Type**: `string` **Default**: `5s` ### [](#fetch_min_bytes)`fetch_min_bytes` The minimum number of bytes that a broker tries to send during a fetch. This field is equivalent to the Java setting `fetch.min.bytes`. **Type**: `string` **Default**: `1B` ### [](#heartbeat_interval)`heartbeat_interval` When you specify a `consumer_group`, `heartbeat_interval` sets how frequently a consumer group member should send heartbeats to Apache Kafka. Apache Kafka uses heartbeats to make sure that a group member’s session is active. You must set `heartbeat_interval` to less than one-third of `session_timeout`. This field is equivalent to the Java `heartbeat.interval.ms` setting and accepts Go duration format strings such as `10s` or `2m`. **Type**: `string` **Default**: `3s` ### [](#instance_id)`instance_id` When you specify a [`consumer_group`](#consumer_group), assign a unique value to `instance_id` to define the group’s static membership, which can prevent unnecessary rebalances during reconnections. When you assign an instance ID, the client does not automatically leave the consumer group when it disconnects. To remove the client, you must use an external admin command on behalf of the instance ID. **Type**: `string` **Default**: `""` ### [](#max_yield_batch_bytes)`max_yield_batch_bytes` The maximum size (in bytes) for each batch yielded by this input. This value must be less than or equal to the `partition_buffer_bytes`. If using Redpanda output, this value should not be greater than the `max_message_bytes` option value (1MB by default), and for high-throughput scenarios they should be equal. **Type**: `string` **Default**: `32KB` ### [](#metadata_max_age)`metadata_max_age` The maximum period of time after which metadata is refreshed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. Lower values provide more responsive topic and partition discovery but may increase broker load. Higher values reduce broker queries but can delay detection of topology changes. **Type**: `string` **Default**: `1m` ### [](#partition_buffer_bytes)`partition_buffer_bytes` A buffer size (in bytes) for each consumed partition, which allows the internal queuing of records before they are flushed. Increasing this value may improve throughput but results in higher memory utilization. Each buffer can grow slightly beyond this value. **Type**: `string` **Default**: `1MB` ### [](#rack_id)`rack_id` A rack specifies where the client is physically located, and changes fetch requests to consume from the closest replica as opposed to the leader replica. **Type**: `string` **Default**: `""` ### [](#rebalance_timeout)`rebalance_timeout` When you specify a [`consumer_group`](#consumer_group), `rebalance_timeout` sets a time limit for all consumer group members to complete their work and commit offsets after a rebalance has begun. The timeout excludes the time taken to detect a failed or late heartbeat, which indicates a rebalance is required. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `45s` ### [](#regexp_topics_exclude)`regexp_topics_exclude[]` A list of regular expression patterns for excluding topics when regex mode is enabled (using `regexp_topics_include` or the deprecated `regexp_topics` boolean). Topics matching any of these patterns will be excluded from consumption, even if they match include patterns. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so use `^` and `$` for exact matching. Exclude patterns are applied after include patterns, providing fine-grained control over topic selection. Example: `regexp_topics_exclude: ["^_", ".**-temp$", ".**-test.*"]` excludes topics starting with underscore, ending with `-temp`, or containing `-test`. **Type**: `array` ### [](#regexp_topics_include)`regexp_topics_include[]` A list of regular expression patterns for matching topics to consume from. When specified, the client will periodically refresh the list of matching topics based on the `metadata_max_age` interval. Each pattern is a full regular expression evaluated against the complete topic name. Patterns are not anchored by default, so `logs_.` **matches `my-logs_events` and `logs_errors`. Use `^logs_.`**`$` to match only topics starting with `logs_`. This field enables regex mode (replacing the deprecated `regexp_topics` boolean) and cannot be used together with explicit `topics` lists. Use `regexp_topics_exclude` to filter out specific patterns from the matched topics. Example: `regexp_topics_include: ["events_.**", "logs_.**"]` consumes from all topics starting with `events_` or `logs_`. **Type**: `array` ```yaml # Examples: regexp_topics_include: - logs_.* - metrics_.* # --- regexp_topics_include: - "events_[0-9]+" ``` ### [](#request_timeout_overhead)`request_timeout_overhead` Grants an additional buffer or overhead to requests that have timeout fields defined. This field is based on the behavior of Apache Kafka’s `request.timeout.ms` parameter. **Type**: `string` **Default**: `10s` ### [](#sasl)`sasl[]` Specify one or more methods or mechanisms of SASL authentication. They are tried in order. If the broker supports the first SASL mechanism, all connections use it. If the first mechanism fails, the client picks the first supported mechanism. If the broker does not support any client mechanisms, all connections fail. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM based authentication as specified in RFC5802. | | none | Disable sasl authentication | ### [](#sasl-password)`sasl[].password` A password to provide for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` A username to provide for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to in order. Use commas to separate multiple addresses in a single list item. Optional when `seed_brokers` is configured in a top-level `redpanda` block. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#session_timeout)`session_timeout` When you specify a `consumer_group`, `session_timeout` sets the maximum interval between heartbeats sent by a consumer group member to the broker. If a broker doesn’t receive a heartbeat from a group member before the timeout expires, it removes the member from the consumer group and initiates a rebalance. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1m` ### [](#start_offset)`start_offset` Specify the offset from which this input starts or restarts consuming messages. Restarts occur when the `OffsetOutOfRange` error is seen during a fetch. **Type**: `string` **Default**: `earliest` | Option | Summary | | --- | --- | | committed | Prevents consuming a partition in a group if the partition has no prior commits. Corresponds to Kafka’s auto.offset.reset=none option | | earliest | Start from the earliest offset. Corresponds to Kafka’s auto.offset.reset=earliest option. | | latest | Start from the latest offset. Corresponds to Kafka’s auto.offset.reset=latest option. | ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timely_nacks_maximum_wait)`timely_nacks_maximum_wait` EXPERIMENTAL: Specify a maximum period of time in which each message can be consumed and awaiting either acknowledgement or rejection before rejection is instead forced. This can be useful for avoiding situations where certain downstream components can result in blocked confirmation of delivery that exceeds SLAs. Accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic_lag_refresh_period)`topic_lag_refresh_period` The interval between refresh cycles. During each cycle, this input queries the Redpanda Connect server to calculate the topic lag minus the number of produced messages that remain to be read from each topic/partition pair by the specified consumer group. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `5s` ### [](#topics)`topics[]` A list of topics to consume from. Use commas to separate multiple topics in a single element. When a `consumer_group` is specified, partitions are automatically distributed across consumers of a topic. Otherwise, all partitions are consumed. Alternatively, you can specify explicit partitions to consume by using a colon after the topic name. For example, `foo:0` would consume the partition `0` of the topic foo. This syntax supports ranges. For example, `foo:0-10` would consume partitions `0` through to `10` inclusive. It is also possible to specify an explicit offset to consume from by adding another colon after the partition. For example, `foo:0:10` would consume the partition `0` of the topic `foo` starting from the offset `10`. If the offset is not present (or remains unspecified) then the field `start_offset` determines which offset to start from. **Type**: `array` ```yaml # Examples: topics: - foo - bar # --- topics: - things.* # --- topics: - "foo,bar" # --- topics: - "foo:0" - "bar:1" - "bar:3" # --- topics: - "foo:0,bar:1,bar:3" # --- topics: - "foo:0-5" ``` ### [](#transaction_isolation_level)`transaction_isolation_level` The isolation level for handling transactional messages. This setting determines how transactions are processed and affects data consistency guarantees. **Type**: `string` **Default**: `read_uncommitted` | Option | Summary | | --- | --- | | read_committed | If set, only committed transactional records are processed. | | read_uncommitted | If set, then uncommitted records are processed. | ### [](#unordered_processing)`unordered_processing` Allows consumers to process messages of any given partition in parallel, which may result in unordered processing. This option enables asynchronous publishing at the output level. The maximum parallelization of each partition is determined by the `checkpoint_limit` field. **Type**: `object` ### [](#unordered_processing-batching)`unordered_processing.batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/) that applies to individual topic partitions in order to batch messages together before flushing them for processing. Batching can be beneficial for performance and useful for windowed processing, and doing so preserves the ordering of topic partitions. **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#unordered_processing-batching-byte_size)`unordered_processing.batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#unordered_processing-batching-check)`unordered_processing.batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#unordered_processing-batching-count)`unordered_processing.batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#unordered_processing-batching-period)`unordered_processing.batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#unordered_processing-batching-processors)`unordered_processing.batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#unordered_processing-checkpoint_limit)`unordered_processing.checkpoint_limit` Determines how many messages of the same partition can be processed in parallel before applying back pressure. When a message of a given offset is delivered to the output the offset is only allowed to be committed when all messages of prior offsets have also been delivered, this ensures at-least-once delivery guarantees. However, this mechanism also increases the likelihood of duplicates in the event of crashes or server faults, reducing the checkpoint limit will mitigate this. **Type**: `int` **Default**: `1024` ### [](#unordered_processing-enabled)`unordered_processing.enabled` Whether to enable the unordered processing of messages from a given partition. **Type**: `bool` **Default**: `false` --- # Page 84: resource **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/resource.md --- # resource > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: resource latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/resource page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/resource.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/resource.adoc description: Resource is an input type that channels messages from a resource input, identified by its name. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Resource is an input type that channels messages from a resource input, identified by its name. ```yml inputs: label: "" resource: "" ``` Resources allow you to tidy up deeply nested configs. For example, the config: ```yaml input: broker: inputs: - kafka: addresses: [ TODO ] topics: [ foo ] consumer_group: foogroup - gcp_pubsub: project: bar subscription: baz ``` Could also be expressed as: ```yaml input: broker: inputs: - resource: foo - resource: bar input_resources: - label: foo kafka: addresses: [ TODO ] topics: [ foo ] consumer_group: foogroup - label: bar gcp_pubsub: project: bar subscription: baz ``` Resources also allow you to reference a single input in multiple places, such as multiple streams mode configs, or multiple entries in a broker input. However, when a resource is referenced more than once the messages it produces are distributed across those references, so each message will only be directed to a single reference, not all of them. --- # Page 85: salesforce_cdc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_cdc.md --- # salesforce_cdc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: salesforce_cdc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/salesforce_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/salesforce_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/salesforce_cdc.adoc description: Subscribes to one or more Salesforce Pub/Sub topics in parallel and emits a message per event. page-git-created-date: "2026-05-26" page-git-modified-date: "2026-08-11" --- Subscribes to one or more Salesforce Pub/Sub topics in parallel and emits a message per event. Topics may be: - `/data/ChangeEvent` — per-sObject CDC channel. - `/data/ChangeEvents` — CDC firehose (every CDC-enabled sObject). - `/event/__e` — custom Platform Event. - `/event/` — standard Platform Event (e.g. `LoginEventStream`). - A bare sObject name (e.g. `Account`) is shorthand for `/data/AccountChangeEvent`. Optionally runs a REST snapshot for the CDC sObjects before opening the streaming subscriptions, so the pipeline sees the current state plus continuous changes. Per-topic replay state persists in a cache resource so each subscription resumes across restarts independently. ## [](#when-to-use-this-input)When to use this input Use `salesforce_cdc` for: - Continuous ingestion with both historical state (snapshot) and live changes (CDC). - Real-time custom or standard Platform Events. - Mixed CDC + Platform Event pipelines under a single component. Use a different Salesforce input instead if: - You only need a one-off extract or periodic SOQL query — use [`salesforce`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce/). - You need a GraphQL query (cross-object in one request) — use [`salesforce_graphql`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_graphql/). ### Common ```yml inputs: label: "" salesforce_cdc: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 topics: [] # No default (required) stream_snapshot: true replay_preset: latest snapshot_max_batch_size: 2000 stream_batch_size: 100 max_parallel_snapshot_objects: 1 checkpoint_cache: "" # No default (required) checkpoint_cache_key: salesforce_cdc checkpoint_limit: 1024 auto_replay_nacks: true batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ### Advanced ```yml inputs: label: "" salesforce_cdc: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 topics: [] # No default (required) stream_snapshot: true replay_preset: latest snapshot_max_batch_size: 2000 stream_batch_size: 100 max_parallel_snapshot_objects: 1 checkpoint_cache: "" # No default (required) checkpoint_cache_key: salesforce_cdc checkpoint_limit: 1024 auto_replay_nacks: true grpc: reconnect_base_delay: 500ms reconnect_max_delay: 30s reconnect_max_attempts: 0 shutdown_timeout: 10s buffer_size: 1000 http: timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] proxy_url: "" disable_http2: false tps_limit: 0 tps_burst: 1 backoff: initial_interval: 1s max_interval: 30s max_retries: 3 tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s http: max_idle_conns: 100 max_idle_conns_per_host: 0 max_conns_per_host: 64 idle_conn_timeout: 1m30s tls_handshake_timeout: 10s expect_continue_timeout: 1s response_header_timeout: 0s disable_keep_alives: false disable_compression: false max_response_header_bytes: 1048576 max_response_body_bytes: 10485760 write_buffer_size: 4096 read_buffer_size: 4096 h2: strict_max_concurrent_requests: false max_decoder_header_table_size: 4096 max_encoder_header_table_size: 4096 max_read_frame_size: 16384 max_receive_buffer_per_connection: 1048576 max_receive_buffer_per_stream: 1048576 send_ping_timeout: 0s ping_timeout: 15s write_byte_timeout: 0s access_log_level: "" access_log_body_limit: 0 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#metadata)Metadata Every emitted message has: - `topic`: The full Pub/Sub topic path (e.g. "/event/Order\_\_e"). - `replay_id`: The Pub/Sub replay ID in hex (streaming events only). CDC events also carry: - `operation`: "read" for snapshot rows; "create", "update", "delete", or "undelete" for CDC events. - `sobject`: The sObject API name (e.g. "Account"). - `record_ids`: Comma-separated record IDs affected by the event (when present). Platform Events also carry: - `event_uuid`: The Salesforce `EventUuid` extracted from the payload (the canonical dedup key), when present. ## [](#fields)Fields ### [](#api_version)`api_version` Salesforce REST API version to target, prefixed with `v`. Affects endpoint paths (`/services/data/{api_version}/…​`) and available fields/objects. Must be supported by your org — check Setup → Company Information. Older versions may lack recent fields. **Type**: `string` **Default**: `v65.0` ```yaml # Examples: api_version: v65.0 # --- api_version: v62.0 ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#checkpoint_cache)`checkpoint_cache` Name of the cache resource used to persist snapshot cursor and per-topic replay IDs across restarts. The cache must be declared under the top-level `cache_resources` block. Choose a durable cache (Redis, Postgres, DynamoDB) for production; in-memory caches lose checkpoints on restart. **Type**: `string` ```yaml # Examples: checkpoint_cache: persistent_cache ``` ### [](#checkpoint_cache_key)`checkpoint_cache_key` Key inside the checkpoint cache where this input’s state is stored. Change when running multiple `salesforce_cdc` inputs against the same cache resource to avoid collisions. **Type**: `string` **Default**: `salesforce_cdc` ### [](#checkpoint_limit)`checkpoint_limit` Maximum number of unacknowledged batches in flight (per topic) before that topic pauses reading. Prevents unbounded memory growth when downstream components stall. Higher values increase throughput in steady state; lower values bound memory under backpressure. **Type**: `int` **Default**: `1024` ### [](#client_id)`client_id` Consumer Key of the Salesforce Connected App authorized for the OAuth Client Credentials flow. Create the Connected App under Setup → App Manager → New Connected App, enable OAuth settings, enable the Client Credentials Flow under `Flow Enablement`, then copy the Consumer Key from `Manage Consumer Details`. **Type**: `string` ### [](#client_secret)`client_secret` Consumer Secret of the Salesforce Connected App, paired with `client_id`. Sensitive — prefer environment variable interpolation (`${SALESFORCE_CLIENT_SECRET}`) over inlining. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#grpc)`grpc` gRPC transport tuning for the Pub/Sub API connection. **Type**: `object` ### [](#grpc-buffer_size)`grpc.buffer_size` Size of the internal gRPC event receive buffer. **Type**: `int` **Default**: `1000` ### [](#grpc-reconnect_base_delay)`grpc.reconnect_base_delay` Base delay for gRPC reconnection backoff. **Type**: `string` **Default**: `500ms` ### [](#grpc-reconnect_max_attempts)`grpc.reconnect_max_attempts` Maximum number of gRPC reconnection attempts. 0 means unlimited. This budget also governs transient schema-fetch failures: an event whose schema fetch keeps failing at the same replay position is retried (by reconnect-and-redeliver, or inline before the first event is delivered) this many times before the topic fails permanently. Deterministic schema failures - a schema that is missing, inaccessible, or fails to compile - and payloads that repeatedly fail to decode give up after a small fixed number of attempts regardless of this setting, since retrying cannot change the outcome. **Type**: `int` **Default**: `0` ### [](#grpc-reconnect_max_delay)`grpc.reconnect_max_delay` Maximum delay for gRPC reconnection backoff. **Type**: `string` **Default**: `30s` ### [](#grpc-shutdown_timeout)`grpc.shutdown_timeout` Timeout for graceful gRPC client shutdown. **Type**: `string` **Default**: `10s` ### [](#http)`http` HTTP client configuration for Salesforce REST calls (OAuth token endpoint and, where applicable, data queries). **Type**: `object` ### [](#http-access_log_body_limit)`http.access_log_body_limit` Maximum bytes of request/response body to include in logs. 0 to skip body logging. **Type**: `int` **Default**: `0` ### [](#http-access_log_level)`http.access_log_level` Log level for HTTP request/response logging. Empty disables logging. **Type**: `string` **Default**: `""` **Options**: `` `, `TRACE ``, `DEBUG`, `INFO`, `WARN`, `ERROR` ### [](#http-backoff)`http.backoff` Adaptive backoff configuration for 429 (Too Many Requests) responses. Always active. **Type**: `object` ### [](#http-backoff-initial_interval)`http.backoff.initial_interval` Initial interval between retries on 429 responses. **Type**: `string` **Default**: `1s` ### [](#http-backoff-max_interval)`http.backoff.max_interval` Maximum interval between retries on 429 responses. **Type**: `string` **Default**: `30s` ### [](#http-backoff-max_retries)`http.backoff.max_retries` Maximum number of retries on 429 responses. **Type**: `int` **Default**: `3` ### [](#http-disable_http2)`http.disable_http2` Disable HTTP/2 and force HTTP/1.1. **Type**: `bool` **Default**: `false` ### [](#http-http)`http.http` HTTP transport settings controlling connection pooling, timeouts, and HTTP/2. **Type**: `object` ### [](#http-http-disable_compression)`http.http.disable_compression` Disable automatic decompression of gzip responses. **Type**: `bool` **Default**: `false` ### [](#http-http-disable_keep_alives)`http.http.disable_keep_alives` Disable HTTP keep-alive connections; each request uses a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-expect_continue_timeout)`http.http.expect_continue_timeout` Maximum time to wait for a server’s 100-continue response before sending the body. 0 means the body is sent immediately. **Type**: `string` **Default**: `1s` ### [](#http-http-h2)`http.http.h2` HTTP/2-specific transport settings. Only applied when HTTP/2 is enabled. **Type**: `object` ### [](#http-http-h2-max_decoder_header_table_size)`http.http.h2.max_decoder_header_table_size` Upper limit in bytes for the HPACK header table used to decode headers from the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_encoder_header_table_size)`http.http.h2.max_encoder_header_table_size` Upper limit in bytes for the HPACK header table used to encode headers sent to the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_read_frame_size)`http.http.h2.max_read_frame_size` Largest HTTP/2 frame this endpoint will read. Valid range: 16 KiB to 16 MiB. **Type**: `int` **Default**: `16384` ### [](#http-http-h2-max_receive_buffer_per_connection)`http.http.h2.max_receive_buffer_per_connection` Maximum flow-control window size in bytes for data received on a connection. Must be at least 64 KiB and less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-max_receive_buffer_per_stream)`http.http.h2.max_receive_buffer_per_stream` Maximum flow-control window size in bytes for data received on a single stream. Must be less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-ping_timeout)`http.http.h2.ping_timeout` Timeout waiting for a PING response before closing the connection. **Type**: `string` **Default**: `15s` ### [](#http-http-h2-send_ping_timeout)`http.http.h2.send_ping_timeout` Idle timeout after which a PING frame is sent to verify connection health. 0 disables health checks. **Type**: `string` **Default**: `0s` ### [](#http-http-h2-strict_max_concurrent_requests)`http.http.h2.strict_max_concurrent_requests` When true, new requests block when a connection’s concurrency limit is reached instead of opening a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-h2-write_byte_timeout)`http.http.h2.write_byte_timeout` Timeout for writing data to a connection. The timer resets whenever bytes are written. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-idle_conn_timeout)`http.http.idle_conn_timeout` How long an idle connection remains in the pool before being closed. 0 disables the timeout. **Type**: `string` **Default**: `1m30s` ### [](#http-http-max_conns_per_host)`http.http.max_conns_per_host` Maximum total connections (active + idle) per host. 0 means unlimited. **Type**: `int` **Default**: `64` ### [](#http-http-max_idle_conns)`http.http.max_idle_conns` Maximum total number of idle (keep-alive) connections across all hosts. 0 means unlimited. **Type**: `int` **Default**: `100` ### [](#http-http-max_idle_conns_per_host)`http.http.max_idle_conns_per_host` Maximum idle connections to keep per host. 0 (the default) uses GOMAXPROCS+1. **Type**: `int` **Default**: `0` ### [](#http-http-max_response_body_bytes)`http.http.max_response_body_bytes` Maximum bytes of response body the client will read. The response body is wrapped with a limit reader; reads beyond this cap return EOF. 0 disables the limit. **Type**: `int` **Default**: `10485760` ### [](#http-http-max_response_header_bytes)`http.http.max_response_header_bytes` Maximum bytes of response headers to allow. **Type**: `int` **Default**: `1048576` ### [](#http-http-read_buffer_size)`http.http.read_buffer_size` Size in bytes of the per-connection read buffer. **Type**: `int` **Default**: `4096` ### [](#http-http-response_header_timeout)`http.http.response_header_timeout` Maximum time to wait for response headers after writing the full request. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-tls_handshake_timeout)`http.http.tls_handshake_timeout` Maximum time to wait for a TLS handshake to complete. 0 disables the timeout. **Type**: `string` **Default**: `10s` ### [](#http-http-write_buffer_size)`http.http.write_buffer_size` Size in bytes of the per-connection write buffer. **Type**: `int` **Default**: `4096` ### [](#http-proxy_url)`http.proxy_url` HTTP proxy URL. Empty string disables proxying. **Type**: `string` **Default**: `""` ### [](#http-tcp)`http.tcp` TCP socket configuration. **Type**: `object` ### [](#http-tcp-connect_timeout)`http.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-tcp-keep_alive)`http.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#http-tcp-keep_alive-count)`http.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#http-tcp-keep_alive-idle)`http.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#http-tcp-keep_alive-interval)`http.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#http-tcp-tcp_user_timeout)`http.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-timeout)`http.timeout` HTTP request timeout. **Type**: `string` **Default**: `5s` ### [](#http-tls)`http.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#http-tls-client_certs)`http.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#http-tls-client_certs-cert)`http.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-cert_file)`http.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key)`http.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key_file)`http.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-password)`http.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#http-tls-enable_renegotiation)`http.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#http-tls-enabled)`http.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#http-tls-root_cas)`http.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#http-tls-root_cas_file)`http.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#http-tls-skip_cert_verify)`http.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#http-tps_burst)`http.tps_burst` Maximum burst size for rate limiting. **Type**: `int` **Default**: `1` ### [](#http-tps_limit)`http.tps_limit` Rate limit in requests per second. 0 disables rate limiting. **Type**: `float` **Default**: `0` ### [](#max_parallel_snapshot_objects)`max_parallel_snapshot_objects` Number of sObjects snapshotted concurrently during the REST snapshot phase. Each in-flight snapshot consumes one HTTP connection and Salesforce API call quota. Default 1 serializes the work — raise when snapshotting many sObjects and your API limits permit. **Type**: `int` **Default**: `1` ### [](#org_url)`org_url` Salesforce instance base URL for your org, protocol included and no trailing slash. Used as the base for both the OAuth token endpoint and REST queries. Production orgs use `[https://{my-domain}.my.salesforce.com](https://{my-domain}.my.salesforce.com)`; sandboxes use `[https://{my-domain}.sandbox.my.salesforce.com](https://{my-domain}.sandbox.my.salesforce.com)`. Legacy instance URLs (`[https://na123.salesforce.com](https://na123.salesforce.com)`) still work but My Domain URLs are strongly recommended by Salesforce. **Type**: `string` ```yaml # Examples: org_url: https://acme.my.salesforce.com # --- org_url: https://acme--staging.sandbox.my.salesforce.com ``` ### [](#replay_preset)`replay_preset` Initial replay position used per topic only on first run (when no checkpoint exists in the cache); ignored once a topic’s replay ID has been written. - `latest`: Start from new events only; any changes between prior run and Connect are skipped. - `earliest`: Replay from the retention start (24h standard, 72h with enhanced retention). Use to recover missed events after outages. **Type**: `string` **Default**: `latest` **Options**: `latest`, `earliest` ### [](#snapshot_max_batch_size)`snapshot_max_batch_size` Page size for the REST snapshot query — records per `/query` response. Must be between 200 and 2000 per Salesforce REST API limits. Larger pages reduce HTTP round trips; smaller pages reduce peak memory per fetch. **Type**: `int` **Default**: `2000` ```yaml # Examples: snapshot_max_batch_size: 2000 # --- snapshot_max_batch_size: 500 ``` ### [](#stream_batch_size)`stream_batch_size` Number of events requested per gRPC `Fetch` call, per topic. Higher values improve throughput at the cost of peak batch memory; lower values give steadier latency under load. **Type**: `int` **Default**: `100` ```yaml # Examples: stream_batch_size: 100 # --- stream_batch_size: 500 ``` ### [](#stream_snapshot)`stream_snapshot` When true (default), paginate a full REST snapshot of every CDC sObject in `topics` before opening any streaming subscription. When false, skip the snapshot and start streaming immediately. Platform Event topics (`/event/…​`) are always skipped — they have no REST equivalent. **Type**: `bool` **Default**: `true` ### [](#topics)`topics[]` Pub/Sub topics to subscribe to. Each entry is one of: a bare sObject name (`Account` → `/data/AccountChangeEvent`), an explicit CDC channel (`/data/AccountChangeEvent`), the CDC firehose (`/data/ChangeEvents`), or a Platform Event topic (`/event/Order__e`, `/event/LoginEventStream`). Each topic gets its own gRPC subscription with an independent replay cursor. **Type**: `array` ```yaml # Examples: topics: - Account - Contact # --- topics: - /data/ChangeEvents # --- topics: - Account - /event/Order__e # --- topics: - Opportunity - MyCustom__c - /event/Sync_Requested__e ``` --- # Page 86: salesforce_graphql **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_graphql.md --- # salesforce_graphql > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: salesforce_graphql latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/salesforce_graphql page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/salesforce_graphql.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/salesforce_graphql.adoc description: Executes a GraphQL query against the Salesforce UIAPI and emits one message per record. page-git-created-date: "2026-05-26" page-git-modified-date: "2026-08-11" --- Executes a GraphQL query against the Salesforce UIAPI (`POST /services/data/{api_version}/graphql`), walks the response tree, and emits one message per record. When records are exhausted the input shuts down, letting the pipeline terminate gracefully (or the next input in a [sequence](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/) to take over). ## [](#when-to-use-this-input)When to use this input Use `salesforce_graphql` for: - Cross-object queries in a single request (parent + children + grandchildren). - Response shapes that already match your downstream schema. - Queries benefiting from GraphQL’s field-level selection and nesting. - Adopting Salesforce’s UIAPI / future-forward query surface. Use a different Salesforce input instead if: - You only need single-object SELECTs — use [`salesforce`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce/) (simpler, no GraphQL schema knowledge needed). - You need continuous change events — use [`salesforce_cdc`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_cdc/). When the query selects an `edges`/`pageInfo` connection, the input transparently paginates by injecting `after: ""` into the query string between requests. The query must include `pageInfo { hasNextPage endCursor }` for pagination to terminate cleanly. Responses without an `edges` array are emitted as a single message and the input completes. ### Common ```yml inputs: label: "" salesforce_graphql: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 query: "" # No default (required) variables: "" # No default (optional) auto_replay_nacks: true ``` ### Advanced ```yml inputs: label: "" salesforce_graphql: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 query: "" # No default (required) variables: "" # No default (optional) auto_replay_nacks: true http: timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] proxy_url: "" disable_http2: false tps_limit: 0 tps_burst: 1 backoff: initial_interval: 1s max_interval: 30s max_retries: 3 tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s http: max_idle_conns: 100 max_idle_conns_per_host: 0 max_conns_per_host: 64 idle_conn_timeout: 1m30s tls_handshake_timeout: 10s expect_continue_timeout: 1s response_header_timeout: 0s disable_keep_alives: false disable_compression: false max_response_header_bytes: 1048576 max_response_body_bytes: 10485760 write_buffer_size: 4096 read_buffer_size: 4096 h2: strict_max_concurrent_requests: false max_decoder_header_table_size: 4096 max_encoder_header_table_size: 4096 max_read_frame_size: 16384 max_receive_buffer_per_connection: 1048576 max_receive_buffer_per_stream: 1048576 send_ping_timeout: 0s ping_timeout: 15s write_byte_timeout: 0s access_log_level: "" access_log_body_limit: 0 ``` ## [](#metadata)Metadata This input adds no Salesforce-specific metadata. GraphQL response shapes vary by query, so record identity and context travel in the message body. ## [](#fields)Fields ### [](#api_version)`api_version` Salesforce REST API version to target, prefixed with `v`. Affects endpoint paths (`/services/data/{api_version}/…​`) and available fields/objects. Must be supported by your org — check Setup → Company Information. Older versions may lack recent fields. **Type**: `string` **Default**: `v65.0` ```yaml # Examples: api_version: v65.0 # --- api_version: v62.0 ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#client_id)`client_id` Consumer Key of the Salesforce Connected App authorized for the OAuth Client Credentials flow. Create the Connected App under Setup → App Manager → New Connected App, enable OAuth settings, enable the Client Credentials Flow under `Flow Enablement`, then copy the Consumer Key from `Manage Consumer Details`. **Type**: `string` ### [](#client_secret)`client_secret` Consumer Secret of the Salesforce Connected App, paired with `client_id`. Sensitive — prefer environment variable interpolation (`${SALESFORCE_CLIENT_SECRET}`) over inlining. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#http)`http` HTTP client configuration for Salesforce REST calls (OAuth token endpoint and, where applicable, data queries). **Type**: `object` ### [](#http-access_log_body_limit)`http.access_log_body_limit` Maximum bytes of request/response body to include in logs. 0 to skip body logging. **Type**: `int` **Default**: `0` ### [](#http-access_log_level)`http.access_log_level` Log level for HTTP request/response logging. Empty disables logging. **Type**: `string` **Default**: `""` **Options**: `` `, `TRACE ``, `DEBUG`, `INFO`, `WARN`, `ERROR` ### [](#http-backoff)`http.backoff` Adaptive backoff configuration for 429 (Too Many Requests) responses. Always active. **Type**: `object` ### [](#http-backoff-initial_interval)`http.backoff.initial_interval` Initial interval between retries on 429 responses. **Type**: `string` **Default**: `1s` ### [](#http-backoff-max_interval)`http.backoff.max_interval` Maximum interval between retries on 429 responses. **Type**: `string` **Default**: `30s` ### [](#http-backoff-max_retries)`http.backoff.max_retries` Maximum number of retries on 429 responses. **Type**: `int` **Default**: `3` ### [](#http-disable_http2)`http.disable_http2` Disable HTTP/2 and force HTTP/1.1. **Type**: `bool` **Default**: `false` ### [](#http-http)`http.http` HTTP transport settings controlling connection pooling, timeouts, and HTTP/2. **Type**: `object` ### [](#http-http-disable_compression)`http.http.disable_compression` Disable automatic decompression of gzip responses. **Type**: `bool` **Default**: `false` ### [](#http-http-disable_keep_alives)`http.http.disable_keep_alives` Disable HTTP keep-alive connections; each request uses a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-expect_continue_timeout)`http.http.expect_continue_timeout` Maximum time to wait for a server’s 100-continue response before sending the body. 0 means the body is sent immediately. **Type**: `string` **Default**: `1s` ### [](#http-http-h2)`http.http.h2` HTTP/2-specific transport settings. Only applied when HTTP/2 is enabled. **Type**: `object` ### [](#http-http-h2-max_decoder_header_table_size)`http.http.h2.max_decoder_header_table_size` Upper limit in bytes for the HPACK header table used to decode headers from the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_encoder_header_table_size)`http.http.h2.max_encoder_header_table_size` Upper limit in bytes for the HPACK header table used to encode headers sent to the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_read_frame_size)`http.http.h2.max_read_frame_size` Largest HTTP/2 frame this endpoint will read. Valid range: 16 KiB to 16 MiB. **Type**: `int` **Default**: `16384` ### [](#http-http-h2-max_receive_buffer_per_connection)`http.http.h2.max_receive_buffer_per_connection` Maximum flow-control window size in bytes for data received on a connection. Must be at least 64 KiB and less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-max_receive_buffer_per_stream)`http.http.h2.max_receive_buffer_per_stream` Maximum flow-control window size in bytes for data received on a single stream. Must be less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-ping_timeout)`http.http.h2.ping_timeout` Timeout waiting for a PING response before closing the connection. **Type**: `string` **Default**: `15s` ### [](#http-http-h2-send_ping_timeout)`http.http.h2.send_ping_timeout` Idle timeout after which a PING frame is sent to verify connection health. 0 disables health checks. **Type**: `string` **Default**: `0s` ### [](#http-http-h2-strict_max_concurrent_requests)`http.http.h2.strict_max_concurrent_requests` When true, new requests block when a connection’s concurrency limit is reached instead of opening a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-h2-write_byte_timeout)`http.http.h2.write_byte_timeout` Timeout for writing data to a connection. The timer resets whenever bytes are written. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-idle_conn_timeout)`http.http.idle_conn_timeout` How long an idle connection remains in the pool before being closed. 0 disables the timeout. **Type**: `string` **Default**: `1m30s` ### [](#http-http-max_conns_per_host)`http.http.max_conns_per_host` Maximum total connections (active + idle) per host. 0 means unlimited. **Type**: `int` **Default**: `64` ### [](#http-http-max_idle_conns)`http.http.max_idle_conns` Maximum total number of idle (keep-alive) connections across all hosts. 0 means unlimited. **Type**: `int` **Default**: `100` ### [](#http-http-max_idle_conns_per_host)`http.http.max_idle_conns_per_host` Maximum idle connections to keep per host. 0 (the default) uses GOMAXPROCS+1. **Type**: `int` **Default**: `0` ### [](#http-http-max_response_body_bytes)`http.http.max_response_body_bytes` Maximum bytes of response body the client will read. The response body is wrapped with a limit reader; reads beyond this cap return EOF. 0 disables the limit. **Type**: `int` **Default**: `10485760` ### [](#http-http-max_response_header_bytes)`http.http.max_response_header_bytes` Maximum bytes of response headers to allow. **Type**: `int` **Default**: `1048576` ### [](#http-http-read_buffer_size)`http.http.read_buffer_size` Size in bytes of the per-connection read buffer. **Type**: `int` **Default**: `4096` ### [](#http-http-response_header_timeout)`http.http.response_header_timeout` Maximum time to wait for response headers after writing the full request. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-tls_handshake_timeout)`http.http.tls_handshake_timeout` Maximum time to wait for a TLS handshake to complete. 0 disables the timeout. **Type**: `string` **Default**: `10s` ### [](#http-http-write_buffer_size)`http.http.write_buffer_size` Size in bytes of the per-connection write buffer. **Type**: `int` **Default**: `4096` ### [](#http-proxy_url)`http.proxy_url` HTTP proxy URL. Empty string disables proxying. **Type**: `string` **Default**: `""` ### [](#http-tcp)`http.tcp` TCP socket configuration. **Type**: `object` ### [](#http-tcp-connect_timeout)`http.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-tcp-keep_alive)`http.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#http-tcp-keep_alive-count)`http.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#http-tcp-keep_alive-idle)`http.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#http-tcp-keep_alive-interval)`http.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#http-tcp-tcp_user_timeout)`http.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-timeout)`http.timeout` HTTP request timeout. **Type**: `string` **Default**: `5s` ### [](#http-tls)`http.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#http-tls-client_certs)`http.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#http-tls-client_certs-cert)`http.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-cert_file)`http.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key)`http.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key_file)`http.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-password)`http.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#http-tls-enable_renegotiation)`http.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#http-tls-enabled)`http.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#http-tls-root_cas)`http.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#http-tls-root_cas_file)`http.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#http-tls-skip_cert_verify)`http.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#http-tps_burst)`http.tps_burst` Maximum burst size for rate limiting. **Type**: `int` **Default**: `1` ### [](#http-tps_limit)`http.tps_limit` Rate limit in requests per second. 0 disables rate limiting. **Type**: `float` **Default**: `0` ### [](#org_url)`org_url` Salesforce instance base URL for your org, protocol included and no trailing slash. Used as the base for both the OAuth token endpoint and REST queries. Production orgs use `[https://{my-domain}.my.salesforce.com](https://{my-domain}.my.salesforce.com)`; sandboxes use `[https://{my-domain}.sandbox.my.salesforce.com](https://{my-domain}.sandbox.my.salesforce.com)`. Legacy instance URLs (`[https://na123.salesforce.com](https://na123.salesforce.com)`) still work but My Domain URLs are strongly recommended by Salesforce. **Type**: `string` ```yaml # Examples: org_url: https://acme.my.salesforce.com # --- org_url: https://acme--staging.sandbox.my.salesforce.com ``` ### [](#query)`query` The GraphQL query document as a single string. Must target the Salesforce UIAPI schema (`uiapi.query.` **or `uiapi.mutation.`**). Typically follows the UIAPI convention of nested `edges { node { field { value } } }`. Variables are referenced with `$name` and supplied via the `variables` field. For automatic pagination, include `pageInfo { hasNextPage endCursor }` in the relevant connection. **Type**: `string` ```yaml # Examples: query: query Accounts { uiapi { query { Account { edges { node { Id { value } Name { value } } } pageInfo { hasNextPage endCursor } } } } } # --- query: query Accounts($first: Int) { uiapi { query { Account(first: $first) { edges { node { Id { value } Name { value } } } pageInfo { hasNextPage endCursor } } } } } ``` ### [](#variables)`variables` Optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) whose result must be an object whose keys are GraphQL variable names referenced by `query`. Values pass through as JSON in the request body’s `variables` field. The mapping is evaluated once at startup with no message context — use `env()`, `now()`, `cache()`, or static literals. **Type**: `string` ```yaml # Examples: variables: root = {"first": 100} # --- variables: root = {"since": now().ts_format("2006-01-02T15:04:05Z"), "limit": 500} ``` --- # Page 87: salesforce **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce.md --- # salesforce > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: salesforce latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/salesforce page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/salesforce.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/salesforce.adoc description: Runs a SOQL query against the Salesforce REST API, paginates through all result pages, and emits one message per record. page-git-created-date: "2026-05-28" page-git-modified-date: "2026-08-11" --- Runs a SOQL query against the Salesforce REST API, paginates through all result pages, and emits one message per record. When results are exhausted the input shuts down, letting the pipeline terminate gracefully (or the next input in a [sequence](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/) to take over). ## [](#when-to-use-this-input)When to use this input Use `salesforce` for: - One-shot extracts (e.g. dump all Accounts into a warehouse). - Periodic full-table refreshes via a scheduled pipeline or [sequence](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/). - Backfills and ad-hoc queries. - Warming up a downstream pipeline before switching to CDC. Use a different Salesforce input instead if: - You need continuous change events — use [`salesforce_cdc`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_cdc/). - You need a GraphQL query (cross-object in one request) — use [`salesforce_graphql`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_graphql/). ### Common ```yml inputs: label: "" salesforce: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 object: "" # No default (required) columns: [] # No default (required) where: "" # No default (optional) args_mapping: "" # No default (optional) auto_replay_nacks: true ``` ### Advanced ```yml inputs: label: "" salesforce: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 object: "" # No default (required) columns: [] # No default (required) where: "" # No default (optional) args_mapping: "" # No default (optional) prefix: "" # No default (optional) suffix: "" # No default (optional) auto_replay_nacks: true http: timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] proxy_url: "" disable_http2: false tps_limit: 0 tps_burst: 1 backoff: initial_interval: 1s max_interval: 30s max_retries: 3 tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s http: max_idle_conns: 100 max_idle_conns_per_host: 0 max_conns_per_host: 64 idle_conn_timeout: 1m30s tls_handshake_timeout: 10s expect_continue_timeout: 1s response_header_timeout: 0s disable_keep_alives: false disable_compression: false max_response_header_bytes: 1048576 max_response_body_bytes: 10485760 write_buffer_size: 4096 read_buffer_size: 4096 h2: strict_max_concurrent_requests: false max_decoder_header_table_size: 4096 max_encoder_header_table_size: 4096 max_read_frame_size: 16384 max_receive_buffer_per_connection: 1048576 max_receive_buffer_per_stream: 1048576 send_ping_timeout: 0s ping_timeout: 15s write_byte_timeout: 0s access_log_level: "" access_log_body_limit: 0 ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `sobject`: The sObject API name (e.g. "Account"). ## [](#fields)Fields ### [](#api_version)`api_version` Salesforce REST API version to target, prefixed with `v`. Affects endpoint paths (`/services/data/{api_version}/…​`) and available fields/objects. Must be supported by your org — check Setup → Company Information. Older versions may lack recent fields. **Type**: `string` **Default**: `v65.0` ```yaml # Examples: api_version: v65.0 # --- api_version: v62.0 ``` ### [](#args_mapping)`args_mapping` Optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) whose result must be an array of values matching the count of `?` placeholders in `where`. Values are SOQL-escaped: strings become quoted literals, timestamps become ISO-8601, booleans and numbers pass through. The mapping is evaluated once at startup with no message context — use `now()`, `env()`, or `cache()`. **Type**: `string` ```yaml # Examples: args_mapping: root = [ (now() - "1h").ts_format("2006-01-02T15:04:05Z") ] # --- args_mapping: root = [ "Active", (now() - "24h").ts_format("2006-01-02T15:04:05Z") ] ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#client_id)`client_id` Consumer Key of the Salesforce Connected App authorized for the OAuth Client Credentials flow. Create the Connected App under Setup → App Manager → New Connected App, enable OAuth settings, enable the Client Credentials Flow under `Flow Enablement`, then copy the Consumer Key from `Manage Consumer Details`. **Type**: `string` ### [](#client_secret)`client_secret` Consumer Secret of the Salesforce Connected App, paired with `client_id`. Sensitive — prefer environment variable interpolation (`${SALESFORCE_CLIENT_SECRET}`) over inlining. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#columns)`columns[]` Ordered list of field API names to retrieve. SOQL does not accept `*` — every field must be listed explicitly. Standard fields use their documented names; custom fields end with `__c`. Relationship fields traverse parents via dot notation (`Account.Name`, `Owner.Manager.Email`) up to 5 levels deep. Requesting a non-existent or non-queryable field fails at Connect time with a SOQL compile error. **Type**: `array` ```yaml # Examples: columns: - Id - Name - LastModifiedDate # --- columns: - Id - Account.Name - Owner.Email # --- columns: - Id - MyCustom__c ``` ### [](#http)`http` HTTP client configuration for Salesforce REST calls (OAuth token endpoint and, where applicable, data queries). **Type**: `object` ### [](#http-access_log_body_limit)`http.access_log_body_limit` Maximum bytes of request/response body to include in logs. 0 to skip body logging. **Type**: `int` **Default**: `0` ### [](#http-access_log_level)`http.access_log_level` Log level for HTTP request/response logging. Empty disables logging. **Type**: `string` **Default**: `""` **Options**: `` `, `TRACE ``, `DEBUG`, `INFO`, `WARN`, `ERROR` ### [](#http-backoff)`http.backoff` Adaptive backoff configuration for 429 (Too Many Requests) responses. Always active. **Type**: `object` ### [](#http-backoff-initial_interval)`http.backoff.initial_interval` Initial interval between retries on 429 responses. **Type**: `string` **Default**: `1s` ### [](#http-backoff-max_interval)`http.backoff.max_interval` Maximum interval between retries on 429 responses. **Type**: `string` **Default**: `30s` ### [](#http-backoff-max_retries)`http.backoff.max_retries` Maximum number of retries on 429 responses. **Type**: `int` **Default**: `3` ### [](#http-disable_http2)`http.disable_http2` Disable HTTP/2 and force HTTP/1.1. **Type**: `bool` **Default**: `false` ### [](#http-http)`http.http` HTTP transport settings controlling connection pooling, timeouts, and HTTP/2. **Type**: `object` ### [](#http-http-disable_compression)`http.http.disable_compression` Disable automatic decompression of gzip responses. **Type**: `bool` **Default**: `false` ### [](#http-http-disable_keep_alives)`http.http.disable_keep_alives` Disable HTTP keep-alive connections; each request uses a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-expect_continue_timeout)`http.http.expect_continue_timeout` Maximum time to wait for a server’s 100-continue response before sending the body. 0 means the body is sent immediately. **Type**: `string` **Default**: `1s` ### [](#http-http-h2)`http.http.h2` HTTP/2-specific transport settings. Only applied when HTTP/2 is enabled. **Type**: `object` ### [](#http-http-h2-max_decoder_header_table_size)`http.http.h2.max_decoder_header_table_size` Upper limit in bytes for the HPACK header table used to decode headers from the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_encoder_header_table_size)`http.http.h2.max_encoder_header_table_size` Upper limit in bytes for the HPACK header table used to encode headers sent to the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_read_frame_size)`http.http.h2.max_read_frame_size` Largest HTTP/2 frame this endpoint will read. Valid range: 16 KiB to 16 MiB. **Type**: `int` **Default**: `16384` ### [](#http-http-h2-max_receive_buffer_per_connection)`http.http.h2.max_receive_buffer_per_connection` Maximum flow-control window size in bytes for data received on a connection. Must be at least 64 KiB and less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-max_receive_buffer_per_stream)`http.http.h2.max_receive_buffer_per_stream` Maximum flow-control window size in bytes for data received on a single stream. Must be less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-ping_timeout)`http.http.h2.ping_timeout` Timeout waiting for a PING response before closing the connection. **Type**: `string` **Default**: `15s` ### [](#http-http-h2-send_ping_timeout)`http.http.h2.send_ping_timeout` Idle timeout after which a PING frame is sent to verify connection health. 0 disables health checks. **Type**: `string` **Default**: `0s` ### [](#http-http-h2-strict_max_concurrent_requests)`http.http.h2.strict_max_concurrent_requests` When true, new requests block when a connection’s concurrency limit is reached instead of opening a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-h2-write_byte_timeout)`http.http.h2.write_byte_timeout` Timeout for writing data to a connection. The timer resets whenever bytes are written. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-idle_conn_timeout)`http.http.idle_conn_timeout` How long an idle connection remains in the pool before being closed. 0 disables the timeout. **Type**: `string` **Default**: `1m30s` ### [](#http-http-max_conns_per_host)`http.http.max_conns_per_host` Maximum total connections (active + idle) per host. 0 means unlimited. **Type**: `int` **Default**: `64` ### [](#http-http-max_idle_conns)`http.http.max_idle_conns` Maximum total number of idle (keep-alive) connections across all hosts. 0 means unlimited. **Type**: `int` **Default**: `100` ### [](#http-http-max_idle_conns_per_host)`http.http.max_idle_conns_per_host` Maximum idle connections to keep per host. 0 (the default) uses GOMAXPROCS+1. **Type**: `int` **Default**: `0` ### [](#http-http-max_response_body_bytes)`http.http.max_response_body_bytes` Maximum bytes of response body the client will read. The response body is wrapped with a limit reader; reads beyond this cap return EOF. 0 disables the limit. **Type**: `int` **Default**: `10485760` ### [](#http-http-max_response_header_bytes)`http.http.max_response_header_bytes` Maximum bytes of response headers to allow. **Type**: `int` **Default**: `1048576` ### [](#http-http-read_buffer_size)`http.http.read_buffer_size` Size in bytes of the per-connection read buffer. **Type**: `int` **Default**: `4096` ### [](#http-http-response_header_timeout)`http.http.response_header_timeout` Maximum time to wait for response headers after writing the full request. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-tls_handshake_timeout)`http.http.tls_handshake_timeout` Maximum time to wait for a TLS handshake to complete. 0 disables the timeout. **Type**: `string` **Default**: `10s` ### [](#http-http-write_buffer_size)`http.http.write_buffer_size` Size in bytes of the per-connection write buffer. **Type**: `int` **Default**: `4096` ### [](#http-proxy_url)`http.proxy_url` HTTP proxy URL. Empty string disables proxying. **Type**: `string` **Default**: `""` ### [](#http-tcp)`http.tcp` TCP socket configuration. **Type**: `object` ### [](#http-tcp-connect_timeout)`http.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-tcp-keep_alive)`http.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#http-tcp-keep_alive-count)`http.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#http-tcp-keep_alive-idle)`http.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#http-tcp-keep_alive-interval)`http.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#http-tcp-tcp_user_timeout)`http.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-timeout)`http.timeout` HTTP request timeout. **Type**: `string` **Default**: `5s` ### [](#http-tls)`http.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#http-tls-client_certs)`http.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#http-tls-client_certs-cert)`http.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-cert_file)`http.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key)`http.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key_file)`http.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-password)`http.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#http-tls-enable_renegotiation)`http.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#http-tls-enabled)`http.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#http-tls-root_cas)`http.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#http-tls-root_cas_file)`http.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#http-tls-skip_cert_verify)`http.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#http-tps_burst)`http.tps_burst` Maximum burst size for rate limiting. **Type**: `int` **Default**: `1` ### [](#http-tps_limit)`http.tps_limit` Rate limit in requests per second. 0 disables rate limiting. **Type**: `float` **Default**: `0` ### [](#object)`object` The sObject API name to SELECT from. Case-sensitive; uses the API name, not the display label. Standard objects use the noun (`Account`, `Opportunity`); custom objects end with `_c_`_; Big Objects end with_ `b`; External Objects end with `__x`. Confirm the exact API name in Setup → Object Manager. **Type**: `string` ```yaml # Examples: object: Account # --- object: Contact # --- object: MyCustom__c ``` ### [](#org_url)`org_url` Salesforce instance base URL for your org, protocol included and no trailing slash. Used as the base for both the OAuth token endpoint and REST queries. Production orgs use `[https://{my-domain}.my.salesforce.com](https://{my-domain}.my.salesforce.com)`; sandboxes use `[https://{my-domain}.sandbox.my.salesforce.com](https://{my-domain}.sandbox.my.salesforce.com)`. Legacy instance URLs (`[https://na123.salesforce.com](https://na123.salesforce.com)`) still work but My Domain URLs are strongly recommended by Salesforce. **Type**: `string` ```yaml # Examples: org_url: https://acme.my.salesforce.com # --- org_url: https://acme--staging.sandbox.my.salesforce.com ``` ### [](#prefix)`prefix` Optional SOQL fragment inserted before the SELECT keyword. Rarely needed — provided for forward compatibility with future SOQL extensions or Bulk API framing. **Type**: `string` ### [](#suffix)`suffix` Optional SOQL fragment appended after the WHERE clause. Typical uses: `ORDER BY` for deterministic pagination, `LIMIT` to cap result size, `FOR REFERENCE` / `FOR VIEW` to mark records for Chatter tracking. **Type**: `string` ```yaml # Examples: suffix: ORDER BY LastModifiedDate DESC # --- suffix: ORDER BY Id LIMIT 1000 # --- suffix: ORDER BY CreatedDate DESC LIMIT 10000 ``` ### [](#where)`where` Optional SOQL WHERE body, without the `WHERE` keyword. `?` placeholders are substituted client-side from `args_mapping` with SOQL literal escaping (quoted strings, ISO-8601 datetimes). Supports the full WHERE grammar: `AND`/`OR`/`NOT`, `LIKE`, `IN`, date literals (`TODAY`, `LAST_N_DAYS:7`), subqueries. Date/datetime comparisons require ISO-8601 with explicit timezone. **Type**: `string` ```yaml # Examples: where: LastModifiedDate > ? # --- where: Status__c = ? AND CreatedDate > ? # --- where: OwnerId IN (?, ?) ``` --- # Page 88: schema_registry **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/schema_registry.md --- # schema_registry > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: schema_registry latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/schema_registry page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/schema_registry.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/schema_registry.adoc description: Reads schemas from SchemaRegistry. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Reads schemas from a schema registry. You can use this connector to extract and back up schemas during a data migration. This input uses the [Franz Kafka Schema Registry client](https://github.com/twmb/franz-go/tree/master/pkg/sr). #### Common ```yml inputs: label: "" schema_registry: url: "" # No default (required) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" schema_registry: url: "" # No default (required) include_deleted: false subject_filter: "" fetch_in_order: true tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] auto_replay_nacks: true oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `schema_registry_subject` - `schema_registry_subject_compatibility_level` - `schema_registry_version` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#example)Example This example reads all schemas from a schema registry that are associated with subjects matching the `^foo.*` filter, including deleted schemas. ```yaml input: schema_registry: url: http://localhost:8081 include_deleted: true subject_filter: ^foo.* ``` ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#basic_auth)`basic_auth` Configure basic authentication for requests from this component to your schema registry. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` The username of the account credentials to authenticate as. Used together with `password` for basic authentication. **Type**: `string` **Default**: `""` ### [](#fetch_in_order)`fetch_in_order` Indicate whether to fetch all schemas from the schema registry service and sort them by ID. Set this value to `true` if you use schemas that refer to other schemas (schema references). **Type**: `bool` **Default**: `true` ### [](#include_deleted)`include_deleted` Include deleted entities. **Type**: `bool` **Default**: `false` ### [](#jwt)`jwt` (beta) Configure JSON Web Token (JWT) authentication for secure data transmission from your schema registry to this component. This feature is in beta and may change in future releases. **Type**: `object` ### [](#jwt-claims)`jwt.claims` Values used to pass the identity of the authenticated entity to the service provider. In this case, between this component and the schema registry. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` The key/value pairs that identify the type of token and signing algorithm. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` A PEM-encoded file containing a private key that is formatted using either PKCS1 or PKCS8 standards. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` The method used to sign the token, such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#oauth)`oauth` Configure OAuth version 1.0 to give this component authorized access to your schema registry. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` The value this component can use to gain access to the data in the schema registry. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` The secret that establishes ownership of the `oauth.access_token` in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` The value used to identify this component or client to your schema registry. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` The secret that establishes ownership of the consumer key in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#subject_filter)`subject_filter` Include only subjects which match the regular expression filter, or leave this field value blank to select all subjects. **Type**: `string` **Default**: `""` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#url)`url` The base URL of the schema registry service. **Type**: `string` --- # Page 89: sequence **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence.md --- # sequence > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sequence latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/sequence page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/sequence.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/sequence.adoc description: Reads messages from a sequence of child inputs, starting with the first and once that input gracefully terminates starts consuming from the next, and so on. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Reads messages from a sequence of child inputs, starting with the first and once that input gracefully terminates starts consuming from the next, and so on. #### Common ```yml inputs: label: "" sequence: inputs: [] # No default (required) ``` #### Advanced ```yml inputs: label: "" sequence: sharded_join: type: none id_path: "" iterations: 1 merge_strategy: array inputs: [] # No default (required) ``` This input is useful for consuming from inputs that have an explicit end but must not be consumed in parallel. ## [](#examples)Examples ### [](#end-of-stream-message)End of Stream Message A common use case for sequence might be to generate a message at the end of our main input. With the following config once the records within `./dataset.csv` are exhausted our final payload `{"status":"finished"}` will be routed through the pipeline. ```yaml input: sequence: inputs: - file: paths: [ ./dataset.csv ] scanner: csv: {} - generate: count: 1 mapping: 'root = {"status":"finished"}' ``` ### [](#joining-data-simple)Joining Data (Simple) Redpanda Connect can be used to join unordered data from fragmented datasets in memory by specifying a common identifier field and a number of sharded iterations. For example, given two CSV files, the first called "main.csv", which contains rows of user data: ```csv uuid,name,age AAA,Melanie,34 BBB,Emma,28 CCC,Geri,45 ``` And the second called "hobbies.csv" that, for each user, contains zero or more rows of hobbies: ```csv uuid,hobby CCC,pokemon go AAA,rowing AAA,golf ``` We can parse and join this data into a single dataset: ```json {"uuid":"AAA","name":"Melanie","age":34,"hobbies":["rowing","golf"]} {"uuid":"BBB","name":"Emma","age":28} {"uuid":"CCC","name":"Geri","age":45,"hobbies":["pokemon go"]} ``` With the following config: ```yaml input: sequence: sharded_join: type: full-outer id_path: uuid merge_strategy: array inputs: - file: paths: - ./hobbies.csv - ./main.csv scanner: csv: {} ``` ### [](#joining-data-advanced)Joining Data (Advanced) In this example we are able to join unordered and fragmented data from a combination of CSV files and newline-delimited JSON documents by specifying multiple sequence inputs with their own processors for extracting the structured data. The first file "main.csv" contains straight forward CSV data: ```csv uuid,name,age AAA,Melanie,34 BBB,Emma,28 CCC,Geri,45 ``` And the second file called "hobbies.ndjson" contains JSON documents, one per line, that associate an identifier with an array of hobbies. However, these data objects are in a nested format: ```json {"document":{"uuid":"CCC","hobbies":[{"type":"pokemon go"}]}} {"document":{"uuid":"AAA","hobbies":[{"type":"rowing"},{"type":"golf"}]}} ``` And so we will want to map these into a flattened structure before the join, and then we will end up with a single dataset that looks like this: ```json {"uuid":"AAA","name":"Melanie","age":34,"hobbies":["rowing","golf"]} {"uuid":"BBB","name":"Emma","age":28} {"uuid":"CCC","name":"Geri","age":45,"hobbies":["pokemon go"]} ``` With the following config: ```yaml input: sequence: sharded_join: type: full-outer id_path: uuid iterations: 10 merge_strategy: array inputs: - file: paths: [ ./main.csv ] scanner: csv: {} - file: paths: [ ./hobbies.ndjson ] scanner: lines: {} processors: - mapping: | root.uuid = this.document.uuid root.hobbies = this.document.hobbies.map_each(this.type) ``` ## [](#fields)Fields ### [](#inputs)`inputs[]` An array of inputs to read from sequentially. **Type**: `array` ### [](#sharded_join)`sharded_join` EXPERIMENTAL: Provides a way to perform outer joins of arbitrarily structured and unordered data resulting from the input sequence, even when the overall size of the data surpasses the memory available on the machine. When configured the sequence of inputs will be consumed one or more times according to the number of iterations, and when more than one iteration is specified each iteration will process an entirely different set of messages by sharding them by the ID field. Increasing the number of iterations reduces the memory consumption at the cost of needing to fully parse the data each time. Each message must be structured (JSON or otherwise processed into a structured form) and the fields will be aggregated with those of other messages sharing the ID. At the end of each iteration the joined messages are flushed downstream before the next iteration begins, hence keeping memory usage limited. **Type**: `object` ### [](#sharded_join-id_path)`sharded_join.id_path` A [dot path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) that points to a common field within messages of each fragmented data set and can be used to join them. Messages that are not structured or are missing this field will be dropped. This field must be set in order to enable joins. **Type**: `string` **Default**: `""` ### [](#sharded_join-iterations)`sharded_join.iterations` The total number of iterations (shards), increasing this number will increase the overall time taken to process the data, but reduces the memory used in the process. The real memory usage required is significantly higher than the real size of the data and therefore the number of iterations should be at least an order of magnitude higher than the available memory divided by the overall size of the dataset. **Type**: `int` **Default**: `1` ### [](#sharded_join-merge_strategy)`sharded_join.merge_strategy` The chosen strategy to use when a data join would otherwise result in a collision of field values. The strategy `array` means non-array colliding values are placed into an array and colliding arrays are merged. The strategy `replace` replaces old values with new values. The strategy `keep` keeps the old value. **Type**: `string` **Default**: `array` **Options**: `array`, `replace`, `keep` ### [](#sharded_join-type)`sharded_join.type` The type of join to perform. A `full-outer` ensures that all identifiers seen in any of the input sequences are sent, and is performed by consuming all input sequences before flushing the joined results. An `outer` join consumes all input sequences but only writes data joined from the last input in the sequence, similar to a left or right outer join. With an `outer` join if an identifier appears multiple times within the final sequence input it will be flushed each time it appears. `full-outter` and `outter` have been deprecated in favour of `full-outer` and `outer`. **Type**: `string` **Default**: `none` **Options**: `none`, `full-outer`, `outer`, `full-outter`, `outter` --- # Page 90: sftp **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sftp.md --- # sftp > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sftp latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/sftp page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/sftp.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/sftp.adoc description: Consumes files from an SFTP server. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consumes files from an SFTP server. #### Common ```yml inputs: label: "" sftp: address: "" # No default (required) credentials: username: "" password: "" host_public_key_file: "" # No default (optional) host_public_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key: "" # No default (optional) private_key_pass: "" paths: [] # No default (required) auto_replay_nacks: true scanner: to_the_end: {} watcher: enabled: false minimum_age: 1s poll_interval: 1s cache: "" ``` #### Advanced ```yml inputs: label: "" sftp: address: "" # No default (required) connection_timeout: 30s credentials: username: "" password: "" host_public_key_file: "" # No default (optional) host_public_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key: "" # No default (optional) private_key_pass: "" max_sftp_sessions: 10 paths: [] # No default (required) auto_replay_nacks: true scanner: to_the_end: {} delete_on_finish: false watcher: enabled: false minimum_age: 1s poll_interval: 1s cache: "" ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `sftp_path` - `sftp_mod_time` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#address)`address` The address (hostname or IP address) of the SFTP server to connect to. **Type**: `string` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#connection_timeout)`connection_timeout` The connection timeout to use when connecting to the target server. **Type**: `string` **Default**: `30s` ### [](#credentials)`credentials` The credentials required to log in to the SFTP server. This can include a username and password, or a private key for secure access. **Type**: `object` ### [](#credentials-host_public_key)`credentials.host_public_key` The raw contents of the SFTP server’s public key, used for host key verification. **Type**: `string` ### [](#credentials-host_public_key_file)`credentials.host_public_key_file` The path to the SFTP server’s public key file, used for host key verification. **Type**: `string` ### [](#credentials-password)`credentials.password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#credentials-private_key)`credentials.private_key` The private key used to authenticate with the SFTP server. This field provides an alternative to the [`private_key_file`](#credentials-private_key_file). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-private_key_file)`credentials.private_key_file` The path to a private key file used to authenticate with the SFTP server. You can also provide a private key using the [`private_key`](#credentials-private_key) field. **Type**: `string` ### [](#credentials-private_key_pass)`credentials.private_key_pass` A passphrase for the private key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#credentials-username)`credentials.username` The username required to authenticate with the SFTP server. **Type**: `string` **Default**: `""` ### [](#delete_on_finish)`delete_on_finish` Whether to delete files from the server once they are processed. **Type**: `bool` **Default**: `false` ### [](#max_sftp_sessions)`max_sftp_sessions` The maximum number of SFTP sessions. **Type**: `int` **Default**: `10` ### [](#paths)`paths[]` A list of paths to consume sequentially. Glob patterns are supported. **Type**: `array` ### [](#scanner)`scanner` The [scanner](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/about/) by which the stream of bytes consumed will be broken out into individual messages. Scanners are useful for processing large sources of data without holding the entirety of it within memory. For example, the `csv` scanner allows you to process individual CSV rows without loading the entire CSV file in memory at once. **Type**: `scanner` **Default**: ```yaml to_the_end: {} ``` ### [](#watcher)`watcher` An experimental mode whereby the input will periodically scan the target paths for new files and consume them, when all files are consumed the input will continue polling for new files. **Type**: `object` ### [](#watcher-cache)`watcher.cache` A [cache resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) for storing the paths of files already consumed. **Type**: `string` **Default**: `""` ### [](#watcher-enabled)`watcher.enabled` Whether file watching is enabled. **Type**: `bool` **Default**: `false` ### [](#watcher-minimum_age)`watcher.minimum_age` The minimum period of time since a file was last updated before attempting to consume it. Increasing this period decreases the likelihood that a file will be consumed whilst it is still being written to. **Type**: `string` **Default**: `1s` ```yaml # Examples: minimum_age: 10s # --- minimum_age: 1m # --- minimum_age: 10m ``` ### [](#watcher-poll_interval)`watcher.poll_interval` The interval between each attempt to scan the target paths for new files. **Type**: `string` **Default**: `1s` ```yaml # Examples: poll_interval: 100ms # --- poll_interval: 1s ``` --- # Page 91: slack_users **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/slack_users.md --- # slack_users > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: slack_users latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/slack_users page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/slack_users.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/slack_users.adoc page-git-created-date: "2025-05-02" page-git-modified-date: "2026-05-26" --- Returns [the full profile](https://api.slack.com/methods/users.list#examples) of all users in your Slack organization using the API method [users.list](https://api.slack.com/methods/users.list). Optionally, you can filter the list of returned users by team ID. This input is useful when you need to: - Join user information to Slack posts. - Ingest user information into a data lakehouse to create joins with other fields. ```yml inputs: label: "" slack_users: bot_token: "" # No default (required) team_id: "" auto_replay_nacks: true ``` ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#bot_token)`bot_token` Your [Slack bot user’s OAuth token](https://api.slack.com/concepts/token-types), which must have the [`users.read` scope](https://api.slack.com/scopes/users:read) to access your Slack organization. **Type**: `string` ### [](#team_id)`team_id` The encoded ID of a Slack team by which to filter the list of returned users, which you can get from the [`team.info` Slack API method](https://api.slack.com/methods/team.info). If `team_id` is left empty, users from all teams within the organization are returned. **Type**: `string` **Default**: `""` --- # Page 92: slack **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/slack.md --- # slack > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: slack latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/slack page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/slack.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/slack.adoc page-git-created-date: "2025-05-02" page-git-modified-date: "2026-05-26" --- Connects to Slack using [Socket Mode](https://api.slack.com/apis/socket-mode), and can receive events, interactions (automated and user-initiated), and slash commands. This input is useful for: - Building bots that can query or write data. - Sending events to data warehouses. You could also try pairing this input with Redpanda Connect’s AI processors, which use the prefixes `cohere`, `openai`, and `ollama`. ```yml inputs: label: "" slack: app_token: "" # No default (required) bot_token: "" # No default (required) auto_replay_nacks: true ``` See also: [Examples](#examples) ## [](#metadata)Metadata Each message emitted from this input has an `@type` metadata flag to indicate the event type, either `"events_api"`, `"interactions"`, or `"slash_commands"`. ## [](#fields)Fields ### [](#app_token)`app_token` The app-level token to use to authenticate and connect to Slack. **Type**: `string` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether to automatically replay messages that are rejected (nacked) at the output level. If the cause of rejections is persistent, leaving this option enabled can result in back pressure. Set `auto_replay_nacks` to `false` to delete rejected messages. Disabling auto replays can greatly improve memory efficiency of high throughput streams, as the original shape of the data is discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#bot_token)`bot_token` Your Slack bot user’s OAuth token, which must have the [`connections.write` scope](https://api.slack.com/scopes/connections:write) to access your Slack app’s [Socket Mode WebSocket URL](https://api.slack.com/methods/apps.connections.open). **Type**: `string` ## [](#examples)Examples ### [](#echo-slackbot)Echo Slackbot A slackbot that echo messages from other users ```yaml input: slack: app_token: "${APP_TOKEN:xapp-demo}" bot_token: "${BOT_TOKEN:xoxb-demo}" pipeline: processors: - mutation: | # ignore hidden or non message events if this.event.type != "message" || (this.event.hidden | false) { root = deleted() } # Don't respond to our own messages if this.authorizations.any(auth -> auth.user_id == this.event.user) { root = deleted() } output: slack_post: bot_token: "${BOT_TOKEN:xoxb-demo}" channel_id: "${!this.event.channel}" thread_ts: "${!this.event.ts}" text: "ECHO: ${!this.event.text}" ``` --- # Page 93: spicedb_watch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/spicedb_watch.md --- # spicedb_watch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: spicedb_watch page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/spicedb_watch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/spicedb_watch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/spicedb_watch.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-11-19" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Consumes messages from the [Watch API](https://buf.build/authzed/api/docs/main:authzed.api.v1#authzed.api.v1.WatchService.Watch) of a [SpiceDB](https://authzed.com/docs/spicedb/getting-started/discovering-spicedb) instance. This input is useful if you have downstream applications that need to react to real-time changes in data managed by SpiceDB. #### Common ```yml inputs: label: "" spicedb_watch: endpoint: "" # No default (required) bearer_token: "" cache: "" # No default (required) ``` #### Advanced ```yml inputs: label: "" spicedb_watch: endpoint: "" # No default (required) bearer_token: "" max_receive_message_bytes: 4MB cache: "" # No default (required) cache_key: authzed.com/spicedb/watch/last_zed_token tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] ``` ## [](#authentication)Authentication For this input to authenticate with your SpiceDB instance, you must provide: - The [`endpoint`](#endpoint) of the SpiceDB instance - A [bearer token](#bearer_token) ## [](#configure-a-cache)Configure a cache You must use a cache resource to store the [ZedToken](https://authzed.com/docs/spicedb/concepts/consistency#zedtokens) (ID) of the latest message consumed and acknowledged by this input. Ideally, the cache should persist across restarts. This means that every time the input is initialized, it starts reading from the newest data updates. The following example uses a [`redis` cache](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/redis/). ```yml # Example input: label: "" spicedb_watch: endpoint: grpc.authzed.com:443 bearer_token: "" cache: "spicedb_cache" cache_resources: - label: "spicedb_cache" redis: url: redis://:6379 ``` To learn more about cache configuration, see the [Caches section](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/), which includes a range of cache components. ## [](#fields)Fields ### [](#bearer_token)`bearer_token` The SpiceDB bearer token to use to authenticate with your SpiceDB instance. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: bearer_token: t_your_token_here_1234567deadbeef ``` ### [](#cache)`cache` The [cache resource](#configure-a-cache) that you must configure to store the ZedToken (ID) of the last message processed. The ZedToken is stored in the cache within the `ACK` function of the message. This means that a ZedToken is only stored when a message is successfully routed through all processors and outputs in the data pipeline. **Type**: `string` ### [](#cache_key)`cache_key` The key identifier to use when storing the ZedToken (ID) of the last message received. **Type**: `string` **Default**: `authzed.com/spicedb/watch/last_zed_token` ### [](#endpoint)`endpoint` The endpoint of your SpiceDB instance. **Type**: `string` ```yaml # Examples: endpoint: grpc.authzed.com:443 ``` ### [](#max_receive_message_bytes)`max_receive_message_bytes` The maximum message size (in bytes) this input can receive. If a message exceeds this limit, an `rpc error` is written to the Redpanda Connect logs. **Type**: `string` **Default**: `4MB` ```yaml # Examples: max_receive_message_bytes: 100MB # --- max_receive_message_bytes: 50mib ``` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` --- # Page 94: splunk **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/splunk.md --- # splunk > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: splunk latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/splunk page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/splunk.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/splunk.adoc description: Consumes messages from Splunk. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consumes messages from Splunk. #### Common ```yml inputs: label: "" splunk: url: "" # No default (required) user: "" # No default (required) password: "" # No default (required) query: "" # No default (required) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" splunk: url: "" # No default (required) user: "" # No default (required) password: "" # No default (required) query: "" # No default (required) tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] auto_replay_nacks: true ``` ## [](#fields)Fields ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#password)`password` Splunk account password. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#query)`query` Splunk search query. **Type**: `string` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` Full HTTP Search API endpoint URL. **Type**: `string` ```yaml # Examples: url: https://foobar.splunkcloud.com/services/search/v2/jobs/export ``` ### [](#user)`user` Splunk account user. **Type**: `string` --- # Page 95: sql_raw **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sql_raw.md --- # sql_raw > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql_raw latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/sql_raw page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/sql_raw.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/sql_raw.adoc description: Executes a select query and creates a message for each row received. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a select query and creates a message for each row received. #### Common ```yml inputs: label: "" sql_raw: driver: "" # No default (required) dsn: "" # No default (required) query: "" # No default (required) args_mapping: "" # No default (optional) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" sql_raw: driver: "" # No default (required) dsn: "" # No default (required) query: "" # No default (required) args_mapping: "" # No default (optional) auto_replay_nacks: true init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) ``` When the rows from the query are exhausted, this input shuts down, allowing the pipeline to gracefully terminate or for the next input in a [sequence](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/) to execute. ## [](#examples)Examples ### [](#consumes-an-sql-table-using-a-query-as-an-input)Consumes an SQL table using a query as an input. Here we perform an aggregate over a list of names in a table that are less than 3600 seconds old. ```yaml input: sql_raw: driver: postgres dsn: postgres://foouser:foopass@localhost:5432/testdb?sslmode=disable query: "SELECT name, count(*) FROM person WHERE last_updated < $1 GROUP BY name;" args_mapping: | root = [ now().ts_unix() - 3600 ] ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that includes the same number of values in an array as the placeholder arguments in the [`query`](#query) field. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#query)`query` The query to execute. The style of placeholder to use depends on the driver, some drivers require question marks (`?`) whereas others expect incrementing dollar signs (`$1`, `$2`, and so on) or colons (`:1`, `:2` and so on). The style to use is outlined in this table: | Driver | Placeholder Style | | --- | --- | | clickhouse | Dollar sign ($) | | gocosmos | Colon (:) | | mysql | Question mark (?) | | mssql | Question mark (?) | | oracle | Colon (:) | | postgres | Dollar sign ($) | | snowflake | Question mark (?) | | spanner | Question mark (?) | | sqlite | Question mark (?) | | trino | Question mark (?) | **Type**: `string` ```yaml # Examples: query: SELECT * FROM footable WHERE user_id = $1; ``` --- # Page 96: sql_select **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sql_select.md --- # sql_select > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql_select latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/sql_select page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/sql_select.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/sql_select.adoc description: Executes a select query and creates a message for each row received. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a select query and creates a message for each row received. #### Common ```yml inputs: label: "" sql_select: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) where: "" # No default (optional) args_mapping: "" # No default (optional) auto_replay_nacks: true ``` #### Advanced ```yml inputs: label: "" sql_select: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) where: "" # No default (optional) args_mapping: "" # No default (optional) prefix: "" # No default (optional) suffix: "" # No default (optional) auto_replay_nacks: true init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) ``` Once the rows from the query are exhausted this input shuts down, allowing the pipeline to gracefully terminate (or the next input in a [sequence](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/sequence/) to execute). ## [](#examples)Examples ### [](#consume-a-table-postgresql)Consume a Table (PostgreSQL) Here we define a pipeline that will consume all rows from a table created within the last hour by comparing the unix timestamp stored in the row column "created\_at": ```yaml input: sql_select: driver: postgres dsn: postgres://foouser:foopass@localhost:5432/testdb?sslmode=disable table: footable columns: [ '*' ] where: created_at >= ? args_mapping: | root = [ now().ts_unix() - 3600 ] ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of placeholder arguments in the field `where`. **Type**: `string` ```yaml # Examples: args_mapping: root = [ "article", now().ts_format("2006-01-02") ] ``` ### [](#auto_replay_nacks)`auto_replay_nacks` Whether messages that are rejected (nacked) at the output level should be automatically replayed indefinitely, eventually resulting in back pressure if the cause of the rejections is persistent. If set to `false` these messages will instead be deleted. Disabling auto replays can greatly improve memory efficiency of high throughput streams as the original shape of the data can be discarded immediately upon consumption and mutation. **Type**: `bool` **Default**: `true` ### [](#columns)`columns[]` A list of columns to select. **Type**: `array` ```yaml # Examples: columns: - "*" # --- columns: - foo - bar - baz ``` ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#prefix)`prefix` An optional prefix to prepend to the select query (before SELECT). **Type**: `string` ### [](#suffix)`suffix` An optional suffix to append to the select query. **Type**: `string` ### [](#table)`table` The table to select from. **Type**: `string` ```yaml # Examples: table: foo ``` ### [](#where)`where` An optional where clause to add. Placeholder arguments are populated with the `args_mapping` field. Placeholders should always be question marks, and will automatically be converted to dollar syntax when the postgres or clickhouse drivers are used. **Type**: `string` ```yaml # Examples: where: type = ? and created_at > ? # --- where: user_id = ? ``` --- # Page 97: timeplus **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/timeplus.md --- # timeplus > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: timeplus latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/inputs/timeplus page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/inputs/timeplus.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/inputs/timeplus.adoc description: Executes a query on Timeplus Enterprise and creates a message from each row received. page-git-created-date: "2024-11-19" page-git-modified-date: "2026-05-26" --- Executes a streaming or table query on [Timeplus Enterprise (Cloud or Self-Hosted)](https://docs.timeplus.com/) or the `timeplusd` component, and creates a structured message for each table row received. If you execute a streaming query, this input runs until the query terminates. For table queries, it shuts down after all rows returned by the query are exhausted. ```yml inputs: label: "" timeplus: query: "" # No default (required) url: tcp://localhost:8463 workspace: "" # No default (optional) apikey: "" # No default (optional) username: "" # No default (optional) password: "" # No default (optional) ``` ## [](#examples)Examples ### [](#from-timeplus-enterprise-cloud-via-http)From Timeplus Enterprise Cloud via HTTP You will need to create API Key on Timeplus Enterprise Cloud Web console first and then set the `apikey` field. ```yaml input: timeplus: url: https://us-west-2.timeplus.cloud workspace: my_workspace_id query: select * from iot apikey: ``` ### [](#from-timeplus-enterprise-self-hosted-via-http)From Timeplus Enterprise (self-hosted) via HTTP For self-hosted Timeplus Enterprise, you will need to specify the username and password as well as the URL of the App server ```yaml input: timeplus: url: http://localhost:8000 workspace: my_workspace_id query: select * from iot username: username password: pw ``` ### [](#from-timeplus-enterprise-self-hosted-via-tcp)From Timeplus Enterprise (self-hosted) via TCP Make sure the the schema of url is tcp ```yaml input: timeplus: url: tcp://localhost:8463 query: select * from iot username: timeplus password: timeplus ``` ## [](#fields)Fields ### [](#apikey)`apikey` The API key for the Timeplus Enterprise REST API. You need to generate the key in the web console of Timeplus Enterprise (Cloud). This field is required if you are reading messages from Timeplus Enterprise (Cloud). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#password)`password` The password for the Timeplus application server. This field is required if you are reading messages from Timeplus Enterprise (Self-Hosted) or `timeplusd`. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#query)`query` The query to execute on Timeplus Enterprise (Cloud or Self-Hosted) or `timeplusd`. **Type**: `string` ```yaml # Examples: query: select * from iot # --- query: select count(*) from table(iot) ``` ### [](#url)`url` The URL of your Timeplus instance, which should always include the schema and host. **Type**: `string` **Default**: `tcp://localhost:8463` ### [](#username)`username` The username for the Timeplus application server. This field is required if you are reading messages from Timeplus Enterprise (Self-Hosted) or `timeplusd`. **Type**: `string` ### [](#workspace)`workspace` The ID of the workspace you want to read messages from. This field is required if you are connecting to Timeplus Enterprise (Cloud or Self-Hosted) using HTTP. **Type**: `string` --- # Page 98: Logger **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/logger/about.md --- # Logger > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Logger latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/logger/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/logger/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/logger/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Redpanda Connect logging prints to stdout (or stderr if your output is stdout) and is formatted as [logfmt](https://brandur.org/logfmt) by default. Use these configuration options to change both the logging formats as well as the destination of logs. #### Common ```yaml # Common config fields, showing default values logger: level: INFO format: logfmt add_timestamp: false static_fields: '@service': redpanda-connect ``` #### Advanced ```yaml # All config fields, showing default values logger: level: INFO format: logfmt add_timestamp: false level_name: level timestamp_name: time message_name: msg static_fields: '@service': redpanda-connect file: path: "" rotate: false rotate_max_age_days: 0 ``` ## [](#fields)Fields The schema of the `logger` section is as follows: ### [](#level)`level` Set the minimum severity level for emitting logs. **Type**: `string` **Default**: `"INFO"` Options: `OFF` , `FATAL` , `ERROR` , `WARN` , `INFO` , `DEBUG` , `TRACE` , `ALL` , `NONE` ### [](#format)`format` Set the format of emitted logs. **Type**: `string` **Default**: `"logfmt"` Options: `json` , `logfmt` ### [](#add_timestamp)`add_timestamp` Whether to include timestamps in logs. **Type**: `bool` **Default**: `false` ### [](#level_name)`level_name` The name of the level field added to logs when the `format` is `json`. **Type**: `string` **Default**: `"level"` ### [](#timestamp_name)`timestamp_name` The name of the timestamp field added to logs when `add_timestamp` is set to `true` and the `format` is `json`. **Type**: `string` **Default**: `"time"` ### [](#message_name)`message_name` The name of the message field added to logs when the `format` is `json`. **Type**: `string` **Default**: `"msg"` ### [](#static_fields)`static_fields` A map of key/value pairs to add to each structured log. **Type**: `object` **Default**: `{"@service":"redpanda-connect"}` ### [](#file)`file` Experimental: Specify fields for optionally writing logs to a file. **Type**: `object` ### [](#file-path)`file.path` The file path to write logs to, if the file does not exist it will be created. Leave this field empty or unset to disable file based logging. **Type**: `string` **Default**: `""` ### [](#file-rotate)`file.rotate` Whether to rotate log files automatically. **Type**: `bool` **Default**: `false` ### [](#file-rotate_max_age_days)`file.rotate_max_age_days` The maximum number of days to retain old log files based on the timestamp encoded in their filename, after which they are deleted. Setting to zero disables this mechanism. **Type**: `int` **Default**: `0` --- # Page 99: Metrics **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about.md --- # Metrics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Metrics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/metrics/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/metrics/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/metrics/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Redpanda Connect emits lots of metrics in order to expose how components configured within your pipeline are behaving. You can configure exactly where these metrics end up with the config field `metrics`, which describes a metrics format and destination. For example, if you wished to push them via the StatsD protocol you could use this configuration: ```yaml metrics: statsd: address: localhost:8125 flush_period: 100ms ``` Redpanda Connect automatically [exports detailed metrics](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/monitor-connect/) for each component of your data pipeline to a Prometheus endpoint. ## [](#timings)Timings It’s worth noting that timing metrics within Redpanda Connect are measured in nanoseconds and are therefore named with a `_ns` suffix. However, some exporters do not support this level of precision and are downgraded, or have the unit converted for convenience. In these cases the exporter documentation outlines the conversion and why it is made. ## [](#metric-names)Metric names Each major Redpanda Connect component type emits one or more metrics with the name prefixed by the type. These metrics are intended to provide an overview of behavior, performance and health. Some specific component implementations may provide their own unique metrics on top of these standardized ones, these extra metrics can be found listed on their respective documentation pages. ## [](#inputs)Inputs - `input_received`: A count of the number of messages received by the input. - `input_latency_ns`: Measures the roundtrip latency in nanoseconds from the point at which a message is read up to the moment the message has either been acknowledged by an output, has been stored within a buffer, or has been rejected (nacked). - `batch_created`: A count of each time an input-level batch has been created using a batching policy. Includes a label `mechanism` describing the particular mechanism that triggered it, one of; `count`, `size`, `period`, `check`. - `input_connection_up`: For continuous stream based inputs represents a count of the number of the times the input has successfully established a connection to the target source. For poll based inputs that do not retain an active connection this value will increment once. - `input_connection_failed`: For continuous stream based inputs represents a count of the number of times the input has failed to establish a connection to the target source. - `input_connection_lost`: For continuous stream based inputs represents a count of the number of times the input has lost a previously established connection to the target source. > ⚠️ **CAUTION** > > The behavior of connection metrics may differ based on input type due to certain libraries and protocols obfuscating the concept of a single connection. ### [](#buffers)Buffers - `buffer_received`: A count of the number of messages written to the buffer. - `buffer_batch_received`: A count of the number of message batches written to the buffer. - `buffer_sent`: A count of the number of messages read from the buffer. - `buffer_batch_sent`: A count of the number of message batches read from the buffer. - `buffer_latency_ns`: Measures the roundtrip latency in nanoseconds from the point at which a message is read from the buffer up to the moment it has been acknowledged by the output. - `batch_created`: A count of each time a buffer-level batch has been created using a batching policy. Includes a label `mechanism` describing the particular mechanism that triggered it, one of; `count`, `size`, `period`, `check`. ### [](#processors)Processors - `processor_received`: A count of the number of messages the processor has been executed upon. - `processor_batch_received`: A count of the number of message batches the processor has been executed upon. - `processor_sent`: A count of the number of messages the processor has returned. - `processor_batch_sent`: A count of the number of message batches the processor has returned. - `processor_error`: A count of the number of times the processor has errored. In cases where an error is batch-wide the count is incremented by one, and therefore would not match the number of messages. - `processor_latency_ns`: Latency of message processing in nanoseconds. When a processor acts upon a batch of messages this latency measures the time taken to process all messages of the batch. ### [](#outputs)Outputs - `output_sent`: A count of the number of messages sent by the output. - `output_batch_sent`: A count of the number of message batches sent by the output. - `output_error`: A count of the number of send attempts that have failed. On failed batched sends this count is incremented once only. - `output_latency_ns`: Latency of writes in nanoseconds. This metric may not be populated by outputs that are pull-based such as the `http_server`. - `batch_created`: A count of each time an output-level batch has been created using a batching policy. Includes a label `mechanism` describing the particular mechanism that triggered it, one of; `count`, `size`, `period`, `check`. - `output_connection_up`: For continuous stream based outputs represents a count of the number of the times the output has successfully established a connection to the target sink. For poll based outputs that do not retain an active connection this value will increment once. - `output_connection_failed`: For continuous stream based outputs represents a count of the number of times the output has failed to establish a connection to the target sink. - `output_connection_lost`: For continuous stream based outputs represents a count of the number of times the output has lost a previously established connection to the target sink. > ⚠️ **CAUTION** > > The behavior of connection metrics may differ based on output type due to certain libraries and protocols obfuscating the concept of a single connection. ### [](#caches)Caches All cache metrics have a label `operation` denoting the operation that triggered the metric series, one of; `add`, `get`, `set` or `delete`. - `cache_success`: A count of the number of successful cache operations. - `cache_error`: A count of the number of cache operations that resulted in an error. - `cache_latency_ns`: Latency of operations in nanoseconds. - `cache_not_found`: A count of the number of get operations that yielded no value due to the item not being found. This count is separate from `cache_error`. - `cache_duplicate`: A count of the number of add operations that were aborted due to the key already existing. This count is separate from `cache_error`. ### [](#rate-limits)Rate limits - `rate_limit_checked`: A count of the number of times the rate limit has been probed. - `rate_limit_triggered`: A count of the number of times the rate limit has been triggered by a probe. - `rate_limit_error`: A count of the number of times the rate limit has errored when probed. ## [](#metric-labels)Metric labels The standard metric names are unique to the component type, but a benthos config may consist of any number of component instantiations. In order to provide a metrics series that is unique for each instantiation Redpanda Connect adds labels (or tags) that uniquely identify the instantiation. These labels are as follows: ### [](#path)`path` The `path` label contains a string representation of the position of a component instantiation within a config in a format that would locate it within a Bloblang mapping, beginning at `root`. This path is a best attempt and may not exactly represent the source component position in all cases and is intended to be used for assisting observability only. This is the highest cardinality label since paths will change as configs are updated and expanded. It is therefore worth removing this label with a [mapping](#metric-mapping) in cases where you wish to restrict the number of unique metric series. ### [](#label)`label` The `label` label contains the unique label configured for a component emitting the metric series, or is empty for components that do not have a configured label. This is the most useful label for uniquely identifying a series for a component. ### [](#stream)`stream` The `stream` label is present in a metric series emitted from a stream config executed when Redpanda Connect is running in streams mode, and is populated with the stream name. ## [](#example)Example The following Redpanda Connect configuration: ```yaml input: label: foo http_server: {} pipeline: processors: - mapping: | root.message = this root.meta.link_count = this.links.length() root.user.age = this.user.age.number() output: label: bar stdout: {} metrics: prometheus: {} ``` Would produce the following metrics series: ```text input_latency_ns{label="foo",path="root.input"} input_received{endpoint="post",label="foo",path="root.input"} input_received{endpoint="websocket",label="foo",path="root.input"} processor_batch_received{label="",path="root.pipeline.processors.0"} processor_batch_sent{label="",path="root.pipeline.processors.0"} processor_error{label="",path="root.pipeline.processors.0"} processor_latency_ns{label="",path="root.pipeline.processors.0"} processor_received{label="",path="root.pipeline.processors.0"} processor_sent{label="",path="root.pipeline.processors.0"} output_batch_sent{label="bar",path="root.output"} output_connection_failed{label="bar",path="root.output"} output_connection_lost{label="bar",path="root.output"} output_connection_up{label="bar",path="root.output"} output_error{label="bar",path="root.output"} output_latency_ns{label="bar",path="root.output"} output_sent{label="bar",path="root.output"} ``` ## [](#metric-mapping)Metric mapping Since Redpanda Connect emits a large variety of metrics it is often useful to restrict or modify the metrics that are emitted. This can be done using the [Bloblang mapping language](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) in the field `metrics.mapping`. This is a mapping executed for each metric that is registered within the Redpanda Connect service and allows you to delete an entire series, modify the series name and delete or modify individual labels. Within the mapping the input document (referenced by the keyword `this`) is a string value containing the metric name, and the resulting document (referenced by the keyword `root`) must be a string value containing the resulting name. As is standard in Bloblang mappings, if the value of `root` is not assigned within the mapping then the metric name remains unchanged. If the value of `root` is `deleted()` then the metric series is dropped. Labels can be referenced as metadata values with the function `meta`, where if the label does not exist in the series being mapped the value `null` is returned. Labels can be changed by using meta assignments, and can be assigned `deleted()` in order to remove them. For example, the following mapping removes all but the `label` label entirely, which reduces the cardinality of each series. It also renames the `label` (for some reason) so that labels containing meows now contain woofs. Finally, the mapping restricts the metrics emitted to only three series; one for the input count, one for processor errors, and one for the output count, it does this by looking up metric names in a static array of allowed names, and if not present the `root` is assigned `deleted()`: ```yaml metrics: mapping: | # Delete all pre-existing labels meta = deleted() # Re-add the `label` label with meows replaced with woofs meta label = meta("label").replace("meow", "woof") # Delete all metric series that aren't in our list root = if ![ "input_received", "processor_error", "output_sent", ].contains(this) { deleted() } prometheus: use_histogram_timing: false ``` --- # Page 100: none **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/none.md --- # none > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: none latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/metrics/none page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/metrics/none.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/metrics/none.adoc description: Disable metrics entirely. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Disable metrics entirely. ```yml # Config fields, showing default values metrics: none: {} mapping: "" ``` --- # Page 101: open_telemetry_collector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/open_telemetry_collector.md --- # open_telemetry_collector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: open_telemetry_collector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/metrics/open_telemetry_collector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/metrics/open_telemetry_collector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/metrics/open_telemetry_collector.adoc description: Send metrics to an Open Telemetry collector. page-git-created-date: "2026-05-28" page-git-modified-date: "2026-08-11" --- Send metrics to an [Open Telemetry collector](https://opentelemetry.io/docs/collector/). Exports Redpanda Connect metrics to one or more OpenTelemetry Collector endpoints over HTTP or gRPC for aggregation and onward export. Metrics are encoded using the OpenTelemetry Metrics protocol and can be sent to any collector endpoint that supports OTLP. You can configure multiple collector endpoints (both HTTP and gRPC simultaneously). All configured endpoints will receive the same metrics data. This is useful for redundancy or sending metrics to multiple observability platforms. ### Common ```yml metrics: open_telemetry_collector: service: benthos http: [] # No default (required) grpc: [] # No default (required) ``` ### Advanced ```yml metrics: open_telemetry_collector: service: benthos http: [] # No default (required) grpc: [] # No default (required) tags: {} ``` ## [](#fields)Fields ### [](#grpc)`grpc[]` A list of grpc collectors. **Type**: `array` ### [](#grpc-address)`grpc[].address` The endpoint of a collector to send events to. **Type**: `string` ```yaml # Examples: address: localhost:4317 ``` ### [](#grpc-secure)`grpc[].secure` Connect to the collector with client transport security **Type**: `bool` **Default**: `false` ### [](#http)`http[]` A list of http collectors. **Type**: `array` ### [](#http-address)`http[].address` The endpoint of a collector to send events to. **Type**: `string` ```yaml # Examples: address: localhost:4318 ``` ### [](#http-secure)`http[].secure` Connect to the collector over HTTPS **Type**: `bool` **Default**: `false` ### [](#service)`service` The name of the service in metrics. **Type**: `string` **Default**: `benthos` ### [](#tags)`tags` A map of tags to add to all exported spans and metrics. **Type**: `object` **Default**: `{}` ## [](#usage)Usage The most common setup uses a local OpenTelemetry Collector running as a sidecar or daemon, which then forwards metrics to your observability backend: ```yaml metrics: open_telemetry_collector: service: my-service-name grpc: - address: localhost:4317 ``` For production deployments with remote collectors, enable TLS: ```yaml metrics: open_telemetry_collector: service: my-service-name grpc: - address: otel-collector.example.com:4317 secure: true ``` Use the `tags` field to add labels to all exported metrics for filtering and grouping in your observability platform: ```yaml metrics: open_telemetry_collector: service: my-service-name grpc: - address: localhost:4317 tags: environment: production cluster: kafka-01 ``` --- # Page 102: prometheus **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/prometheus.md --- # prometheus > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: prometheus latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/metrics/prometheus page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/metrics/prometheus.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/metrics/prometheus.adoc description: Host endpoints (/metrics and /stats) for Prometheus scraping. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Host endpoints (`/metrics` and `/stats`) for Prometheus scraping. ### Common ```yml metrics: prometheus: {} ``` ### Advanced ```yml metrics: prometheus: use_histogram_timing: false histogram_buckets: [] summary_quantiles_objectives: - error: 0.05 quantile: 0.5 - error: 0.01 quantile: 0.9 - error: 0.001 quantile: 0.99 add_process_metrics: false add_go_metrics: false push_url: "" # No default (optional) push_interval: "" # No default (optional) push_job_name: benthos_push push_basic_auth: username: "" password: "" file_output_path: "" ``` ## [](#fields)Fields ### [](#add_go_metrics)`add_go_metrics` Whether to export Go runtime metrics such as GC pauses in addition to Redpanda Connect metrics. **Type**: `bool` **Default**: `false` ### [](#add_process_metrics)`add_process_metrics` Whether to export process metrics such as CPU and memory usage in addition to Redpanda Connect metrics. **Type**: `bool` **Default**: `false` ### [](#file_output_path)`file_output_path` An optional file path to write all prometheus metrics on service shutdown. **Type**: `string` **Default**: `""` ### [](#histogram_buckets)`histogram_buckets[]` Timing metrics histogram buckets (in seconds). If left empty defaults to DefBuckets ([https://pkg.go.dev/github.com/prometheus/client\_golang/prometheus#pkg-variables](https://pkg.go.dev/github.com/prometheus/client_golang/prometheus#pkg-variables)). Applicable when `use_histogram_timing` is set to `true`. **Type**: `array` **Default**: `[]` ### [](#push_basic_auth)`push_basic_auth` The Basic Authentication credentials. **Type**: `object` ### [](#push_basic_auth-password)`push_basic_auth.password` The Basic Authentication password. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#push_basic_auth-username)`push_basic_auth.username` The Basic Authentication username. **Type**: `string` **Default**: `""` ### [](#push_interval)`push_interval` The period of time between each push when sending metrics to a Push Gateway. **Type**: `string` ### [](#push_job_name)`push_job_name` An identifier for push jobs. **Type**: `string` **Default**: `benthos_push` ### [](#push_url)`push_url` An optional [Push Gateway URL](#push-gateway) to push metrics to. **Type**: `string` ### [](#summary_quantiles_objectives)`summary_quantiles_objectives[]` A list of timing metrics summary buckets (as quantiles). Applicable when `use_histogram_timing` is set to `false`. **Type**: `array` **Default**: ```yaml - error: 0.05 quantile: 0.5 - error: 0.01 quantile: 0.9 - error: 0.001 quantile: 0.99 ``` ```yaml # Examples: summary_quantiles_objectives: - error: 0.05 quantile: 0.5 - error: 0.01 quantile: 0.9 - error: 0.001 quantile: 0.99 ``` ### [](#summary_quantiles_objectives-error)`summary_quantiles_objectives[].error` Permissible margin of error for quantile calculations. Precise calculations in a streaming context (without prior knowledge of the full dataset) can be resource-intensive. To balance accuracy with computational efficiency, an error margin is introduced. For instance, if the 90th quantile (`0.9`) is determined to be `100ms` with a 1% error margin (`0.01`), the true value will fall within the `[99ms, 101ms]` range.) **Type**: `float` **Default**: `0` ### [](#summary_quantiles_objectives-quantile)`summary_quantiles_objectives[].quantile` Quantile value. **Type**: `float` **Default**: `0` ### [](#use_histogram_timing)`use_histogram_timing` Whether to export timing metrics as a histogram, if `false` a summary is used instead. When exporting histogram timings the delta values are converted from nanoseconds into seconds in order to better fit within bucket definitions. For more information on histograms and summaries refer to: [https://prometheus.io/docs/practices/histograms/](https://prometheus.io/docs/practices/histograms/). **Type**: `bool` **Default**: `false` ## [](#push-gateway)Push gateway The field `push_url` is optional and when set will trigger a push of metrics to a [Prometheus Push Gateway](https://prometheus.io/docs/instrumenting/pushing/) once Redpanda Connect shuts down. It is also possible to specify a `push_interval` which results in periodic pushes. The Push Gateway is useful for when Redpanda Connect instances are short lived. Do not include the "/metrics/jobs/…​" path in the push URL. If the Push Gateway requires HTTP Basic Authentication it can be configured with `push_basic_auth`. --- # Page 103: Outputs **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about.md --- # Outputs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Outputs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- An output config section looks like this: ```yaml output: label: my_s3_output aws_s3: bucket: TODO path: '${! meta("kafka_topic") }/${! json("message.id") }.json' # Optional list of processing steps processors: - mapping: '{"message":this,"meta":{"link_count":this.links.length()}}' ``` ## [](#back-pressure)Back pressure Redpanda Connect outputs apply back pressure to components upstream. This means if your output target starts blocking traffic Redpanda Connect will gracefully stop consuming until the issue is resolved. ## [](#retries)Retries When a Redpanda Connect output fails to send a message the error is propagated back up to the input, where depending on the protocol it will either be pushed back to the source as a Noack (e.g. AMQP) or will be reattempted indefinitely with the commit withheld until success (e.g. Kafka). It’s possible to instead have Redpanda Connect indefinitely retry an output until success with a [`retry`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/retry/) output. Some other outputs, such as the [`broker`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/broker/), might also retry indefinitely depending on their configuration. ## [](#dead-letter-queues)Dead letter queues It’s possible to create fallback outputs for when an output target fails using a [`fallback`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/) output: ```yaml output: fallback: - aws_sqs: url: https://sqs.us-west-2.amazonaws.com/TODO/TODO max_in_flight: 20 - http_client: url: http://backup:1234/dlq verb: POST ``` ## [](#multiplexing-outputs)Multiplexing outputs There are a few different ways of multiplexing in Redpanda Connect, here’s a quick run through: ### [](#interpolation-multiplexing)Interpolation multiplexing Some output fields support [field interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/), which is a super easy way to multiplex messages based on their contents in situations where you are multiplexing to the same service. For example, multiplexing against Kafka topics is a common pattern: ```yaml output: kafka: addresses: [ TODO:6379 ] topic: ${! meta("target_topic") } ``` Refer to the field documentation for a given output to see if it support interpolation. ### [](#switch-multiplexing)Switch multiplexing A more advanced form of multiplexing is to route messages to different output configurations based on a query. This is easy with the [`switch` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/switch/): ```yaml output: switch: cases: - check: this.type == "foo" output: amqp_1: urls: [ amqps://guest:guest@localhost:5672/ ] target_address: queue:/the_foos - check: this.type == "bar" output: gcp_pubsub: project: dealing_with_mike topic: mikes_bars - output: redis_streams: url: tcp://localhost:6379 stream: everything_else processors: - mapping: | root = this root.type = this.type.not_null() | "unknown" ``` ## [](#labels)Labels Outputs have an optional field `label` that can uniquely identify them in observability data such as metrics and logs. This can be useful when running configs with multiple outputs, otherwise their metrics labels will be generated based on their composition. For more information check out the [metrics documentation](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/). --- # Page 104: amqp_0_9 **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/amqp_0_9.md --- # amqp_0_9 > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: amqp_0_9 latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/amqp_0_9 page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/amqp_0_9.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/amqp_0_9.adoc description: Sends messages to an AMQP (0.91) exchange. AMQP is a messaging protocol used by various message brokers, including RabbitMQ.Connects to an AMQP (0.91) queue. AMQP is a messaging protocol used by various message brokers, including RabbitMQ. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends messages to an AMQP (0.91) exchange. AMQP is a messaging protocol used by various message brokers, including RabbitMQ. #### Common ```yml outputs: label: "" amqp_0_9: urls: [] # No default (required) exchange: "" # No default (required) key: "" type: "" metadata: exclude_prefixes: [] max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" amqp_0_9: urls: [] # No default (required) exchange: "" # No default (required) exchange_declare: enabled: false type: direct durable: true arguments: "" # No default (optional) key: "" type: "" content_type: application/octet-stream content_encoding: "" correlation_id: "" reply_to: "" expiration: "" message_id: "" user_id: "" app_id: "" metadata: exclude_prefixes: [] priority: "" max_in_flight: 64 persistent: false mandatory: false immediate: false timeout: "" tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] ``` The metadata fields from each message are delivered as headers. TLS is automatically enabled when connecting to an `amqps` URL. However, you can customize [TLS settings](#tls) if required. You can use [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) to dynamically set values for the following fields: `key`, `exchange`, and `type`. ## [](#fields)Fields ### [](#app_id)`app_id` Set an application ID for each message using a dynamic interpolated expression. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#content_encoding)`content_encoding` The content encoding attribute of each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#content_type)`content_type` The MIME type of each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `application/octet-stream` ### [](#correlation_id)`correlation_id` Set a unique correlation ID for each message using a dynamic interpolated expression to help match messages to responses. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#exchange)`exchange` The AMQP exchange to publish messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#exchange_declare)`exchange_declare` Passively declares the [target exchange](#exchange) to check whether an exchange with the specified name exists and is configured correctly. If the exchange exists, then the passive declaration verifies that fields specified in this object match its properties. If the target exchange does not exist, this output creates it. **Type**: `object` ### [](#exchange_declare-arguments)`exchange_declare.arguments` Arguments for server-specific implementations of the exchange (optional). You can use arguments to configure additional parameters for exchange types that require them. **Type**: `object` ```yaml # Examples: arguments: alternate-exchange: my-ae ``` ### [](#exchange_declare-durable)`exchange_declare.durable` Whether the declared exchange is durable. **Type**: `bool` **Default**: `true` ### [](#exchange_declare-enabled)`exchange_declare.enabled` Whether to enable exchange declaration. **Type**: `bool` **Default**: `false` ### [](#exchange_declare-type)`exchange_declare.type` The type of the exchange, which determines how messages are routed to queues. > 📝 **NOTE** > > Dots (`.`) in message keys are only enforced in routing keys and message types for `topic` exchanges. **Type**: `string` **Default**: `direct` **Options**: `direct`, `fanout`, `topic`, `headers`, `x-custom` ### [](#expiration)`expiration` Set the TTL of each message in milliseconds. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#immediate)`immediate` Whether to set the immediate flag on published messages. When set to `true`, if there are no active consumers for a queue, the message is dropped instead of waiting. **Type**: `bool` **Default**: `false` ### [](#key)`key` The binding key to set for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#mandatory)`mandatory` Whether to set the mandatory flag on published messages. When set to `true`, a published message that cannot be routed to any queues is returned to the sender. **Type**: `bool` **Default**: `false` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this number to improve throughput. **Type**: `int` **Default**: `64` ### [](#message_id)`message_id` Set a message ID for each message using a dynamic interpolated expression. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#metadata)`metadata` Configure which metadata values are added to messages as headers. This allows you to pass additional context information along with your messages. **Type**: `object` ### [](#metadata-exclude_prefixes)`metadata.exclude_prefixes[]` Provide a list of explicit metadata key prefixes to exclude when adding metadata to sent messages. **Type**: `array` **Default**: `[]` ### [](#persistent)`persistent` Whether to store delivered messages on disk. By default, message delivery is transient. **Type**: `bool` **Default**: `false` ### [](#priority)`priority` Set the priority of each message using a dynamic interpolated expression. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: priority: 0 # --- priority: ${! meta("amqp_priority") } # --- priority: ${! json("doc.priority") } ``` ### [](#reply_to)`reply_to` Set the name of the queue to which responses are sent using a dynamic interpolated expression. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#timeout)`timeout` The maximum period to wait for a message acknowledgment before abandoning it and attempting a resend. If this value is not set, the system waits indefinitely. **Type**: `string` **Default**: `""` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#type)`type` A custom message type to set for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#urls)`urls[]` A list of URLs to connect to. This input attempts to connect to each URL in the list, in order, until a successful connection is established. It then continues to use that URL until the connection is closed. If an item in the list contains commas, it is split into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "amqp://guest:guest@127.0.0.1:5672/" # --- urls: - "amqp://127.0.0.1:5672/,amqp://127.0.0.2:5672/" # --- urls: - "amqp://127.0.0.1:5672/" - "amqp://127.0.0.2:5672/" ``` ### [](#user_id)`user_id` Set the user ID to the name of the publisher. If this property is set by a publisher, its value must match the name of the user that opened the connection. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` --- # Page 105: arc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/arc.md --- # arc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: arc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/arc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/arc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/arc.adoc description: Writes data to an Arc database via the msgpack ingestion endpoint. page-git-created-date: "2026-04-20" page-git-modified-date: "2026-08-11" --- Writes data to an Arc database via the msgpack ingestion endpoint. This output sends data to an [Arc](https://github.com/Basekick-Labs/arc) columnar analytical database using its high-performance MessagePack ingestion endpoint. Arc supports two payload formats: - **columnar** (default): Transposes batched messages into column arrays. This is the recommended format, offering significantly faster ingestion. - **row**: Sends each message as an individual row record with fields and optional tags. Data is encoded as MessagePack and optionally compressed with zstd (recommended) or gzip before being sent to the Arc endpoint. > 📝 **NOTE** > > In columnar mode, all messages within a single batch must have the same set of fields. Arc validates that all column arrays have equal length and rejects batches with mismatched columns. Schema evolution across separate batches is fully supported. Use row format if messages within a batch have varying schemas. ### Common ```yml outputs: label: "" arc: base_url: "" # No default (required) timeout: 5s token: "" # No default (optional) database: default measurement: "" # No default (required) format: columnar tags_mapping: "" # No default (optional) compression: zstd max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ### Advanced ```yml outputs: label: "" arc: base_url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] proxy_url: "" disable_http2: false tps_limit: 0 tps_burst: 1 backoff: initial_interval: 1s max_interval: 30s max_retries: 3 tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s http: max_idle_conns: 100 max_idle_conns_per_host: 0 max_conns_per_host: 64 idle_conn_timeout: 1m30s tls_handshake_timeout: 10s expect_continue_timeout: 1s response_header_timeout: 0s disable_keep_alives: false disable_compression: false max_response_header_bytes: 1048576 max_response_body_bytes: 10485760 write_buffer_size: 4096 read_buffer_size: 4096 h2: strict_max_concurrent_requests: false max_decoder_header_table_size: 4096 max_encoder_header_table_size: 4096 max_read_frame_size: 16384 max_receive_buffer_per_connection: 1048576 max_receive_buffer_per_stream: 1048576 send_ping_timeout: 0s ping_timeout: 15s write_byte_timeout: 0s access_log_level: "" access_log_body_limit: 0 token: "" # No default (optional) database: default measurement: "" # No default (required) format: columnar timestamp_field: "" timestamp_unit: auto tags_mapping: "" # No default (optional) compression: zstd max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#fields)Fields ### [](#access_log_body_limit)`access_log_body_limit` Maximum bytes of request/response body to include in logs. 0 to skip body logging. **Type**: `int` **Default**: `0` ### [](#access_log_level)`access_log_level` Log level for HTTP request/response logging. Empty disables logging. **Type**: `string` **Default**: `""` **Options**: `` `, `TRACE ``, `DEBUG`, `INFO`, `WARN`, `ERROR` ### [](#backoff)`backoff` Adaptive backoff configuration for 429 (Too Many Requests) responses. Always active. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` Initial interval between retries on 429 responses. **Type**: `string` **Default**: `1s` ### [](#backoff-max_interval)`backoff.max_interval` Maximum interval between retries on 429 responses. **Type**: `string` **Default**: `30s` ### [](#backoff-max_retries)`backoff.max_retries` Maximum number of retries on 429 responses. **Type**: `int` **Default**: `3` ### [](#base_url)`base_url` Base URL of the target service (e.g., [https://api.example.com](https://api.example.com)). TLS is enabled automatically for https URLs. **Type**: `string` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#compression)`compression` Compression algorithm for the request body. `zstd` is recommended for best decompression performance in Arc. **Type**: `string` **Default**: `zstd` **Options**: `zstd`, `gzip`, `none` ### [](#database)`database` The target database name. **Type**: `string` **Default**: `default` ### [](#disable_http2)`disable_http2` Disable HTTP/2 and force HTTP/1.1. **Type**: `bool` **Default**: `false` ### [](#format)`format` The payload format. `columnar` transposes batch messages into column arrays for best performance. `row` sends each message as an individual record. **Type**: `string` **Default**: `columnar` **Options**: `columnar`, `row` ### [](#http)`http` HTTP transport settings controlling connection pooling, timeouts, and HTTP/2. **Type**: `object` ### [](#http-disable_compression)`http.disable_compression` Disable automatic decompression of gzip responses. **Type**: `bool` **Default**: `false` ### [](#http-disable_keep_alives)`http.disable_keep_alives` Disable HTTP keep-alive connections; each request uses a new connection. **Type**: `bool` **Default**: `false` ### [](#http-expect_continue_timeout)`http.expect_continue_timeout` Maximum time to wait for a server’s 100-continue response before sending the body. 0 means the body is sent immediately. **Type**: `string` **Default**: `1s` ### [](#http-h2)`http.h2` HTTP/2-specific transport settings. Only applied when HTTP/2 is enabled. **Type**: `object` ### [](#http-h2-max_decoder_header_table_size)`http.h2.max_decoder_header_table_size` Upper limit in bytes for the HPACK header table used to decode headers from the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-h2-max_encoder_header_table_size)`http.h2.max_encoder_header_table_size` Upper limit in bytes for the HPACK header table used to encode headers sent to the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-h2-max_read_frame_size)`http.h2.max_read_frame_size` Largest HTTP/2 frame this endpoint will read. Valid range: 16 KiB to 16 MiB. **Type**: `int` **Default**: `16384` ### [](#http-h2-max_receive_buffer_per_connection)`http.h2.max_receive_buffer_per_connection` Maximum flow-control window size in bytes for data received on a connection. Must be at least 64 KiB and less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-h2-max_receive_buffer_per_stream)`http.h2.max_receive_buffer_per_stream` Maximum flow-control window size in bytes for data received on a single stream. Must be less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-h2-ping_timeout)`http.h2.ping_timeout` Timeout waiting for a PING response before closing the connection. **Type**: `string` **Default**: `15s` ### [](#http-h2-send_ping_timeout)`http.h2.send_ping_timeout` Idle timeout after which a PING frame is sent to verify connection health. 0 disables health checks. **Type**: `string` **Default**: `0s` ### [](#http-h2-strict_max_concurrent_requests)`http.h2.strict_max_concurrent_requests` When true, new requests block when a connection’s concurrency limit is reached instead of opening a new connection. **Type**: `bool` **Default**: `false` ### [](#http-h2-write_byte_timeout)`http.h2.write_byte_timeout` Timeout for writing data to a connection. The timer resets whenever bytes are written. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-idle_conn_timeout)`http.idle_conn_timeout` How long an idle connection remains in the pool before being closed. 0 disables the timeout. **Type**: `string` **Default**: `1m30s` ### [](#http-max_conns_per_host)`http.max_conns_per_host` Maximum total connections (active + idle) per host. 0 means unlimited. **Type**: `int` **Default**: `64` ### [](#http-max_idle_conns)`http.max_idle_conns` Maximum total number of idle (keep-alive) connections across all hosts. 0 means unlimited. **Type**: `int` **Default**: `100` ### [](#http-max_idle_conns_per_host)`http.max_idle_conns_per_host` Maximum idle connections to keep per host. 0 (the default) uses GOMAXPROCS+1. **Type**: `int` **Default**: `0` ### [](#http-max_response_body_bytes)`http.max_response_body_bytes` Maximum bytes of response body the client will read. The response body is wrapped with a limit reader; reads beyond this cap return EOF. 0 disables the limit. **Type**: `int` **Default**: `10485760` ### [](#http-max_response_header_bytes)`http.max_response_header_bytes` Maximum bytes of response headers to allow. **Type**: `int` **Default**: `1048576` ### [](#http-read_buffer_size)`http.read_buffer_size` Size in bytes of the per-connection read buffer. **Type**: `int` **Default**: `4096` ### [](#http-response_header_timeout)`http.response_header_timeout` Maximum time to wait for response headers after writing the full request. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-tls_handshake_timeout)`http.tls_handshake_timeout` Maximum time to wait for a TLS handshake to complete. 0 disables the timeout. **Type**: `string` **Default**: `10s` ### [](#http-write_buffer_size)`http.write_buffer_size` Size in bytes of the per-connection write buffer. **Type**: `int` **Default**: `4096` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#measurement)`measurement` The measurement (table) name. Supports interpolation functions. **Type**: `string` ```yaml # Examples: measurement: cpu_metrics # --- measurement: ${!metadata("measurement")} # --- measurement: ${!json("type")} ``` ### [](#proxy_url)`proxy_url` HTTP proxy URL. Empty string disables proxying. **Type**: `string` **Default**: `""` ### [](#tags_mapping)`tags_mapping` An optional Bloblang mapping to extract tags from each message. Only used in `row` format. The result must be a `map[string]string`. **Type**: `string` ```yaml # Examples: tags_mapping: root = {"host": this.hostname, "region": this.region} ``` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` HTTP request timeout. **Type**: `string` **Default**: `5s` ### [](#timestamp_field)`timestamp_field` The field name within each message containing the timestamp. If empty, the current time is used. Supports Unix timestamps and RFC3339 strings. **Type**: `string` **Default**: `""` ### [](#timestamp_unit)`timestamp_unit` The unit of a numeric timestamp field. `auto` detects the unit based on magnitude. Ignored when `timestamp_field` is empty. **Type**: `string` **Default**: `auto` **Options**: `us`, `ms`, `s`, `ns`, `auto` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#token)`token` Bearer token for authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#tps_burst)`tps_burst` Maximum burst size for rate limiting. **Type**: `int` **Default**: `1` ### [](#tps_limit)`tps_limit` Rate limit in requests per second. 0 disables rate limiting. **Type**: `float` **Default**: `0` ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). --- # Page 106: aws_dynamodb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/aws_dynamodb.md --- # aws_dynamodb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_dynamodb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/aws_dynamodb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/aws_dynamodb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/aws_dynamodb.adoc description: Inserts items into a DynamoDB table. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Inserts items into a DynamoDB table. #### Common ```yml outputs: label: "" aws_dynamodb: table: "" # No default (required) string_columns: {} json_map_columns: {} max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" aws_dynamodb: table: "" # No default (required) string_columns: {} json_map_columns: {} ttl: "" ttl_key: "" max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) max_retries: 3 backoff: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s ``` The field `string_columns` is a map of column names to string values, where the values are [function interpolated](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) per message of a batch. This allows you to populate string columns of an item by extracting fields within the document payload or metadata like follows: ```yml string_columns: id: ${!json("id")} title: ${!json("body.title")} topic: ${!meta("kafka_topic")} full_content: ${!content()} ``` The field `json_map_columns` is a map of column names to json paths, where the [dot path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) is extracted from each document and converted into a map value. Both an empty path and the path `.` are interpreted as the root of the document. This allows you to populate map columns of an item like follows: ```yml json_map_columns: user: path.to.user whole_document: . ``` A column name can be empty: ```yml json_map_columns: "": . ``` In which case the top level document fields will be written at the root of the item, potentially overwriting previously defined column values. If a path is not found within a document the column will not be populated. ## [](#credentials)Credentials By default Redpanda Connect will use a shared credentials file when connecting to AWS services. It’s also possible to set them explicitly at the component level, allowing you to transfer data across accounts. You can find out more in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#backoff)`backoff` Control time intervals between retry attempts. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` The initial period to wait between retry attempts. The retry interval increases for each failed attempt, up to the `backoff.max_interval` value. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1s` ### [](#backoff-max_elapsed_time)`backoff.max_elapsed_time` The maximum period to wait before retry attempts are abandoned. If zero then no limit is used. **Type**: `string` **Default**: `30s` ### [](#backoff-max_interval)`backoff.max_interval` The maximum period to wait between retry attempts. **Type**: `string` **Default**: `5s` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#json_map_columns)`json_map_columns` A map of column keys to [field paths](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) pointing to value data within messages. **Type**: `object` **Default**: `{}` ```yaml # Examples: json_map_columns: user: path.to.user whole_document: . # --- json_map_columns: "": . ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#max_retries)`max_retries` The maximum number of retries before giving up on the request. If set to zero there is no discrete limit. **Type**: `int` **Default**: `3` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#string_columns)`string_columns` A map of column keys to string values to store. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: string_columns: full_content: ${!content()} id: ${!json("id")} title: ${!json("body.title")} topic: ${!meta("kafka_topic")} ``` ### [](#table)`table` The table to store messages in. **Type**: `string` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#ttl)`ttl` An optional TTL to set for items, calculated from the moment the message is sent. **Type**: `string` **Default**: `""` ### [](#ttl_key)`ttl_key` The column key to place the TTL value within. **Type**: `string` **Default**: `""` --- # Page 107: aws_kinesis_firehose **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/aws_kinesis_firehose.md --- # aws_kinesis_firehose > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_kinesis_firehose latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/aws_kinesis_firehose page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/aws_kinesis_firehose.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/aws_kinesis_firehose.adoc description: Sends messages to a Kinesis Firehose delivery stream. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends messages to a Kinesis Firehose delivery stream. #### Common ```yml outputs: label: "" aws_kinesis_firehose: stream: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" aws_kinesis_firehose: stream: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) max_retries: 0 backoff: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s ``` ## [](#credentials)Credentials By default Redpanda Connect will use a shared credentials file when connecting to AWS services. It’s also possible to set them explicitly at the component level, allowing you to transfer data across accounts. You can find out more in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#backoff)`backoff` Control time intervals between retry attempts. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` The initial period to wait between retry attempts. The retry interval increases for each failed attempt, up to the `backoff.max_interval` value. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1s` ### [](#backoff-max_elapsed_time)`backoff.max_elapsed_time` The maximum period to wait before retry attempts are abandoned. If zero then no limit is used. **Type**: `string` **Default**: `30s` ### [](#backoff-max_interval)`backoff.max_interval` The maximum period to wait between retry attempts. **Type**: `string` **Default**: `5s` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#max_retries)`max_retries` The maximum number of retries before giving up on the request. If set to zero there is no discrete limit. **Type**: `int` **Default**: `0` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#stream)`stream` The stream to publish messages to. **Type**: `string` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` --- # Page 108: aws_kinesis **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/aws_kinesis.md --- # aws_kinesis > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_kinesis latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/aws_kinesis page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/aws_kinesis.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/aws_kinesis.adoc description: Sends messages to a Kinesis stream. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends messages to a Kinesis stream. #### Common ```yml outputs: label: "" aws_kinesis: stream: "" # No default (required) partition_key: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" aws_kinesis: stream: "" # No default (required) partition_key: "" # No default (required) hash_key: "" # No default (optional) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) max_retries: 0 backoff: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s ``` Both the `partition_key`(required) and `hash_key` (optional) fields can be dynamically set using function interpolations described [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). When sending batched messages the interpolations are performed per message part. ## [](#credentials)Credentials By default Redpanda Connect will use a shared credentials file when connecting to AWS services. It’s also possible to set them explicitly at the component level, allowing you to transfer data across accounts. You can find out more in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#backoff)`backoff` Control time intervals between retry attempts. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` The initial period to wait between retry attempts. The retry interval increases for each failed attempt, up to the `backoff.max_interval` value. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1s` ### [](#backoff-max_elapsed_time)`backoff.max_elapsed_time` The maximum period to wait before retry attempts are abandoned. If zero then no limit is used. **Type**: `string` **Default**: `30s` ### [](#backoff-max_interval)`backoff.max_interval` The maximum period to wait between retry attempts. **Type**: `string` **Default**: `5s` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#hash_key)`hash_key` A optional hash key for partitioning messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of parallel message batches to have in flight at any given time. **Type**: `int` **Default**: `64` ### [](#max_retries)`max_retries` The maximum number of retries before giving up on the request. If set to zero there is no discrete limit. **Type**: `int` **Default**: `0` ### [](#partition_key)`partition_key` A required key for partitioning messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#stream)`stream` The stream to publish messages to. Streams can either be specified by their name or full ARN. **Type**: `string` ```yaml # Examples: stream: foo # --- stream: arn:aws:kinesis:*:111122223333:stream/my-stream ``` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` --- # Page 109: aws_s3 **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/aws_s3.md --- # aws_s3 > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_s3 latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/aws_s3 page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/aws_s3.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/aws_s3.adoc description: Sends message parts as objects to an Amazon S3 bucket. Each object is uploaded with the path specified with the path field. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Uploads messages to an Amazon S3 bucket as objects, using the path specified in the `path` field. #### Common ```yml outputs: label: "" aws_s3: bucket: "" # No default (required) path: ${!counter()}-${!timestamp_unix_nano()}.txt tags: {} content_type: application/octet-stream metadata: exclude_prefixes: [] max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" aws_s3: bucket: "" # No default (required) path: ${!counter()}-${!timestamp_unix_nano()}.txt tags: {} content_type: application/octet-stream content_encoding: "" cache_control: "" content_disposition: "" content_language: "" website_redirect_location: "" metadata: exclude_prefixes: [] storage_class: STANDARD kms_key_id: "" checksum_algorithm: "" server_side_encryption: "" force_path_style_urls: false max_in_flight: 64 timeout: 5s object_canned_acl: "" batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` To use a different path for each object, use [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), which is evaluated for each message in a batch. ## [](#metadata)Metadata Metadata fields on messages will be sent as headers, in order to mutate these values (or remove them) check out the [metadata docs](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/metadata/). ## [](#tags)Tags The `tags` field accepts key/value pairs to attach to objects as tags, and the values support [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries): ```yaml output: aws_s3: bucket: TODO path: ${!counter()}-${!timestamp_unix_nano()}.tar.gz tags: Key1: Value1 Timestamp: ${!meta("Timestamp")} ``` ## [](#credentials)Credentials By default, Redpanda Connect uses a shared credentials file when connecting to AWS services. You can also set credentials explicitly at the component level to transfer data across accounts. You can find out more in [AWS credentials](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#batching)Batching It’s common to want to upload messages to S3 as batched archives. The easiest way to do this is to batch your messages at the output level and join the batch of messages with an [`archive`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive/) or [`compress`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/compress/) processor. For example, the following configuration uploads messages as a `.tar.gz` archive of documents: ```yaml output: aws_s3: bucket: TODO path: ${!counter()}-${!timestamp_unix_nano()}.tar.gz batching: count: 100 period: 10s processors: - archive: format: tar - compress: algorithm: gzip ``` This configuration uploads JSON documents as a single large document containing an array of objects: ```yaml output: aws_s3: bucket: TODO path: ${!counter()}-${!timestamp_unix_nano()}.json batching: count: 100 processors: - archive: format: json_array ``` ## [](#bucket-name-format)Bucket name format The `bucket` field accepts a bucket name only, not an ARN. For example, use `my-bucket`, not `arn:aws:s3:::my-bucket`. ## [](#s3-compatible-storage)S3-compatible storage The `endpoint` and `force_path_style_urls` fields let you connect to S3-compatible storage services such as Cloudflare R2, MinIO, or DigitalOcean Spaces. For Cloudflare R2, set `endpoint` to your account endpoint URL and enable `force_path_style_urls`: ```yaml output: aws_s3: bucket: r2-bucket path: ${!uuid_v4()}.json endpoint: https://.r2.cloudflarestorage.com force_path_style_urls: true region: auto credentials: id: secret: ``` Find your account ID in the Cloudflare dashboard under **R2 > Overview > Account Details**. Generate API credentials under **R2 > Manage R2 API Tokens**. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#batching-2)`batching` Configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#bucket)`bucket` The bucket to upload messages to. **Type**: `string` ### [](#cache_control)`cache_control` The cache control to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#checksum_algorithm)`checksum_algorithm` The algorithm used to validate each object during its upload to the Amazon S3 bucket. **Type**: `string` **Default**: `""` **Options**: `CRC32`, `CRC32C`, `SHA1`, `SHA256` ### [](#content_disposition)`content_disposition` The content disposition to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#content_encoding)`content_encoding` An optional content encoding to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#content_language)`content_language` The content language to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#content_type)`content_type` The content type to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `application/octet-stream` ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#force_path_style_urls)`force_path_style_urls` Forces the client API to use path style URLs, which helps when connecting to custom endpoints. **Type**: `bool` **Default**: `false` ### [](#kms_key_id)`kms_key_id` An optional server-side encryption key. **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#metadata-2)`metadata` Specify criteria for which metadata values are attached to objects as headers. **Type**: `object` ### [](#metadata-exclude_prefixes)`metadata.exclude_prefixes[]` Provide a list of explicit metadata key prefixes to be excluded when adding metadata to sent messages. **Type**: `array` **Default**: `[]` ### [](#object_canned_acl)`object_canned_acl` The object canned ACL value. Leave empty to omit the ACL from upload requests, which is required for buckets that have ACLs disabled (the AWS default since 2023). **Type**: `string` **Default**: `""` **Options**: `` `, `private ``, `public-read`, `public-read-write`, `authenticated-read`, `aws-exec-read`, `bucket-owner-read`, `bucket-owner-full-control` ### [](#path)`path` The path of each message to upload. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `${!counter()}-${!timestamp_unix_nano()}.txt` ```yaml # Examples: path: ${!counter()}-${!timestamp_unix_nano()}.txt # --- path: ${!meta("kafka_key")}.json # --- path: ${!json("doc.namespace")}/${!json("doc.id")}.json ``` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#server_side_encryption)`server_side_encryption` An optional server-side encryption algorithm. **Type**: `string` **Default**: `""` ### [](#storage_class)`storage_class` The storage class to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `STANDARD` **Options**: `STANDARD`, `REDUCED_REDUNDANCY`, `GLACIER`, `STANDARD_IA`, `ONEZONE_IA`, `INTELLIGENT_TIERING`, `DEEP_ARCHIVE` ### [](#tags-2)`tags` Key/value pairs to store with the object as tags. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: tags: Key1: Value1 Timestamp: ${!meta("Timestamp")} ``` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` The maximum period to wait on an upload before abandoning it and reattempting. **Type**: `string` **Default**: `5s` ### [](#website_redirect_location)`website_redirect_location` The website redirect location to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` --- # Page 110: aws_sns **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/aws_sns.md --- # aws_sns > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_sns latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/aws_sns page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/aws_sns.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/aws_sns.adoc description: Sends messages to an AWS SNS topic. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends messages to an AWS SNS topic. #### Common ```yml outputs: label: "" aws_sns: topic_arn: "" # No default (required) message_group_id: "" # No default (optional) message_deduplication_id: "" # No default (optional) subject: "" # No default (optional) max_in_flight: 64 metadata: exclude_prefixes: [] ``` #### Advanced ```yml outputs: label: "" aws_sns: topic_arn: "" # No default (required) message_group_id: "" # No default (optional) message_deduplication_id: "" # No default (optional) subject: "" # No default (optional) max_in_flight: 64 metadata: exclude_prefixes: [] timeout: 5s region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` ## [](#credentials)Credentials By default Redpanda Connect will use a shared credentials file when connecting to AWS services. It’s also possible to set them explicitly at the component level, allowing you to transfer data across accounts. You can find out more in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#message_deduplication_id)`message_deduplication_id` An optional deduplication ID to set for messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#message_group_id)`message_group_id` An optional group ID to set for messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#metadata)`metadata` Specify criteria for which metadata values are sent as headers. **Type**: `object` ### [](#metadata-exclude_prefixes)`metadata.exclude_prefixes[]` Provide a list of explicit metadata key prefixes to be excluded when adding metadata to sent messages. **Type**: `array` **Default**: `[]` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#subject)`subject` An optional subject to set for messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` The maximum period to wait on an upload before abandoning it and reattempting. **Type**: `string` **Default**: `5s` ### [](#topic_arn)`topic_arn` The topic to publish to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 111: aws_sqs **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/aws_sqs.md --- # aws_sqs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_sqs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/aws_sqs page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/aws_sqs.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/aws_sqs.adoc description: Sends messages to an SQS queue. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends messages to an SQS queue. #### Common ```yml outputs: label: "" aws_sqs: url: "" # No default (required) message_group_id: "" # No default (optional) message_deduplication_id: "" # No default (optional) delay_seconds: "" # No default (optional) max_in_flight: 64 metadata: exclude_prefixes: [] batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" aws_sqs: url: "" # No default (required) message_group_id: "" # No default (optional) message_deduplication_id: "" # No default (optional) delay_seconds: "" # No default (optional) max_in_flight: 64 metadata: exclude_prefixes: [] batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_records_per_request: 10 region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) max_retries: 0 backoff: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s ``` Metadata values are sent along with the payload as attributes with the data type String. If the number of metadata values in a message exceeds the message attribute limit (10) then the top ten keys ordered alphabetically will be selected. The fields `message_group_id`, `message_deduplication_id` and `delay_seconds` can be set dynamically using [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), which are resolved individually for each message of a batch. ## [](#credentials)Credentials By default Redpanda Connect will use a shared credentials file when connecting to AWS services. It’s also possible to set them explicitly at the component level, allowing you to transfer data across accounts. You can find out more in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#backoff)`backoff` Control time intervals between retry attempts. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` The initial period to wait between retry attempts. The retry interval increases for each failed attempt, up to the `backoff.max_interval` value. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `1s` ### [](#backoff-max_elapsed_time)`backoff.max_elapsed_time` The maximum period to wait before retry attempts are abandoned. If zero then no limit is used. **Type**: `string` **Default**: `30s` ### [](#backoff-max_interval)`backoff.max_interval` The maximum period to wait between retry attempts. **Type**: `string` **Default**: `5s` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#delay_seconds)`delay_seconds` An optional delay time in seconds for message. Value between 0 and 900 This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of parallel message batches to have in flight at any given time. **Type**: `int` **Default**: `64` ### [](#max_records_per_request)`max_records_per_request` The maximum number of records delivered in a single SQS request. Enter only values from `0` to `10`. **Type**: `int` **Default**: `10` ### [](#max_retries)`max_retries` The maximum number of retries before giving up on the request. If set to zero there is no discrete limit. **Type**: `int` **Default**: `0` ### [](#message_deduplication_id)`message_deduplication_id` An optional deduplication ID to set for messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#message_group_id)`message_group_id` An optional group ID to set for messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#metadata)`metadata` Specify criteria for which metadata values are sent as headers. **Type**: `object` ### [](#metadata-exclude_prefixes)`metadata.exclude_prefixes[]` Provide a list of explicit metadata key prefixes to be excluded when adding metadata to sent messages. **Type**: `array` **Default**: `[]` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#url)`url` The URL of the target SQS queue. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 112: azure_blob_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/azure_blob_storage.md --- # azure_blob_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_blob_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/azure_blob_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/azure_blob_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/azure_blob_storage.adoc description: Sends message parts as objects to an Azure Blob Storage Account container. Each object is uploaded with the filename specified with the container field. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends message parts as objects to an Azure Blob Storage Account container. Each object is uploaded with the filename specified with the `container` field. #### Common ```yml outputs: label: "" azure_blob_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" container: "" # No default (required) path: ${!counter()}-${!timestamp_unix_nano()}.txt max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" azure_blob_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" container: "" # No default (required) path: ${!counter()}-${!timestamp_unix_nano()}.txt blob_type: BLOCK public_access_level: PRIVATE max_in_flight: 64 ``` In order to have a different path for each object you should use function interpolations described [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), which are calculated per message of a batch. Supports multiple authentication methods but only one of the following is required: - `storage_connection_string` - `storage_account` and `storage_access_key` - `storage_account` and `storage_sas_token` - `storage_account` to access via [DefaultAzureCredential](https://pkg.go.dev/github.com/Azure/azure-sdk-for-go/sdk/azidentity#DefaultAzureCredential) If multiple are set then the `storage_connection_string` is given priority. If the `storage_connection_string` does not contain the `AccountName` parameter, please specify it in the `storage_account` field. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#blob_type)`blob_type` Block and Append blobs are comprized of blocks, and each blob can support up to 50,000 blocks. The default value is ``"`BLOCK`"``.\` This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `BLOCK` **Options**: `BLOCK`, `APPEND` ### [](#container)`container` The container for uploading the messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: container: messages-${!timestamp("2006")} ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#path)`path` The path of each message to upload. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `${!counter()}-${!timestamp_unix_nano()}.txt` ```yaml # Examples: path: ${!counter()}-${!timestamp_unix_nano()}.json # --- path: ${!meta("kafka_key")}.json # --- path: ${!json("doc.namespace")}/${!json("doc.id")}.json ``` ### [](#public_access_level)`public_access_level` The container’s public access level. The default value is `PRIVATE`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `PRIVATE` **Options**: `PRIVATE`, `BLOB`, `CONTAINER` ### [](#storage_access_key)`storage_access_key` The storage account access key. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_account)`storage_account` The storage account to access. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_connection_string)`storage_connection_string` A storage account connection string. This field is required if `storage_account` and `storage_access_key` / `storage_sas_token` are not set. **Type**: `string` **Default**: `""` ### [](#storage_sas_token)`storage_sas_token` The storage account SAS token. This field is ignored if `storage_connection_string` or `storage_access_key` are set. **Type**: `string` **Default**: `""` --- # Page 113: azure_cosmosdb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/azure_cosmosdb.md --- # azure_cosmosdb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_cosmosdb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/azure_cosmosdb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/azure_cosmosdb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/azure_cosmosdb.adoc description: Creates or updates messages as JSON documents in Azure CosmosDB. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Creates or updates messages as JSON documents in [Azure CosmosDB](https://learn.microsoft.com/en-us/azure/cosmos-db/introduction). ### Common ```yml outputs: label: "" azure_cosmosdb: endpoint: "" # No default (optional) account_key: "" # No default (optional) connection_string: "" # No default (optional) database: "" # No default (required) container: "" # No default (required) partition_keys_map: "" # No default (required) operation: Create item_id: "" # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 64 ``` ### Advanced ```yml outputs: label: "" azure_cosmosdb: endpoint: "" # No default (optional) account_key: "" # No default (optional) connection_string: "" # No default (optional) database: "" # No default (required) container: "" # No default (required) partition_keys_map: "" # No default (required) operation: Create patch_operations: [] # No default (optional) patch_condition: "" # No default (optional) auto_id: true item_id: "" # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 64 ``` When creating documents, each message must have the `id` property (case-sensitive) set (or use `auto_id: true`). It is the unique name that identifies the document, that is, no two documents share the same `id` within a logical partition. The `id` field must not exceed 255 characters. [See details](https://learn.microsoft.com/en-us/rest/api/cosmos-db/documents). The `partition_keys` field must resolve to the same value(s) across the entire message batch. ## [](#credentials)Credentials You can use one of the following authentication mechanisms: - Set the `endpoint` field and the `account_key` field - Set only the `endpoint` field to use [DefaultAzureCredential](https://pkg.go.dev/github.com/Azure/azure-sdk-for-go/sdk/azidentity#DefaultAzureCredential) - Set the `connection_string` field ## [](#batching)Batching CosmosDB limits the maximum batch size to 100 messages and the payload must not exceed 2MB ([details here](https://learn.microsoft.com/en-us/azure/cosmos-db/concepts-limits#per-request-limits)). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#examples)Examples ### [](#create-documents)Create documents Create new documents in the `blobfish` container with partition key `/habitat`. ```yaml output: azure_cosmosdb: endpoint: http://localhost:8080 account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== database: blobbase container: blobfish partition_keys_map: root = json("habitat") operation: Create ``` ### [](#patch-documents)Patch documents Execute the Patch operation on documents from the `blobfish` container. ```yaml output: azure_cosmosdb: endpoint: http://localhost:8080 account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== database: testdb container: blobfish partition_keys_map: root = json("habitat") item_id: ${! json("id") } operation: Patch patch_operations: # Add a new /diet field - operation: Add path: /diet value_map: root = json("diet") # Remove the first location from the /locations array field - operation: Remove path: /locations/0 # Add new location at the end of the /locations array field - operation: Add path: /locations/- value_map: root = "Challenger Deep" ``` ## [](#fields)Fields ### [](#account_key)`account_key` Account key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== ``` ### [](#auto_id)`auto_id` Automatically set the item `id` field to a random UUID v4. If the `id` field is already set, then it will not be overwritten. Setting this to `false` can improve performance, since the messages will not have to be parsed. **Type**: `bool` **Default**: `true` ### [](#batching-2)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#connection_string)`connection_string` Connection string. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: connection_string: AccountEndpoint=https://localhost:8081/;AccountKey=C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw==; ``` ### [](#container)`container` Container. **Type**: `string` ```yaml # Examples: container: testcontainer ``` ### [](#database)`database` Database. **Type**: `string` ```yaml # Examples: database: testdb ``` ### [](#endpoint)`endpoint` CosmosDB endpoint. **Type**: `string` ```yaml # Examples: endpoint: https://localhost:8081 ``` ### [](#item_id)`item_id` ID of item to replace or delete. Only used by the Replace and Delete operations This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: item_id: ${! json("id") } ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#operation)`operation` Operation. **Type**: `string` **Default**: `Create` | Option | Summary | | --- | --- | | Create | Create operation. | | Delete | Delete operation. | | Patch | Patch operation. | | Replace | Replace operation. | | Upsert | Upsert operation. | ### [](#partition_keys_map)`partition_keys_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to a single partition key value or an array of partition key values of type string, integer or boolean. Currently, hierarchical partition keys are not supported so only one value may be provided. **Type**: `string` ```yaml # Examples: partition_keys_map: root = "blobfish" # --- partition_keys_map: root = 41 # --- partition_keys_map: root = true # --- partition_keys_map: root = null # --- partition_keys_map: root = json("blobfish").depth ``` ### [](#patch_condition)`patch_condition` Patch operation condition. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: patch_condition: from c where not is_defined(c.blobfish) ``` ### [](#patch_operations)`patch_operations[]` Patch operations to be performed when `operation: Patch` . **Type**: `array` ### [](#patch_operations-operation)`patch_operations[].operation` Operation. **Type**: `string` **Default**: `Add` | Option | Summary | | --- | --- | | Add | Add patch operation. | | Increment | Increment patch operation. | | Remove | Remove patch operation. | | Replace | Replace patch operation. | | Set | Set patch operation. | ### [](#patch_operations-path)`patch_operations[].path` Path. **Type**: `string` ```yaml # Examples: path: /foo/bar/baz ``` ### [](#patch_operations-value_map)`patch_operations[].value_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to a value of any type that is supported by CosmosDB. **Type**: `string` ```yaml # Examples: value_map: root = "blobfish" # --- value_map: root = 41 # --- value_map: root = true # --- value_map: root = json("blobfish").depth # --- value_map: root = [1, 2, 3] ``` ## [](#cosmosdb-emulator)CosmosDB emulator If you wish to run the CosmosDB emulator that is referenced in the documentation [here](https://learn.microsoft.com/en-us/azure/cosmos-db/linux-emulator), the following Docker command should do the trick: ```bash > docker run --rm -it -p 8081:8081 --name=cosmosdb -e AZURE_COSMOS_EMULATOR_PARTITION_COUNT=10 -e AZURE_COSMOS_EMULATOR_ENABLE_DATA_PERSISTENCE=false mcr.microsoft.com/cosmosdb/linux/azure-cosmos-emulator ``` Note: `AZURE_COSMOS_EMULATOR_PARTITION_COUNT` controls the number of partitions that will be supported by the emulator. The bigger the value, the longer it takes for the container to start up. Additionally, instead of installing the container self-signed certificate which is exposed via `[https://localhost:8081/_explorer/emulator.pem](https://localhost:8081/_explorer/emulator.pem)`, you can run [mitmproxy](https://mitmproxy.org/) like so: ```bash > mitmproxy -k --mode "reverse:https://localhost:8081" ``` Then you can access the CosmosDB UI via `[http://localhost:8080/_explorer/index.html](http://localhost:8080/_explorer/index.html)` and use `[http://localhost:8080](http://localhost:8080)` as the CosmosDB endpoint. --- # Page 114: azure_data_lake_gen2 **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/azure_data_lake_gen2.md --- # azure_data_lake_gen2 > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_data_lake_gen2 latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/azure_data_lake_gen2 page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/azure_data_lake_gen2.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/azure_data_lake_gen2.adoc description: Sends message parts as files to an Azure Data Lake Gen2 filesystem. Each file is uploaded with the filename specified with the path field. page-git-created-date: "2024-11-05" page-git-modified-date: "2026-05-26" --- Sends message parts as files to an [Azure Data Lake Gen2](https://learn.microsoft.com/en-us/azure/storage/blobs/data-lake-storage-introduction) file system. Each file is uploaded with the file name specified in the `path` field. ```yml outputs: label: "" azure_data_lake_gen2: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" filesystem: "" # No default (required) path: ${!counter()}-${!timestamp_unix_nano()}.txt max_in_flight: 64 ``` To specify a different [`path` value](#path) (file name) for each file, use [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). Function interpolations are calculated for each message in a batch. ## [](#authentication-methods)Authentication methods This output supports multiple authentication methods. You must configure at least one method from the following list: - `storage_connection_string` - `storage_account` and `storage_access_key` - `storage_account` and `storage_sas_token` - `storage_account` to access using [DefaultAzureCredential](https://pkg.go.dev/github.com/Azure/azure-sdk-for-go/sdk/azidentity#DefaultAzureCredential) If you configure multiple authentication methods, the `storage_connection_string` takes precedence. ## [](#performance)Performance Sends multiple messages in flight in parallel for improved performance. You can tune the number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#filesystem)`filesystem` The name of the data lake storage file system you want to upload messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: filesystem: messages-${!timestamp("2006")} ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this number to improve throughput until performance plateaus. **Type**: `int` **Default**: `64` ### [](#path)`path` The path (file name) of each message to upload to the data lake storage file system. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `${!counter()}-${!timestamp_unix_nano()}.txt` ```yaml # Examples: path: ${!counter()}-${!timestamp_unix_nano()}.json # --- path: ${!meta("kafka_key")}.json # --- path: ${!json("doc.namespace")}/${!json("doc.id")}.json ``` ### [](#storage_access_key)`storage_access_key` The access key for the storage account. Use this field along with `storage_account` for authentication. This field is ignored when the `storage_connection_string` field is populated. **Type**: `string` **Default**: `""` ### [](#storage_account)`storage_account` The storage account to access. This field is ignored when the `storage_connection_string` field is populated. **Type**: `string` **Default**: `""` ### [](#storage_connection_string)`storage_connection_string` The connection string for the storage account. You must enter a value for this field if no other authentication method is specified. > 📝 **NOTE** > > If the `storage_connection_string` field does not contain the `AccountName` parameter value, specify it in the `storage_account` field. **Type**: `string` **Default**: `""` ### [](#storage_sas_token)`storage_sas_token` The SAS token for the storage account. Use this field along with `storage_account` for authentication. This field is ignored when either the `storage_connection_string` or `storage_access_key` fields are populated. **Type**: `string` **Default**: `""` --- # Page 115: azure_queue_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/azure_queue_storage.md --- # azure_queue_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_queue_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/azure_queue_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/azure_queue_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/azure_queue_storage.adoc description: Sends messages to an Azure Storage Queue. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends messages to an Azure Storage Queue. #### Common ```yml outputs: label: "" azure_queue_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" queue_name: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" azure_queue_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" queue_name: "" # No default (required) ttl: "" max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` Only one authentication method is required, `storage_connection_string` or `storage_account` and `storage_access_key`. If both are set then the `storage_connection_string` is given priority. In order to set the `queue_name` you can use function interpolations described [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), which are calculated per message of a batch. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#max_in_flight)`max_in_flight` The maximum number of parallel message batches to have in flight at any given time. **Type**: `int` **Default**: `64` ### [](#queue_name)`queue_name` The name of the target Queue Storage queue. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#storage_access_key)`storage_access_key` The storage account access key. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_account)`storage_account` The storage account to access. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_connection_string)`storage_connection_string` A storage account connection string. This field is required if `storage_account` and `storage_access_key` / `storage_sas_token` are not set. **Type**: `string` **Default**: `""` ### [](#storage_sas_token)`storage_sas_token` The storage account SAS token. This field is ignored if `storage_connection_string` or `storage_access_key` are set. **Type**: `string` **Default**: `""` ### [](#ttl)`ttl` The TTL of each individual message as a duration string. Defaults to 0, meaning no retention period is set This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: ttl: 60s # --- ttl: 5m # --- ttl: 36h ``` --- # Page 116: azure_table_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/azure_table_storage.md --- # azure_table_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_table_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/azure_table_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/azure_table_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/azure_table_storage.adoc description: Stores messages in an Azure Table Storage table. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores messages in an Azure Table Storage table. #### Common ```yml outputs: label: "" azure_table_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" table_name: "" # No default (required) partition_key: "" row_key: "" properties: {} max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" azure_table_storage: storage_account: "" storage_access_key: "" storage_connection_string: "" storage_sas_token: "" table_name: "" # No default (required) partition_key: "" row_key: "" properties: {} transaction_type: INSERT max_in_flight: 64 timeout: 5s batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` Only one authentication method is required, `storage_connection_string` or `storage_account` and `storage_access_key`. If both are set then the `storage_connection_string` is given priority. In order to set the `table_name`, `partition_key` and `row_key` you can use function interpolations described [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), which are calculated per message of a batch. If the `properties` are not set in the config, all the `json` fields are marshalled and stored in the table, which will be created if it does not exist. The `object` and `array` fields are marshaled as strings. e.g.: The JSON message: ```json { "foo": 55, "bar": { "baz": "a", "bez": "b" }, "diz": ["a", "b"] } ``` Will store in the table the following properties: ```yml foo: '55' bar: '{ "baz": "a", "bez": "b" }' diz: '["a", "b"]' ``` It’s also possible to use function interpolations to get or transform the properties values, e.g.: ```yml properties: device: '${! json("device") }' timestamp: '${! json("timestamp") }' ``` ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#max_in_flight)`max_in_flight` The maximum number of parallel message batches to have in flight at any given time. **Type**: `int` **Default**: `64` ### [](#partition_key)`partition_key` The partition key. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: partition_key: ${! json("date") } ``` ### [](#properties)`properties` A map of properties to store into the table. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ### [](#row_key)`row_key` The row key. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: row_key: ${! json("device")}-${!uuid_v4() } ``` ### [](#storage_access_key)`storage_access_key` The storage account access key. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_account)`storage_account` The storage account to access. This field is ignored if `storage_connection_string` is set. **Type**: `string` **Default**: `""` ### [](#storage_connection_string)`storage_connection_string` A storage account connection string. This field is required if `storage_account` and `storage_access_key` / `storage_sas_token` are not set. **Type**: `string` **Default**: `""` ### [](#storage_sas_token)`storage_sas_token` The storage account SAS token. This field is ignored if `storage_connection_string` or `storage_access_key` are set. **Type**: `string` **Default**: `""` ### [](#table_name)`table_name` The table to store messages into. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: table_name: ${! meta("kafka_topic") } # --- table_name: ${! json("table") } ``` ### [](#timeout)`timeout` The maximum period to wait on an upload before abandoning it and reattempting. **Type**: `string` **Default**: `5s` ### [](#transaction_type)`transaction_type` Type of transaction operation. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `INSERT` **Options**: `INSERT`, `INSERT_MERGE`, `INSERT_REPLACE`, `UPDATE_MERGE`, `UPDATE_REPLACE`, `DELETE` ```yaml # Examples: transaction_type: ${! json("operation") } # --- transaction_type: ${! meta("operation") } # --- transaction_type: INSERT ``` --- # Page 117: broker **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/broker.md --- # broker > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: broker latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/broker page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/broker.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/broker.adoc description: Allows you to route messages to multiple child outputs using a range of brokering patterns. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A meta-output that routes messages to child outputs using a range of brokering [patterns](#patterns). Unlike regular outputs, `broker` doesn’t send messages anywhere by itself. Instead, it wraps other outputs and controls how messages are delivered across them. Use `broker` to fan out the same message to multiple destinations (for example, publishing events to Kafka while also writing them to a database), or to distribute messages across a pool of outputs for load balancing or throughput scaling. The delivery pattern determines whether each message is written to all outputs or routed to a single output, and whether writes happen in parallel or in sequence. > 📝 **NOTE** > > The name `broker` refers to the brokering delivery pattern, not a Redpanda broker (cluster node). #### Common ```yml outputs: label: "" broker: pattern: fan_out outputs: [] # No default (required) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" broker: copies: 1 pattern: fan_out outputs: [] # No default (required) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` [Processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) can be listed to apply across individual outputs or all outputs: ```yaml output: broker: pattern: fan_out outputs: - resource: foo - resource: bar # Processors only applied to messages sent to bar. processors: - resource: bar_processor # Processors applied to messages sent to all brokered outputs. processors: - resource: general_processor ``` ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#copies)`copies` The number of copies of each configured output to spawn. **Type**: `int` **Default**: `1` ### [](#outputs)`outputs[]` A list of child outputs to broker. **Type**: `array` ### [](#pattern)`pattern` The brokering pattern to use. **Type**: `string` **Default**: `fan_out` **Options**: `fan_out`, `fan_out_fail_fast`, `fan_out_sequential`, `fan_out_sequential_fail_fast`, `round_robin`, `greedy` ## [](#patterns)Patterns The broker pattern determines how messages are distributed across outputs. Use `fan_out` (the default) when every output should receive every message. Use `round_robin` or `greedy` when you want to distribute messages across outputs for load balancing rather than duplication. The available patterns are: ### [](#fan_out)`fan_out` With the fan out pattern all outputs will be sent every message that passes through Redpanda Connect in parallel. If an output applies back pressure it will block all subsequent messages, and if an output fails to send a message it will be retried continuously until completion or service shut down. This mechanism is in place in order to prevent one bad output from causing a larger retry loop that results in a good output from receiving unbounded message duplicates. Sometimes it is useful to disable the back pressure or retries of certain fan out outputs and instead drop messages that have failed or were blocked. In this case you can wrap outputs with a [`drop_on` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/drop_on/). ### [](#fan_out_fail_fast)`fan_out_fail_fast` The same as the `fan_out` pattern, except that output failures will not be automatically retried. This pattern should be used with caution as busy retry loops could result in unlimited duplicates being introduced into the non-failure outputs. ### [](#fan_out_sequential)`fan_out_sequential` Similar to the fan out pattern except outputs are written to sequentially, meaning an output is only written to once the preceding output has confirmed receipt of the same message. If an output applies back pressure it will block all subsequent messages, and if an output fails to send a message it will be retried continuously until completion or service shut down. This mechanism is in place in order to prevent one bad output from causing a larger retry loop that results in a good output from receiving unbounded message duplicates. ### [](#fan_out_sequential_fail_fast)`fan_out_sequential_fail_fast` The same as the `fan_out_sequential` pattern, except that output failures will not be automatically retried. This pattern should be used with caution as busy retry loops could result in unlimited duplicates being introduced into the non-failure outputs. ### [](#round_robin)`round_robin` With the round robin pattern each message will be assigned a single output following their order. If an output applies back pressure it will block all subsequent messages. If an output fails to send a message then the message will be re-attempted with the next input, and so on. ### [](#greedy)`greedy` The greedy pattern results in higher output throughput at the cost of potentially disproportionate message allocations to those outputs. Each message is sent to a single output, which is determined by allowing outputs to claim messages as soon as they are able to process them. This results in certain faster outputs potentially processing more messages at the cost of slower outputs. --- # Page 118: cache **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/cache.md --- # cache > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: cache latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/cache page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/cache.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/cache.adoc description: Stores each message in a cache. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Stores each message in a [cache](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/). #### Common ```yml outputs: label: "" cache: target: "" # No default (required) key: ${!count("items")}-${!timestamp_unix_nano()} max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" cache: target: "" # No default (required) key: ${!count("items")}-${!timestamp_unix_nano()} ttl: "" # No default (optional) max_in_flight: 64 ``` Caches are configured as [resources](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/), where there’s a wide variety to choose from. The `target` field must reference a configured cache resource label like follows: ```yaml output: cache: target: foo key: ${!json("document.id")} cache_resources: - label: foo memcached: addresses: - localhost:11211 default_ttl: 60s ``` In order to create a unique `key` value per item you should use function interpolations described in [Bloblang queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#key)`key` The key to store messages by, function interpolation should be used in order to derive a unique key for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `${!count("items")}-${!timestamp_unix_nano()}` ```yaml # Examples: key: ${!count("items")}-${!timestamp_unix_nano()} # --- key: ${!json("doc.id")} # --- key: ${!meta("kafka_key")} ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#target)`target` The target cache to store messages in. **Type**: `string` ### [](#ttl)`ttl` The TTL of each individual item as a duration string. After this period an item will be eligible for removal during the next compaction. Not all caches support per-key TTLs, and those that do not will fall back to their generally configured TTL setting. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: ttl: 60s # --- ttl: 5m # --- ttl: 36h ``` --- # Page 119: cyborgdb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/cyborgdb.md --- # cyborgdb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: cyborgdb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/cyborgdb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/cyborgdb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/cyborgdb.adoc description: Inserts items into a CyborgDB encrypted vector index. page-git-created-date: "2025-10-09" page-git-modified-date: "2026-08-11" --- Inserts items into a CyborgDB encrypted vector index. #### Common ```yaml outputs: label: "" cyborgdb: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) host: "" # No default (required) api_key: "" # No default (required) index_name: redpanda-vectors index_key: "" # No default (required) operation: upsert id: "" # No default (required) vector_mapping: "" # No default (optional) metadata_mapping: "" # No default (optional) ``` #### Advanced ```yaml outputs: label: "" cyborgdb: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) host: "" # No default (required) api_key: "" # No default (required) index_name: redpanda-vectors index_key: "" # No default (required) create_if_missing: false operation: upsert id: "" # No default (required) vector_mapping: "" # No default (optional) metadata_mapping: "" # No default (optional) ``` This output allows you to write vectors to a CyborgDB encrypted index. CyborgDB provides end-to-end encrypted vector storage with automatic dimension detection and index optimization. All vector data is encrypted client-side before being sent to the server, ensuring complete data privacy. The encryption key never leaves your infrastructure. ## [](#fields)Fields ### [](#api_key)`api_key` The API key for authenticating with the CyborgDB service. This key identifies your account and provides access to your CyborgDB indexes. Keep this key secure and avoid exposing it in logs or version control. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#create_if_missing)`create_if_missing` Whether to create the index if it doesn’t exist. When enabled, CyborgDB automatically detects the vector dimensions from your data and optimizes the index configuration for performance. This is useful for development and testing environments. **Type**: `bool` **Default**: `false` ### [](#host)`host` The host URL for the CyborgDB instance. This should include the protocol (https://) and port number if required. For example: `[https://api.cyborgdb.com](https://api.cyborgdb.com)` or `[https://localhost:8080](https://localhost:8080)`. **Type**: `string` ```yaml # Examples: host: api.cyborg.com # --- host: localhost:8000 ``` ### [](#id)`id` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that determines the unique identifier for each vector entry. This ID is used to update existing vectors during upsert operations or to specify which vectors to delete. If not provided, CyborgDB will generate unique IDs automatically. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#index_key)`index_key` The base64-encoded encryption key for the CyborgDB index. This key must be exactly 32 bytes when decoded from base64. All vector data is encrypted client-side using this key before transmission, ensuring complete data privacy. Store this key securely as it cannot be recovered if lost. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: index_key: your-base64-encoded-32-byte-key ``` ### [](#index_name)`index_name` The name of the CyborgDB index to write vectors to. If the index doesn’t exist and `create_if_missing` is enabled, CyborgDB will create it automatically with optimized settings based on your data. **Type**: `string` **Default**: `redpanda-vectors` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#metadata_mapping)`metadata_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that extracts metadata to associate with the vector entry. The metadata can contain any JSON-serializable data that helps identify or categorize the vector. This data is stored encrypted alongside the vector. **Type**: `string` ```yaml # Examples: metadata_mapping: root = @ # --- metadata_mapping: root = metadata() # --- metadata_mapping: root = {"summary": this.summary, "category": this.category} ``` ### [](#operation)`operation` The operation to perform against the CyborgDB index. Supported operations: - `upsert`: Insert new vectors or update existing ones (requires `vector_mapping`) - `delete`: Remove vectors from the index (requires `id`) - `query`: Search for similar vectors (requires `vector_mapping`) **Type**: `string` **Default**: `upsert` **Options**: `upsert`, `delete` ### [](#vector_mapping)`vector_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that extracts the vector from the message. The result must be an array of floating-point numbers representing the vector embeddings. This field is required for `upsert` and `query` operations. **Type**: `string` ```yaml # Examples: vector_mapping: root = this.embeddings_vector # --- vector_mapping: root = [1.2, 0.5, 0.76] ``` --- # Page 120: drop_on **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/drop_on.md --- # drop_on > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: drop_on latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/drop_on page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/drop_on.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/drop_on.adoc description: Attempts to write messages to a child output and if the write fails for one of a list of configurable reasons the message is dropped (acked) instead of being reattempted (or nacked). page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Attempts to write messages to a child output and if the write fails for one of a list of configurable reasons the message is dropped (acked) instead of being reattempted (or nacked). ```yml outputs: label: "" drop_on: error: false error_patterns: [] # No default (optional) back_pressure: "" # No default (optional) output: "" # No default (required) ``` Regular Redpanda Connect outputs will apply back pressure when downstream services aren’t accessible, and Redpanda Connect retries (or nacks) all messages that fail to be delivered. However, in some circumstances, or for certain output types, we instead might want to relax these mechanisms, which is when this output becomes useful. ## [](#fields)Fields ### [](#back_pressure)`back_pressure` An optional duration string that determines the maximum length of time to wait for a given message to be accepted by the child output before the message should be dropped instead. The most common reason for an output to block is when waiting for a lost connection to be re-established. Once a message has been dropped due to back pressure all subsequent messages are dropped immediately until the output is ready to process them again. Note that if `error` is set to `false` and this field is specified then messages dropped due to back pressure will return an error response (are nacked or reattempted). **Type**: `string` ```yaml # Examples: back_pressure: 30s # --- back_pressure: 1m ``` ### [](#error)`error` Whether messages should be dropped when the child output returns an error of any type. For example, this could be when an `http_client` output gets a 4XX response code. In order to instead drop only on specific error patterns use the `error_matches` field instead. **Type**: `bool` **Default**: `false` ### [](#error_patterns)`error_patterns[]` A list of regular expressions (re2) where if the child output returns an error that matches any part of any of these patterns the message will be dropped. **Type**: `array` ```yaml # Examples: error_patterns: - "and that was really bad$" # --- error_patterns: - "roughly [0-9]+ issues occurred" ``` ### [](#output)`output` A child output to wrap with this drop mechanism. **Type**: `output` nclude::connect:components:partial$examples/outputs/drop\_on.adoc\[\] --- # Page 121: drop **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/drop.md --- # drop > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: drop latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/drop page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/drop.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/drop.adoc description: Drop output reference for silently discarding messages in Redpanda Connect pipelines. page-topic-type: reference personas: streaming_developer, app_developer learning-objective-1: Look up drop output syntax and configuration learning-objective-2: Find examples of drop in conditional routing learning-objective-3: Identify use cases for drop output in pipelines page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Silently discards all messages without error or side effects. The `drop` output is a utility component that drops messages from the [pipeline](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#pipeline). Unlike filtering or conditional processors that modify or route messages, `drop` consumes messages and does nothing with them. This is useful for: - **Testing and debugging**: Measure input throughput without output bottlenecks - **Conditional workflows**: Discard messages that don’t meet certain criteria - **Dead letter queue patterns**: Provide a final fallback when all other outputs fail - **Development**: Temporarily disable output while testing pipeline logic Use this reference to: - Look up drop output syntax and configuration - Find examples of drop in conditional routing - Identify use cases for drop output in pipelines ```yaml outputs: label: "" drop: {} ``` ## [](#performance)Performance The `drop` output has minimal overhead and immediately acknowledges messages. This makes it ideal for performance testing, as it removes output processing time from measurements. ## [](#examples)Examples ### [](#conditional-filtering)Conditional filtering Use `drop` with the [`switch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/switch/) output to conditionally discard messages: ```yaml output: switch: cases: - check: this.type == "error" output: label: error_sink kafka: addresses: ["kafka:9092"] topic: errors - check: this.type == "debug" output: label: drop_debug drop: {} # Don't process debug messages in production - output: label: main_sink kafka: addresses: ["kafka:9092"] topic: events ``` ### [](#dead-letter-queue-pattern)Dead letter queue pattern Use `drop` as a last resort in a [`fallback`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/) chain: ```yaml output: fallback: - kafka: addresses: ["kafka:9092"] topic: primary_topic max_in_flight: 1 - kafka: addresses: ["kafka:9092"] topic: dlq_topic - drop: {} # Last resort: drop if both outputs fail ``` ### [](#testing-input-throughput)Testing input throughput Measure how fast your input can consume data without output bottlenecks: ```yaml input: kafka: addresses: ["kafka:9092"] topics: ["test"] output: drop: {} # Measure input consumption speed without output overhead ``` For more patterns on message routing, filtering, and when to use `drop` vs. other approaches, see the [Message Routing Patterns](https://docs.redpanda.com/connect/cookbooks/message_routing/) cookbook. --- # Page 122: elasticsearch_v8 **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/elasticsearch_v8.md --- # elasticsearch_v8 > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: elasticsearch_v8 latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/elasticsearch_v8 page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/elasticsearch_v8.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/elasticsearch_v8.adoc page-git-created-date: "2025-03-12" page-git-modified-date: "2026-05-26" --- Publishes messages into an [Elasticsearch index](https://www.elastic.co/guide/en/elasticsearch/reference/current/documents-indices.html). If the index does not exist, this output creates it using dynamic mapping. > 📝 **NOTE** > > The `elasticsearch_v8` output is based on the the [go-elasticsearch/v8](https://github.com/elastic/go-elasticsearch?tab=readme-ov-file) library. For full information about breaking changes from previous versions, see [Elastic’s Migrating to 8.0 guide](https://www.elastic.co/guide/en/elasticsearch/reference/current/migrating-8.0.html#breaking_80_rest_api_changes). To help configure your own `elasticsearch_v8` output, this page includes [example pipeline configurations](#example-pipelines). ### Common ```yml outputs: label: "" elasticsearch_v8: urls: [] # No default (required) index: "" # No default (required) action: "" # No default (required) id: "" # No default (required) max_in_flight: 64 api_key: "" batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ### Advanced ```yml outputs: label: "" elasticsearch_v8: urls: [] # No default (required) index: "" # No default (required) action: "" # No default (required) id: "" # No default (required) pipeline: "" routing: "" retry_on_conflict: 0 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] max_in_flight: 64 api_key: "" basic_auth: enabled: false username: "" password: "" batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#set-values-dynamically)Set values dynamically You can use [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) to dynamically set values for the [`id`](#id) and [`index`](#index) fields, as well as other fields where [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) are supported. When message batches are sent, interpolations are performed per message. ## [](#performance)Performance For improved performance, this output sends: - Multiple messages in parallel. Adjust the `max_in_flight` field value to tune the maximum number of in-flight messages (or message batches). - Messages as batches. You can configure batches at both input and output level. For more information, see [Message Batching](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#action)`action` The action to perform on each document. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). For more information on how the `update` action works, see [Example pipelines](#example-pipelines). **Type**: `string` ### [](#api_key)`api_key` An API key to authenticate with. If set, it supersedes basic authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth)`basic_auth` Configure basic authentication credentials for connecting to Elasticsearch. When enabled, these credentials are sent with each request to authenticate with the cluster. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#batching)`batching` Configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#id)`id` Define the ID for indexed messages. Use [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) to dynamically create a unique ID for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: id: ${!counter()}-${!timestamp_unix()} ``` ### [](#index)`index` The Elasticsearch index where messages are published. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#pipeline)`pipeline` Specify the ID of a pipeline to preprocess incoming documents before they are published (optional). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#retry_on_conflict)`retry_on_conflict` The number of times to retry an update operation when a version conflict occurs. **Type**: `int` **Default**: `0` ### [](#routing)`routing` The routing key to use for the document. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether to enable TLS for secure connections. Set to `true` to enable TLS encryption. Required to be `true` for other TLS options (like `client_certs`, `root_cas`, etc.) to take effect. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. This output attempts to connect to each URL in the list, in order, until a successful connection is established. If an item in the list contains commas, it is split into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "http://localhost:9200" ``` ## [](#example-pipelines)Example pipelines ### Update documents To update documents in the target index, the top level of the request body must include at least one of the following fields: - `doc`: Performs partial updates on a document. - `upsert`: Updates an existing document or inserts a document if it doesn’t exist. - `script`: Performs an update using a scripting language, such as [Elasticsearch’s Painless scripting language](https://www.elastic.co/guide/en/elasticsearch/reference/current/modules-scripting-painless.html). The following examples show how to configure mapping processors with this output to achieve different types of updates. Example 1: Partial document update ```yaml output: processors: # Sets the metadata ID field to the message ID then # performs a partial update on the document. - mapping: | meta id = this.id root.doc = this elasticsearch_v8: urls: [localhost:9200] # The URL of the Elasticsearch server. index: my_target_index # The name of the Elasticsearch index. id: ${! @id } # Sets the document ID to the value of the metadata ID field. action: update # The action to perform on each document. ``` Example 2: Scripted update ```yaml output: processors: # Sets the metadata ID field to the message ID then # increments the counter field by `1` using a script. - mapping: | meta id = this.id root.script.source = "ctx._source.counter += 1" elasticsearch_v8: urls: [localhost:9200] # The URL of the Elasticsearch server. index: my_target_index # The name of the Elasticsearch index. id: ${! @id } # Sets the document ID to the value of the metadata ID field. action: update # The action to perform on each document. ``` Example 3: Upsert ```yaml output: processors: # Sets the metadata ID field to the message ID. # If the product with the specified ID exists, update its product_price to 100. # If the document does not exist, insert a new document with the ID set to 1 # and the `product_price` set to 50. - mapping: | meta id = this.id root.doc.product_price = 100 root.upsert.product_price = 50 elasticsearch_v8: urls: [localhost:9200] # The URL of the Elasticsearch server. index: my_target_index # The name of the Elasticsearch index. id: ${! @id } # Sets the document ID to the value of the metadata ID field. action: update # The action to perform on each document. ``` For more information on the structures and behaviors of `doc`, `upsert`, and `script` fields, see the [Elasticsearch Update API](https://www.elastic.co/guide/en/elasticsearch/reference/current/docs-update.html). ### Index documents from Redpanda Reads messages from a Redpanda cluster and writes them to an Elasticsearch index using a field from the message as the document ID. ```yaml # Reads messages from a Redpanda cluster. input: redpanda: seed_brokers: [localhost:19092] # The address of the Redpanda broker. topics: ["product_code"] # The topic to consume messages from. consumer_group: "rpcn3" # The consumer group ID. processors: # Sets the metadata ID field to the message ID and # sets the root of the message to the message content. - mapping: | meta id = this.id root = this # Writes messages to the specified Elasticsearch index. output: elasticsearch_v8: urls: ['http://localhost:9200'] # The URL of the Elasticsearch server. index: "product_code" # The name of the Elasticsearch index. action: "index" # The action to perform on each document. id: ${! meta("id") } # Sets the document ID to the value of the metadata ID field. ``` ### Index documents from AWS S3 Reads messages from a AWS S3 bucket and writes them to an Elasticsearch index using the S3 key as the ID for the Elasticsearch document. ```yaml # Reads messages from an AWS S3 bucket. input: aws_s3: bucket: "my_bucket" # The name of the S3 bucket. prefix: "prod_inventory/" # A prefix to filter objects in the bucket. scanner: to_the_end: {} # Scans the bucket to the end. # Writes messages to the specified Elasticsearch index. output: elasticsearch_v8: urls: ['http://localhost:9200'] # The URL of the Elasticsearch server. index: "current_prod_inventory" # The name of the Elasticsearch index. action: "index" # The action to perform on each document. id: ${! meta("s3_key") } # Sets the document ID to the S3 key. ``` --- # Page 123: fallback **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback.md --- # fallback > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: fallback latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/fallback page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/fallback.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/fallback.adoc description: Attempts to send each message to a child output, starting from the first output on the list. If an output attempt fails then the next output in the list is attempted, and so on. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Attempts to send each message to a child output, starting from the first output on the list. If an output attempt fails then the next output in the list is attempted, and so on. ```yml outputs: label: "" fallback: - label: "" stdout: codec: lines - label: "" file: path: /tmp/fallback.txt codec: lines ``` This pattern is useful for triggering events in the case where certain output targets have broken. For example, if you had an output type `http_client` but wished to reroute messages whenever the endpoint becomes unreachable you could use this pattern: ```yaml output: fallback: - http_client: url: http://foo:4195/post/might/become/unreachable retries: 3 retry_period: 1s - http_client: url: http://bar:4196/somewhere/else retries: 3 retry_period: 1s processors: - mapping: 'root = "failed to send this message to foo: " + content()' - file: path: /usr/local/benthos/everything_failed.jsonl ``` ## [](#metadata)Metadata When a given output fails the message routed to the following output will have a metadata value named `fallback_error` containing a string error message outlining the cause of the failure. The content of this string will depend on the particular output and can be used to enrich the message or provide information used to broker the data to an appropriate output using something like a `switch` output. ## [](#batching)Batching When an output within a fallback sequence uses batching, like so: ```yaml output: fallback: - aws_dynamodb: table: foo string_columns: id: ${!json("id")} content: ${!content()} batching: count: 10 period: 1s - file: path: /usr/local/benthos/failed_stuff.jsonl ``` Redpanda Connect makes a best attempt at inferring which specific messages of the batch failed, and only propagates those individual messages to the next fallback tier. However, depending on the output and the error returned it is sometimes not possible to determine the individual messages that failed, in which case the whole batch is passed to the next tier in order to preserve at-least-once delivery guarantees. --- # Page 124: gcp_bigquery_write_api **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/gcp_bigquery_write_api.md --- # gcp_bigquery_write_api > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_bigquery_write_api latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/gcp_bigquery_write_api page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/gcp_bigquery_write_api.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/gcp_bigquery_write_api.adoc description: Streams data into BigQuery using the Storage Write API. page-git-created-date: "2026-05-28" page-git-modified-date: "2026-08-11" --- Streams data into BigQuery using the Storage Write API. Writes messages to a BigQuery table using the Storage Write API. This provides higher throughput and lower latency than the legacy streaming API or load jobs. Messages can be formatted as JSON (default) or raw protobuf bytes. When using JSON format the component automatically fetches the table schema and converts each message to the corresponding proto representation. > ⚠️ **WARNING** > > The proto3 JSON mapping encodes int64 and uint64 values as strings. JSON messages with integer fields must use string values (e.g. `"age": "30"` not `"age": 30`). Otherwise the write will fail with an unmarshalling error. When batching is enabled the table name is resolved from the first message in each batch. All messages in the same batch are written to that table. #### Common ```yml outputs: label: "" gcp_bigquery_write_api: project: "" dataset: "" # No default (required) table: "" # No default (required) message_format: json change_type: "" # No default (optional) change_sequence_number: "" # No default (optional) primary_keys: [] # No default (optional) max_in_flight: 4 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) credentials_json: "" ``` #### Advanced ```yml outputs: label: "" gcp_bigquery_write_api: project: "" dataset: "" # No default (required) table: "" # No default (required) message_format: json write_mode: default_stream change_type: "" # No default (optional) change_sequence_number: "" # No default (optional) primary_keys: [] # No default (optional) auto_create_table: false schema: [] time_partitioning: type: "" # No default (optional) field: "" expiration: 0s require_filter: false clustering: [] max_in_flight: 4 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) credentials_json: "" target_principal: "" delegates: [] stream_idle_timeout: 5m stream_sweep_interval: 1m max_cached_streams: 1024 schema_resolve_timeout: 15s schema_evolution_timeout: 30s endpoint: http: "" grpc: "" ``` ## [](#fields)Fields ### [](#auto_create_table)`auto_create_table` If true and the target table does not exist, the output creates it using the configured `schema`, `time_partitioning`, and `clustering`. AlreadyExists errors from concurrent creators are treated as success. When the table name is interpolated, every auto-created table receives the same schema and partition/clustering settings. **Type**: `bool` **Default**: `false` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#change_sequence_number)`change_sequence_number` Optional Bloblang expression resolving to the `_CHANGE_SEQUENCE_NUMBER` pseudo-column value. Format: 1 to 4 sections of 1 to 16 hexadecimal characters each, separated by `/`. Example: `${! metadata("scn") }` or `${! "0/0/0/0" }`. When unset, BigQuery resolves ordering by arrival time. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#change_type)`change_type` Bloblang expression resolving to the `_CHANGE_TYPE` pseudo-column value for each row. Must resolve to `UPSERT` or `DELETE` (case-insensitive). Required when `write_mode` is `upsert` or `upsert_delete`. Example: `${! metadata("operation") }`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#clustering)`clustering[]` Optional clustering columns (up to 4) applied during `auto_create_table`. All names must appear in `schema`. **Type**: `array` **Default**: `[]` ### [](#credentials_json)`credentials_json` An optional JSON string containing GCP credentials. If empty, credentials are loaded from the environment. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#dataset)`dataset` The BigQuery dataset ID. **Type**: `string` ### [](#delegates)`delegates[]` Optional delegation chain for chained service account impersonation. Each service account must be granted roles/iam.serviceAccountTokenCreator on the next in the chain. **Type**: `array` **Default**: `[]` ### [](#endpoint)`endpoint` Optional endpoint overrides for the BigQuery and Storage Write API clients. **Type**: `object` ### [](#endpoint-grpc)`endpoint.grpc` Override the BigQuery Storage gRPC endpoint. Useful for local emulators. **Type**: `string` **Default**: `""` ### [](#endpoint-http)`endpoint.http` Override the BigQuery HTTP endpoint. Useful for local emulators. **Type**: `string` **Default**: `""` ### [](#max_cached_streams)`max_cached_streams` Soft cap on the number of cached streams. When the cache exceeds this size, the least-recently-used stream is evicted. Set to 0 for unlimited (rely on idle-timeout sweeping only). Relevant when the table field uses interpolation to route to many tables. **Type**: `int` **Default**: `1024` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `4` ### [](#message_format)`message_format` The format of input messages. Use 'json' to have the component convert JSON to proto automatically. Use 'protobuf' to supply raw proto-encoded bytes. **Type**: `string` **Default**: `json` **Options**: `json`, `protobuf` ### [](#primary_keys)`primary_keys[]` Optional list of primary-key column names. Required when `auto_create_table` is true and `write_mode` is `upsert` or `upsert_delete`. A pre-existing table must already declare its PRIMARY KEY — this field cannot add one; when both are set they must match exactly (same columns, same order). Up to 16 columns; composite keys are supported in the same order they are listed. **Type**: `array` ### [](#project)`project` The GCP project ID. If empty, the project is auto-detected from the environment. **Type**: `string` **Default**: `""` ### [](#schema)`schema[]` Column definitions used by `auto_create_table`. Required when `auto_create_table` is true. **Type**: `array` **Default**: `[]` ### [](#schema-fields)`schema[].fields[]` For RECORD columns, the list of nested fields. Same shape as the top-level schema list. **Type**: `array` ### [](#schema-mode)`schema[].mode` Column mode: NULLABLE (default), REQUIRED, or REPEATED. **Type**: `string` **Default**: `NULLABLE` ### [](#schema-name)`schema[].name` Column name. **Type**: `string` ### [](#schema-type)`schema[].type` BigQuery column type (STRING, BYTES, INTEGER/INT64, FLOAT/FLOAT64, NUMERIC, BIGNUMERIC, BOOLEAN/BOOL, TIMESTAMP, DATE, TIME, DATETIME, GEOGRAPHY, JSON, RECORD). **Type**: `string` ### [](#schema_evolution_timeout)`schema_evolution_timeout` Total time budget for a single schema evolution attempt (Metadata + Update across all CAS retries on HTTP 412). Bounds how long the WriteBatch retry loop can be starved by a wedged backend. **Type**: `string` **Default**: `30s` ### [](#schema_resolve_timeout)`schema_resolve_timeout` How long a single BigQuery table-metadata fetch can run before being aborted. Coalesced concurrent resolves share one fetch, so this bounds the time a wedged backend can stall every batch routing to the same table. On the auto\_create\_table path the budget covers Metadata→Create→Metadata, so it needs to absorb transient backend slowness on top of the metadata fetch itself. **Type**: `string` **Default**: `15s` ### [](#stream_idle_timeout)`stream_idle_timeout` How long a cached stream can remain unused before being closed. Relevant when the table field uses interpolation to route to many tables. **Type**: `string` **Default**: `5m` ### [](#stream_sweep_interval)`stream_sweep_interval` How often to check for idle streams to close. **Type**: `string` **Default**: `1m` ### [](#table)`table` The BigQuery table ID. Supports interpolation functions. When batching, resolved from the first message in each batch. **Type**: `string` ### [](#target_principal)`target_principal` Service account email to impersonate. When set, the output obtains tokens acting as this service account. Requires the caller to have roles/iam.serviceAccountTokenCreator on the target. **Type**: `string` **Default**: `""` ### [](#time_partitioning)`time_partitioning` Optional time-partitioning settings applied during `auto_create_table`. Setting `type` is the trigger — when omitted, the block is treated as absent. **Type**: `object` ### [](#time_partitioning-expiration)`time_partitioning.expiration` Optional partition expiration. Zero means no expiration. **Type**: `string` **Default**: `0s` ### [](#time_partitioning-field)`time_partitioning.field` Column to partition on. Must be of type DATE, TIMESTAMP, or DATETIME. If empty, the table uses ingestion-time partitioning (`_PARTITIONTIME`). **Type**: `string` **Default**: `""` ### [](#time_partitioning-require_filter)`time_partitioning.require_filter` If true, queries against the table must filter on the partition column. **Type**: `bool` **Default**: `false` ### [](#time_partitioning-type)`time_partitioning.type` Partitioning granularity. **Type**: `string` **Options**: `DAY`, `HOUR`, `MONTH`, `YEAR` ### [](#write_mode)`write_mode` How the output writes to BigQuery. `default_stream` uses the multiplexed default stream (at-least-once, lowest latency). `pending_stream` allocates a per-batch pending stream that commits atomically, providing exactly-once semantics within a single committed batch. `upsert` writes UPSERT-only rows to a BigQuery CDC-enabled table; the target table must have a PRIMARY KEY. `upsert_delete` allows both UPSERT and DELETE rows. Both CDC modes use the default stream as required by BigQuery. **Type**: `string` **Default**: `default_stream` **Options**: `default_stream`, `pending_stream`, `upsert`, `upsert_delete` --- # Page 125: gcp_bigquery **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/gcp_bigquery.md --- # gcp_bigquery > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_bigquery latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/gcp_bigquery page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/gcp_bigquery.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/gcp_bigquery.adoc description: Sends messages as new rows to a Google Cloud BigQuery table. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Inserts message data as new rows in a Google Cloud BigQuery table. #### Common ```yml outputs: label: "" gcp_bigquery: project: "" job_project: "" dataset: "" # No default (required) table: "" # No default (required) format: NEWLINE_DELIMITED_JSON max_in_flight: 64 job_labels: {} credentials_json: "" csv: header: [] field_delimiter: , allow_jagged_rows: false allow_quoted_newlines: false encoding: UTF-8 skip_leading_rows: 1 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" gcp_bigquery: project: "" job_project: "" dataset: "" # No default (required) table: "" # No default (required) format: NEWLINE_DELIMITED_JSON max_in_flight: 64 write_disposition: WRITE_APPEND create_disposition: CREATE_IF_NEEDED ignore_unknown_values: false max_bad_records: 0 auto_detect: false job_labels: {} credentials_json: "" csv: header: [] field_delimiter: , allow_jagged_rows: false allow_quoted_newlines: false encoding: UTF-8 skip_leading_rows: 1 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#credentials)Credentials By default, Redpanda Connect uses a [shared credentials file](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/gcp/) when connecting to GCP services. ## [](#format)Format The `gcp_bigquery` output currently supports only `NEWLINE_DELIMITED_JSON`, `CSV` and `PARQUET` formats. To learn more about how to use BigQuery with these formats, see the following documentation: - [`NEWLINE_DELIMITED_JSON`](https://cloud.google.com/bigquery/docs/loading-data-cloud-storage-json) - [`CSV`](https://cloud.google.com/bigquery/docs/loading-data-cloud-storage-csv) - [`PARQUET`](https://cloud.google.com/bigquery/docs/loading-data-cloud-storage-parquet) ### [](#newline-delimited-json)Newline-delimited JSON Each JSON message may contain multiple elements separated by newlines. For example, a single message containing: ```json {"key": "1"} {"key": "2"} ``` Is equivalent to two separate messages: ```json {"key": "1"} ``` And: ```json {"key": "2"} ``` The same is true for the CSV format. ### [](#csv)CSV When the field `csv.header` is specified for the `CSV` format, a header row is inserted as the first line of each message batch. If this field is not provided, then the first message of each message batch must include a header line. ### [](#parquet)Parquet Each message sent to this output must be a Parquet file. You can use the [`parquet_encode` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/parquet_encode/) to convert message data into the correct format. For example: ```yaml input: generate: mapping: | root = { "foo": random_int(), "bar": uuid_v4(), "time": now(), } interval: 0 count: 1000 batch_size: 1000 pipeline: processors: - parquet_encode: schema: - name: foo type: INT64 - name: bar type: UTF8 - name: time type: UTF8 default_compression: zstd output: gcp_bigquery: project: "${PROJECT}" dataset: "my_bq_dataset" table: "redpanda_connect_ingest" format: PARQUET ``` ## [](#performance)Performance The `gcp_bigquery` output benefits from sending multiple messages in parallel for improved performance. You can tune the maximum number of in-flight messages (or message batches) with the field `max_in_flight`. This output also sends messages as a batch for improved performance. Redpanda Connect can form batches at both the input and output level. For more information, see [Message Batching](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#auto_detect)`auto_detect` Whether this component automatically infers the options and schema for `CSV` and `NEWLINE_DELIMITED_JSON` sources. If this value is set to `false` and the destination table doesn’t exist, the output throws an insertion error as it is unable to insert data. > ⚠️ **CAUTION** > > This field delegates schema detection to the GCP BigQuery service. For the `CSV` format, values like `no` may be treated as booleans. **Type**: `bool` **Default**: `false` ### [](#batching)`batching` Configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#create_disposition)`create_disposition` Specifies the circumstances under which a destination table is created. - Use `CREATE_IF_NEEDED` to create the destination table if it does not already exist. Tables are created atomically on successful completion of a job. - Use `CREATE_NEVER` if the destination table must already exist. **Type**: `string` **Default**: `CREATE_IF_NEEDED` **Options**: `CREATE_IF_NEEDED`, `CREATE_NEVER` ### [](#credentials_json)`credentials_json` Sets the [Google Service Account Credentials JSON](https://developers.google.com/workspace/guides/create-credentials#create_credentials_for_a_service_account) (optional). > ⚠️ **WARNING** > > When using [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) to populate this field, wrap the function in single quotes, not double quotes. For example, use `'${secrets.GCP_CREDENTIALS_JSON}'` instead of `"${secrets.GCP_CREDENTIALS_JSON}"`. Double quotes cause JSON parsing errors because the credentials already contain JSON content. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#csv-2)`csv` Specify how CSV data is interpreted. **Type**: `object` ### [](#csv-allow_jagged_rows)`csv.allow_jagged_rows` Set to `true` to treat optional missing trailing columns as nulls in CSV data. **Type**: `bool` **Default**: `false` ### [](#csv-allow_quoted_newlines)`csv.allow_quoted_newlines` Whether quoted data sections containing new lines are allowed when reading CSV data. **Type**: `bool` **Default**: `false` ### [](#csv-encoding)`csv.encoding` The character encoding of CSV data. **Type**: `string` **Default**: `UTF-8` **Options**: `UTF-8`, `ISO-8859-1` ### [](#csv-field_delimiter)`csv.field_delimiter` The separator for fields in a CSV file. The output uses this value when reading or exporting data. **Type**: `string` **Default**: `,` ### [](#csv-header)`csv.header[]` A list of values to use as the header for each batch of messages. If not specified, the first line of each message is used as the header. **Type**: `array` **Default**: `[]` ### [](#csv-skip_leading_rows)`csv.skip_leading_rows` The number of rows at the top of a CSV file that BigQuery will skip when reading data. The default value is `1`, which allows Redpanda Connect to add the specified header in the first line of each batch sent to BigQuery. **Type**: `int` **Default**: `1` ### [](#dataset)`dataset` The BigQuery Dataset ID. **Type**: `string` ### [](#format-2)`format` The format of each incoming message. **Type**: `string` **Default**: `NEWLINE_DELIMITED_JSON` **Options**: `NEWLINE_DELIMITED_JSON`, `CSV`, `PARQUET` ### [](#ignore_unknown_values)`ignore_unknown_values` Set this value to `true` to ignore values that do not match the schema: - For the `CSV` format, extra values at the end of a line are ignored. - For the `NEWLINE_DELIMITED_JSON` format, values that do not match any column name are ignored. By default, this value is set to `false`, and records containing unknown values are treated as bad records. Use the `max_bad_records` field to customize how bad records are handled. **Type**: `bool` **Default**: `false` ### [](#job_labels)`job_labels` A list of labels to add to the load job. **Type**: `object` **Default**: `{}` ### [](#job_project)`job_project` Specify the project ID in which jobs are executed. If not set, the `project` value is used. **Type**: `string` **Default**: `""` ### [](#max_bad_records)`max_bad_records` The maximum number of bad records to ignore when reading data and [`ignore_unknown_values`](#ignore_unknown_values) is set to `true`. **Type**: `int` **Default**: `0` ### [](#max_in_flight)`max_in_flight` The maximum number of message batches to have in flight at a given time. Increase this value to improve throughput. **Type**: `int` **Default**: `64` ### [](#project)`project` Specify the project ID of the dataset to insert data into. If not set, the project ID is inferred from the project linked to the service account or read from the `GOOGLE_CLOUD_PROJECT` environment variable. **Type**: `string` **Default**: `""` ### [](#table)`table` The table to insert messages into. **Type**: `string` ### [](#write_disposition)`write_disposition` Specifies how existing data in a destination table is treated. **Type**: `string` **Default**: `WRITE_APPEND` **Options**: `WRITE_APPEND`, `WRITE_EMPTY`, `WRITE_TRUNCATE` --- # Page 126: gcp_cloud_storage **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/gcp_cloud_storage.md --- # gcp_cloud_storage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_cloud_storage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/gcp_cloud_storage page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/gcp_cloud_storage.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/gcp_cloud_storage.adoc description: Sends message parts as objects to a Google Cloud Storage bucket. Each object is uploaded with the path specified with the path field. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends message parts as objects to a Google Cloud Storage bucket. Each object is uploaded with the path specified with the `path` field. #### Common ```yml outputs: label: "" gcp_cloud_storage: bucket: "" # No default (required) path: ${!counter()}-${!timestamp_unix_nano()}.txt content_type: application/octet-stream collision_mode: overwrite timeout: 3s credentials_json: "" max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" gcp_cloud_storage: bucket: "" # No default (required) path: ${!counter()}-${!timestamp_unix_nano()}.txt content_type: application/octet-stream content_encoding: "" collision_mode: overwrite chunk_size: 16777216 timeout: 3s credentials_json: "" max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` In order to have a different path for each object you should use function interpolations described in [Bloblang queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), which are calculated per message of a batch. ## [](#metadata)Metadata Metadata fields on messages will be sent as headers, in order to mutate these values (or remove them) check out the [metadata docs](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/metadata/). ## [](#credentials)Credentials By default Redpanda Connect will use a shared credentials file when connecting to GCP services. You can find out more in [Google Cloud Platform](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/gcp/). ## [](#batching)Batching It’s common to want to upload messages to Google Cloud Storage as batched archives, the easiest way to do this is to batch your messages at the output level and join the batch of messages with an [`archive`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive/) and/or [`compress`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/compress/) processor. For example, if we wished to upload messages as a .tar.gz archive of documents we could achieve that with the following config: ```yaml output: gcp_cloud_storage: bucket: TODO path: ${!counter()}-${!timestamp_unix_nano()}.tar.gz batching: count: 100 period: 10s processors: - archive: format: tar - compress: algorithm: gzip ``` Alternatively, if we wished to upload JSON documents as a single large document containing an array of objects we can do that with: ```yaml output: gcp_cloud_storage: bucket: TODO path: ${!counter()}-${!timestamp_unix_nano()}.json batching: count: 100 processors: - archive: format: json_array ``` ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#batching-2)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#bucket)`bucket` The bucket to upload messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#chunk_size)`chunk_size` An optional chunk size which controls the maximum number of bytes of the object that the Writer will attempt to send to the server in a single request. If ChunkSize is set to zero, chunking will be disabled. **Type**: `int` **Default**: `16777216` ### [](#collision_mode)`collision_mode` Determines how file path collisions should be dealt with. Options are "overwrite", which replaces the existing file with the new one, "append", which appends the message bytes to the original file, "error-if-exists", which returns an error and rejects the message if the file exists, and "ignore", does not modify the original file and drops the message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `overwrite` **Options**: `overwrite`, `append`, `error-if-exists`, `ignore` ### [](#content_encoding)`content_encoding` An optional content encoding to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#content_type)`content_type` The content type to set for each object. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `application/octet-stream` ### [](#credentials_json)`credentials_json` Base64-encoded Google Service Account credentials in JSON format (optional). Use this field to authenticate with Google Cloud services. For more information about creating service account credentials, see [Google’s service account documentation](https://developers.google.com/workspace/guides/create-credentials#create_credentials_for_a_service_account). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of message batches to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#path)`path` The path of each message to upload. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `${!counter()}-${!timestamp_unix_nano()}.txt` ```yaml # Examples: path: ${!counter()}-${!timestamp_unix_nano()}.txt # --- path: ${!meta("kafka_key")}.json # --- path: ${!json("doc.namespace")}/${!json("doc.id")}.json ``` ### [](#timeout)`timeout` The maximum period to wait on an upload before abandoning it and reattempting. **Type**: `string` **Default**: `3s` ```yaml # Examples: timeout: 1s # --- timeout: 500ms ``` --- # Page 127: gcp_pubsub **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/gcp_pubsub.md --- # gcp_pubsub > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_pubsub latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/gcp_pubsub page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/gcp_pubsub.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/gcp_pubsub.adoc description: Sends messages to a GCP Cloud Pub/Sub topic. Metadata from messages are sent as attributes. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends messages to a GCP Cloud Pub/Sub topic. [Metadata](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/metadata/) from messages are sent as attributes. #### Common ```yml outputs: label: "" gcp_pubsub: project: "" # No default (required) credentials_json: "" topic: "" # No default (required) endpoint: "" max_in_flight: 64 count_threshold: 100 delay_threshold: 10ms byte_threshold: 1000000 metadata: exclude_prefixes: [] batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" gcp_pubsub: project: "" # No default (required) credentials_json: "" topic: "" # No default (required) endpoint: "" ordering_key: "" # No default (optional) max_in_flight: 64 count_threshold: 100 delay_threshold: 10ms byte_threshold: 1000000 publish_timeout: 1m0s validate_topic: true metadata: exclude_prefixes: [] flow_control: max_outstanding_bytes: -1 max_outstanding_messages: 1000 limit_exceeded_behavior: block batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` For information on how to set up credentials, see [this guide](https://cloud.google.com/docs/authentication/production). ## [](#troubleshooting)Troubleshooting If you’re consistently seeing `Failed to send message to gcp_pubsub: context deadline exceeded` error logs without any further information it is possible that you are encountering [https://github.com/benthosdev/benthos/issues/1042](https://github.com/benthosdev/benthos/issues/1042), which occurs when metadata values contain characters that are not valid utf-8. This can frequently occur when consuming from Kafka as the key metadata field may be populated with an arbitrary binary value, but this issue is not exclusive to Kafka. If you are blocked by this issue then a work around is to delete either the specific problematic keys: ```yaml pipeline: processors: - mapping: | meta kafka_key = deleted() ``` Or delete all keys with: ```yaml pipeline: processors: - mapping: meta = deleted() ``` ## [](#fields)Fields ### [](#batching)`batching` Configures a batching policy on this output. While the PubSub client maintains its own internal buffering mechanism, preparing larger batches of messages can further trade-off some latency for throughput. **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#byte_threshold)`byte_threshold` Publish a batch when its size in bytes reaches this value. **Type**: `int` **Default**: `1000000` ### [](#count_threshold)`count_threshold` Publish a pubsub buffer when it has this many messages **Type**: `int` **Default**: `100` ### [](#credentials_json)`credentials_json` Base64-encoded Google Service Account credentials in JSON format (optional). Use this field to authenticate with Google Cloud services. For more information about creating service account credentials, see [Google’s service account documentation](https://developers.google.com/workspace/guides/create-credentials#create_credentials_for_a_service_account). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#delay_threshold)`delay_threshold` Publish a non-empty pubsub buffer after this delay has passed. **Type**: `string` **Default**: `10ms` ### [](#endpoint)`endpoint` An optional endpoint to override the default of `pubsub.googleapis.com:443`. This can be used to connect to a region specific pubsub endpoint. For a list of valid values, see [this document](https://cloud.google.com/pubsub/docs/reference/service_apis_overview#list_of_regional_endpoints). **Type**: `string` **Default**: `""` ```yaml # Examples: endpoint: us-central1-pubsub.googleapis.com:443 # --- endpoint: us-west3-pubsub.googleapis.com:443 ``` ### [](#flow_control)`flow_control` For a given topic, configures the PubSub client’s internal buffer for messages to be published. **Type**: `object` ### [](#flow_control-limit_exceeded_behavior)`flow_control.limit_exceeded_behavior` Configures the behavior when trying to publish additional messages while the flow controller is full. The available options are block (default), ignore (disable), and signal\_error (publish results will return an error). **Type**: `string` **Default**: `block` **Options**: `ignore`, `block`, `signal_error` ### [](#flow_control-max_outstanding_bytes)`flow_control.max_outstanding_bytes` Maximum size of buffered messages to be published. If less than or equal to zero, this is disabled. **Type**: `int` **Default**: `-1` ### [](#flow_control-max_outstanding_messages)`flow_control.max_outstanding_messages` Maximum number of buffered messages to be published. If less than or equal to zero, this is disabled. **Type**: `int` **Default**: `1000` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increasing this may improve throughput. **Type**: `int` **Default**: `64` ### [](#metadata)`metadata` Specify criteria for which metadata values are sent as attributes, all are sent by default. **Type**: `object` ### [](#metadata-exclude_prefixes)`metadata.exclude_prefixes[]` Provide a list of explicit metadata key prefixes to be excluded when adding metadata to sent messages. **Type**: `array` **Default**: `[]` ### [](#ordering_key)`ordering_key` The ordering key to use for publishing messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#project)`project` The project ID of the topic to publish to. **Type**: `string` ### [](#publish_timeout)`publish_timeout` The maximum length of time to wait before abandoning a publish attempt for a message. **Type**: `string` **Default**: `1m0s` ```yaml # Examples: publish_timeout: 10s # --- publish_timeout: 5m # --- publish_timeout: 60m ``` ### [](#topic)`topic` The topic to publish to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#validate_topic)`validate_topic` Whether to validate the existence of the topic before publishing. If set to false and the topic does not exist, messages will be lost. **Type**: `bool` **Default**: `true` --- # Page 128: http_client **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/http_client.md --- # http_client > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: http_client latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/http_client page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/http_client.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/http_client.adoc page-git-created-date: "2025-03-04" page-git-modified-date: "2026-05-26" --- Sends messages to a HTTP server. #### Common ```yml outputs: label: "" http_client: url: "" # No default (required) verb: POST headers: {} rate_limit: "" # No default (optional) timeout: 5s max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml # All configuration fields, showing default values output: label: "" http_client: url: "" # No default (required) verb: POST headers: {} metadata: include_prefixes: [] include_patterns: [] dump_request_log_level: "" # Optional oauth: enabled: false consumer_key: "" # Optional consumer_secret: "" # Optional access_token: "" # Optional access_token_secret: "" # Optional oauth2: enabled: false client_key: "" # Optional client_secret: "" # Optional token_url: "" # Optional scopes: [] endpoint_params: {} basic_auth: enabled: false username: "" # Optional password: "" # Optional jwt: enabled: false private_key_file: "" # Optional signing_method: "" # Optional claims: {} headers: {} tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] rate_limit: "" # No default (optional) timeout: 5s retry_period: 1s max_retry_backoff: 300s retries: 3 backoff_on: - 429 drop_on: [] successful_on: [] proxy_url: "" # No default (optional) batch_as_multipart: false max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" # Optional check: "" # Optional processors: [] # No default (optional) multipart: [] ``` ## [](#message-sends)Message sends The body of the request sent to the HTTP server is the raw contents of the message payload. If the message has multiple parts (is a batch), the request is sent according to [RFC1341](https://www.w3.org/Protocols/rfc1341/7_2_Multipart.html). To disable this behavior, set the [`batch_as_multipart`](#batch_as_multipart) field to `false`. When message retries are exhausted, this output rejects a message. Typically, a pipeline then continues attempts to send the message until it succeeds, whilst applying back pressure. ## [](#dynamic-url-and-header-settings)Dynamic URL and header settings You can set the [`url`](#url) and [`headers`](#headers) values dynamically using [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#performance)Performance For improved performance, this output sends: - Multiple messages in parallel. Adjust the `max_in_flight` field value to tune the maximum number of in-flight messages (or message batches). - Messages as batches. You can configure batches at both input and output level. For more information, see [Message Batching](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#backoff_on)`backoff_on[]` A list of status codes that indicate a request failure and trigger retries with an increasing backoff period between attempts. **Type**: `array` **Default**: ```yaml - 429 ``` ### [](#basic_auth)`basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#batch_as_multipart)`batch_as_multipart` When set to `true`, sends all message in a batch as a single request using [RFC1341](https://www.w3.org/Protocols/rfc1341/7_2_Multipart.html). When set to `false`, sends messages in a batch as individual requests. **Type**: `bool` **Default**: `false` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#disable_http2)`disable_http2` Whether to disable HTTP/2. By default, HTTP/2 is enabled. **Type**: `bool` **Default**: `false` ### [](#drop_on)`drop_on[]` A list of status codes that indicate a request failure where the input should not attempt retries. This helps avoid unnecessary retries for requests that are unlikely to succeed. > 📝 **NOTE** > > In these cases, the _request_ is dropped, but the _message_ that triggered the request is retained. **Type**: `array` **Default**: `[]` ### [](#dump_request_log_level)`dump_request_log_level` EXPERIMENTAL: Set the logging level for the request and response payloads of each HTTP request. **Type**: `string` **Default**: `""` **Options**: `TRACE`, `DEBUG`, `INFO`, `WARN`, `ERROR`, `FATAL`, \`\` ### [](#follow_redirects)`follow_redirects` Whether or not to transparently follow redirects, i.e. responses with 300-399 status codes. If disabled, the response message will contain the body, status, and headers from the redirect response and the processor will not make a request to the URL set in the Location header of the response. **Type**: `bool` **Default**: `true` ### [](#headers)`headers` A map of headers to add to the request. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: Content-Type: application/octet-stream traceparent: ${! tracing_span().traceparent } ``` ### [](#jwt)`jwt` (beta) Configure JSON Web Token (JWT) authentication. This feature is in beta and may change in future releases. JWT tokens provide secure, stateless authentication between services. **Type**: `object` ### [](#jwt-claims)`jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` Additional key-value pairs to include in the JWT header (optional). These headers provide extra metadata for JWT processing. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` Path to a file containing the PEM-encoded private key using PKCS#1 or PKCS#8 format. The private key must be compatible with the algorithm specified in the `signing_method` field. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` The cryptographic algorithm used to sign the JWT token. Supported algorithms include RS256, RS384, RS512, and EdDSA. This algorithm must be compatible with the private key specified in the `private_key_file` field. **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#max_retry_backoff)`max_retry_backoff` The maximum period to wait between failed requests. **Type**: `string` **Default**: `300s` ### [](#metadata)`metadata` Specify matching rules that determine which metadata keys to add to the HTTP request as headers (optional). **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#multipart)`multipart[]` EXPERIMENTAL: Create explicit multipart HTTP requests by specifying an array of parts to add to a request. Each part consists of content headers and a data field, which can be populated dynamically. If populated, this field overrides the [default request creation behavior](#message-sends). **Type**: `array` **Default**: `[]` ### [](#multipart-body)`multipart[].body` The body of the individual message part. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: body: ${! this.data.part1 } ``` ### [](#multipart-content_disposition)`multipart[].content_disposition` The content disposition of the individual message part. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: content_disposition: form-data; name="bin"; filename='${! @AttachmentName } ``` ### [](#multipart-content_type)`multipart[].content_type` The content type of the individual message part. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: content_type: application/bin ``` ### [](#oauth)`oauth` Configure OAuth version 1.0 authentication for secure API access. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` The value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` The secret that establishes ownership of the `oauth.access_token` in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` The secret that establishes ownership of the consumer key in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2)`oauth2` Allows you to specify open authentication using OAuth version 2 and the client credentials token flow. **Type**: `object` ### [](#oauth2-client_key)`oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#oauth2-client_secret)`oauth2.client_secret` The secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth2-enabled)`oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2-endpoint_params)`oauth2.endpoint_params` A list of endpoint parameters specified as arrays of strings (optional). **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: bar: - woof foo: - meow - quack ``` ### [](#oauth2-scopes)`oauth2.scopes[]` A list of requested permissions (optional). **Type**: `array` **Default**: `[]` ### [](#oauth2-token_url)`oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#proxy_url)`proxy_url` A HTTP proxy URL (optional). **Type**: `string` ### [](#rate_limit)`rate_limit` A [rate limit](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) to throttle requests by (optional). **Type**: `string` ### [](#retries)`retries` The maximum number of retry attempts to make. **Type**: `int` **Default**: `3` ### [](#retry_period)`retry_period` The initial period to wait between failed requests before retrying. **Type**: `string` **Default**: `1s` ### [](#successful_on)`successful_on[]` A list of HTTP status codes that should be considered as successful, even if they are not 2XX codes. This is useful for handling cases where non-2XX codes indicate that the request was processed successfully, such as `303 See Other` or `409 Conflict`. By default, all 2XX codes are considered successful unless they are specified in `backoff_on` or `drop_on` fields. **Type**: `array` **Default**: `[]` ### [](#timeout)`timeout` A static timeout to apply to requests. **Type**: `string` **Default**: `5s` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL to connect to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#verb)`verb` A verb to connect with. **Type**: `string` **Default**: `POST` ```yaml # Examples: verb: POST # --- verb: GET # --- verb: DELETE ``` --- # Page 129: iceberg **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/iceberg.md --- # iceberg > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: iceberg latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/iceberg page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/iceberg.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/iceberg.adoc description: Fan out Redpanda topics to Apache Iceberg tables using the REST catalog API. page-git-created-date: "2026-03-05" page-git-modified-date: "2026-08-11" --- Fan out Redpanda topics to Apache Iceberg tables using the REST catalog API. This output is well suited for migrating fanout pipelines from Kafka Connect to Redpanda Connect, and supports: - Multiple storage backends (S3, GCS, Azure) - Automatic table creation with schema detection - Partition transforms (year, month, day, hour, bucket, truncate) - Schema evolution (automatic column addition) - Transaction retry logic for concurrent writes ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ### Common ```yml outputs: label: "" iceberg: catalog: url: "" # No default (required) warehouse: "" # No default (optional) auth: oauth2: server_uri: /v1/oauth/tokens client_id: "" # No default (required) client_secret: "" # No default (required) scope: "" # No default (optional) bearer: "" # No default (optional) aws_sigv4: region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) service: "" # No default (optional) headers: "" # No default (optional) tls_skip_verify: false namespace: "" # No default (required) table: "" # No default (required) storage: aws_s3: bucket: "" # No default (required) region: "" # No default (optional) endpoint: "" # No default (optional) force_path_style_urls: false credentials: id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) gcp_cloud_storage: bucket: "" # No default (required) endpoint: "" # No default (optional) credentials_type: "" # No default (optional) credentials_file: "" # No default (optional) credentials_json: "" # No default (optional) azure_blob_storage: storage_account: "" # No default (required) container: "" # No default (required) endpoint: "" # No default (optional) storage_sas_token: "" # No default (optional) storage_connection_string: "" # No default (optional) storage_access_key: "" # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 4 ``` ### Advanced ```yml outputs: label: "" iceberg: catalog: url: "" # No default (required) warehouse: "" # No default (optional) auth: oauth2: server_uri: /v1/oauth/tokens client_id: "" # No default (required) client_secret: "" # No default (required) scope: "" # No default (optional) bearer: "" # No default (optional) aws_sigv4: region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) service: "" # No default (optional) headers: "" # No default (optional) tls_skip_verify: false namespace: "" # No default (required) table: "" # No default (required) case_sensitive_columns: true row_operation: insert identifier_fields: [] merge_strategy: merge-on-read storage: aws_s3: bucket: "" # No default (required) region: "" # No default (optional) endpoint: "" # No default (optional) force_path_style_urls: false credentials: id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) gcp_cloud_storage: bucket: "" # No default (required) endpoint: "" # No default (optional) credentials_type: "" # No default (optional) credentials_file: "" # No default (optional) credentials_json: "" # No default (optional) azure_blob_storage: storage_account: "" # No default (required) container: "" # No default (required) endpoint: "" # No default (optional) storage_sas_token: "" # No default (optional) storage_connection_string: "" # No default (optional) storage_access_key: "" # No default (optional) schema_evolution: enabled: false partition_spec: () table_location: "" # No default (optional) schema_metadata: "" new_column_type_mapping: "" # No default (optional) require_schema_metadata: false commit: manifest_merge_enabled: true max_snapshot_age: 24h max_retries: 3 cleanup_on_failure: true parquet: string_encoding: delta_length_byte_array batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 4 ``` ## [](#catalog-integration)Catalog integration This output works with REST catalog implementations including Apache Polaris, AWS Glue Data Catalog, and Databricks Unity Catalog. ### [](#apache-polaris)Apache Polaris To use with [Apache Polaris](https://polaris.apache.org): - Set `catalog.url` to the Polaris REST endpoint (e.g., `[http://localhost:8181/api/catalog](http://localhost:8181/api/catalog)`). - Set `catalog.warehouse` to the catalog name configured in Polaris. - Configure `catalog.auth.oauth2` with client credentials granted access to the catalog. ### [](#aws-glue-data-catalog)AWS Glue Data Catalog To use with AWS Glue Data Catalog: - Set `catalog.url` to `[https://glue..amazonaws.com/iceberg](https://glue.\.amazonaws.com/iceberg)` (the REST client appends the API version automatically). - Set `catalog.warehouse` to your AWS account ID (the Glue catalog identifier). - Set `schema_evolution.table_location` to an S3 prefix (e.g., `s3://my-bucket/`) since Glue does not automatically assign table locations. - Configure `catalog.auth.aws_sigv4` with the appropriate region and set `service` to `glue`. - Configure `storage.aws_s3` with the same bucket and region. ### [](#azure-blob-storage-adls-gen2)Azure Blob Storage (ADLS Gen2) To use with Azure Data Lake Storage Gen2: - Configure `storage.azure_blob_storage` with your storage account name and container. - Authenticate using one of: `storage_access_key` (shared key), `storage_sas_token`, or `storage_connection_string`. - The storage account must have hierarchical namespace (HNS) enabled for ADLS Gen2 compatibility. ## [](#type-mapping)Type mapping | Bloblang type | Iceberg type | | --- | --- | | string | string | | bytes | binary | | bool | boolean | | number | double | | timestamp | timestamp (with timezone) | | object | struct | | array | list | ## [](#row-level-operations)Row-level operations By default, this output is append-only: every message becomes a new row (`row_operation: insert`), and existing configurations are unaffected. Set `row_operation` to apply a per-message operation, with `identifier_fields` defining the row identity: - `insert`: Append the row. This is an unconditional append, and is not keyed or deduplicated. - `upsert`: Replace any existing rows matching `identifier_fields`, then append this row. This is equivalent to Iceberg’s Flink UPSERT mode. - `delete`: Remove rows matching `identifier_fields`. `row_operation` supports interpolation, so the operation can be driven by the data itself, for example by mapping a change-data-capture stream’s operation field. No CDC-specific format is assumed. The field is named `row_operation` to distinguish it from Iceberg’s snapshot-level operation. `upsert` and `delete` require `identifier_fields` and use Iceberg merge-on-read equality deletes, which require table format version 2. A version-1 table is automatically upgraded to version 2 on the first `upsert` or `delete`. This upgrade is irreversible. ### [](#identifier-fields)Identifier fields `identifier_fields` must reference existing table columns of a primitive, non-floating-point type. A static `upsert` or `delete` is validated at startup. An interpolated `row_operation` is validated for each message at write time, so an empty `identifier_fields` is not caught until the first `upsert` or `delete` message arrives. Identifier columns of a temporal type (`timestamp`, `timestamptz`, `date`, `time`) must arrive as time values, not bare numbers. A numeric epoch is ambiguous as a delete key and is rejected at write time, so convert it to a timestamp upstream. If the table is partitioned, every partition source column must be one of the `identifier_fields`, because equality deletes are partition-scoped. When this output auto-creates a table through `schema_evolution`, the `identifier_fields` columns are created as required and registered as the table’s Iceberg identifier-field-ids, so downstream engines and other writers see the primary key. As a result, a null or missing value in an identifier column is rejected on write, even for `insert`. Identifier columns must therefore be present at creation, either in the first message or declared through `schema_metadata`. Pre-existing tables are never modified. ### [](#batching-and-ordering)Batching and ordering Within a single batch, the last `upsert` or `delete` for each `identifier_fields` key wins. Each batch containing an `upsert` or `delete` is committed as its own snapshot. These commits are never coalesced, which is required for correctness, so a high-throughput mutation workload produces one snapshot per batch. Size batches accordingly, and run regular table maintenance (snapshot expiry and compaction) to keep metadata manageable. Pure `insert`\-only batches keep the original append fast path, which does coalesce commits. Ordering only holds within a batch. With more than one batch in flight, concurrent batches can commit out of order, so a stale `upsert` might overwrite a newer one for the same key. Set `max_in_flight: 1` for keyed (change-data-capture) workloads to preserve per-key order. This is enforced by config linting whenever `row_operation` is anything other than a static `insert`. > ⚠️ **CAUTION** > > `insert` is an unconditional append, and is not keyed or deduplicated. For keyed data, including change-data-capture, map create and read events to `upsert`, never `insert`. Mixing `insert` with `upsert` or `delete` on the same key in one batch produces duplicate rows. ### [](#merge-strategies)Merge strategies `merge_strategy` controls how `upsert` and `delete` mutations are written to disk, which determines both write cost and which query engines can read the table. - `merge-on-read` (the default) writes Iceberg v2 equality-delete files and applies them at read time. Writes stay cheap and streaming-friendly, but only catalog-native or Flink-world engines, such as Apache Polaris, Flink, Trino, and Spark, can read equality deletes. Engine-backed catalogs such as Snowflake and the Databricks Unity Catalog cannot. It requires table format version 2, so a version-1 table is automatically upgraded to version 2 on the first `upsert` or `delete`. This upgrade is irreversible. - `copy-on-write` rewrites whole data files instead, so the table only ever holds plain data files, with no delete files. Every engine that reads Iceberg data files can read the result, including Snowflake, the Databricks Unity Catalog, Polaris, Flink, and Trino. Because it only ever writes plain data files, it works on version-1 or version-2 tables without forcing the version upgrade. Choose `merge-on-read` for streaming or high-throughput mutation into a lake read by Polaris, Flink, Trino, or Spark, where write cost must stay low. Choose `copy-on-write` when the table must be readable by Snowflake, the Databricks Unity Catalog, or another engine that cannot read equality deletes, and the workload is batch or moderate-throughput so the write amplification is acceptable. `copy-on-write` rewrites every data file that contains a touched key, so sort the table by the identifier key and use large batches to keep the number of rewritten files as low as possible. It also materializes the whole new-row batch in memory while the batch commits, so size keyed batches to stay within your process memory budget. If a commit fails, `commit.cleanup_on_failure` (enabled by default) removes the files that commit had already written, to limit orphaned objects in storage. Disable it only if you’d rather leave those files for Iceberg’s own orphan-file maintenance, such as snapshot expiry and `remove_orphan_files`, to reclaim. ### [](#change-data-capture-example)Change-data-capture example Materialize a change-data-capture stream into an Iceberg table. The mapping derives the row operation from the source’s operation field (here Debezium’s `op`, where `c`, `r`, and `u` map to `upsert`, and `d` maps to `delete`) and selects the row image, while `identifier_fields` is the primary key: ```yaml input: redpanda: seed_brokers: [ localhost:9092 ] topics: [ dbserver.inventory.customers ] consumer_group: iceberg_sink pipeline: processors: - mapping: | meta op = match this.op { "d" => "delete", _ => "upsert", } # Debezium puts the row image in 'after', or 'before' for deletes. root = this.after | this.before output: iceberg: catalog: url: http://localhost:8181/api/catalog namespace: inventory table: customers row_operation: ${! metadata("op") } identifier_fields: [ id ] # Keyed writes must stay ordered. A single batch in flight prevents # concurrent batches from committing a stale update over a newer one. max_in_flight: 1 storage: aws_s3: bucket: my-iceberg-data region: us-east-1 ``` ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#case_sensitive_columns)`case_sensitive_columns` Controls how message field names are matched against table column names, and how column references in the partition spec are resolved. When `true` (the default), names must match exactly. When `false`, matching is case-insensitive — set this when your downstream catalog or query engine treats column names as case-insensitive (the iceberg specification’s recommended convention) so that, for example, a message keyed `"COLUMN"` lands in an existing `column` rather than triggering schema evolution. Ambiguous case-only duplicates in the input are rejected. **Type**: `bool` **Default**: `true` ### [](#catalog)`catalog` REST catalog configuration. **Type**: `object` ### [](#catalog-auth)`catalog.auth` Authentication configuration for the REST catalog. Only one authentication method can be active at a time. **Type**: `object` ### [](#catalog-auth-aws_sigv4)`catalog.auth.aws_sigv4` AWS SigV4 authentication (for AWS Glue Data Catalog or API Gateway). **Type**: `object` ### [](#catalog-auth-aws_sigv4-credentials)`catalog.auth.aws_sigv4.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#catalog-auth-aws_sigv4-credentials-from_ec2_role)`catalog.auth.aws_sigv4.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#catalog-auth-aws_sigv4-credentials-id)`catalog.auth.aws_sigv4.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#catalog-auth-aws_sigv4-credentials-profile)`catalog.auth.aws_sigv4.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#catalog-auth-aws_sigv4-credentials-role)`catalog.auth.aws_sigv4.credentials.role` A role ARN to assume. **Type**: `string` ### [](#catalog-auth-aws_sigv4-credentials-role_external_id)`catalog.auth.aws_sigv4.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#catalog-auth-aws_sigv4-credentials-secret)`catalog.auth.aws_sigv4.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#catalog-auth-aws_sigv4-credentials-token)`catalog.auth.aws_sigv4.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#catalog-auth-aws_sigv4-endpoint)`catalog.auth.aws_sigv4.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#catalog-auth-aws_sigv4-region)`catalog.auth.aws_sigv4.region` The AWS region to target. **Type**: `string` ### [](#catalog-auth-aws_sigv4-service)`catalog.auth.aws_sigv4.service` AWS service name for SigV4 signing. **Type**: `string` ### [](#catalog-auth-aws_sigv4-tcp)`catalog.auth.aws_sigv4.tcp` TCP socket configuration. **Type**: `object` ### [](#catalog-auth-aws_sigv4-tcp-connect_timeout)`catalog.auth.aws_sigv4.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#catalog-auth-aws_sigv4-tcp-keep_alive)`catalog.auth.aws_sigv4.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#catalog-auth-aws_sigv4-tcp-keep_alive-count)`catalog.auth.aws_sigv4.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#catalog-auth-aws_sigv4-tcp-keep_alive-idle)`catalog.auth.aws_sigv4.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#catalog-auth-aws_sigv4-tcp-keep_alive-interval)`catalog.auth.aws_sigv4.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#catalog-auth-aws_sigv4-tcp-tcp_user_timeout)`catalog.auth.aws_sigv4.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#catalog-auth-bearer)`catalog.auth.bearer` Static bearer token for authentication. For testing only, not recommended for production. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#catalog-auth-oauth2)`catalog.auth.oauth2` OAuth2 authentication configuration. **Type**: `object` ### [](#catalog-auth-oauth2-client_id)`catalog.auth.oauth2.client_id` OAuth2 client identifier. **Type**: `string` ### [](#catalog-auth-oauth2-client_secret)`catalog.auth.oauth2.client_secret` OAuth2 client secret. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#catalog-auth-oauth2-scope)`catalog.auth.oauth2.scope` OAuth2 scope to request. **Type**: `string` ### [](#catalog-auth-oauth2-server_uri)`catalog.auth.oauth2.server_uri` OAuth2 token endpoint URI. **Type**: `string` **Default**: `/v1/oauth/tokens` ### [](#catalog-headers)`catalog.headers` Custom HTTP headers to include in all requests to the catalog. **Type**: `object` ```yaml # Examples: headers: X-Api-Key: your-api-key ``` ### [](#catalog-tls_skip_verify)`catalog.tls_skip_verify` Skip TLS certificate verification. Not recommended for production. **Type**: `bool` **Default**: `false` ### [](#catalog-url)`catalog.url` The REST catalog endpoint URL. **Type**: `string` ```yaml # Examples: url: http://localhost:8181/api/catalog # --- url: https://polaris.example.com/api/catalog # --- url: https://glue.us-east-1.amazonaws.com/iceberg ``` ### [](#catalog-warehouse)`catalog.warehouse` The REST catalog warehouse. **Type**: `string` ```yaml # Examples: warehouse: redpanda-catalog ``` ### [](#commit)`commit` Commit behavior configuration. **Type**: `object` ### [](#commit-cleanup_on_failure)`commit.cleanup_on_failure` Whether to remove the files a failed commit had already written. A commit’s data files (and, under `merge-on-read`, its equality-delete files) are written to storage before the catalog commit, so a failed commit leaves them referenced by no snapshot; the default `true` removes them best-effort to limit orphaned objects. The same sweep reclaims the superseded files of earlier attempts when a `copy-on-write` commit succeeds on retry (each attempt re-stages fresh rewrites). This applies to every write path: append, `merge-on-read`, and `copy-on-write`. Cleanup is already skipped automatically whenever a commit’s outcome is ambiguous, since such a commit may still land server-side, and deleting files a landed snapshot references would corrupt the table. Those leftovers are always deferred to table maintenance regardless of this setting. Setting this to `false` disables cleanup entirely, as a safety valve: writes then never delete anything, and every failed commit leaves its files orphaned in storage for Iceberg orphan-file maintenance (snapshot expiry plus `remove_orphan_files`) to reclaim. See the [Merge strategies](#merge-strategies) section above for that maintenance guidance. **Type**: `bool` **Default**: `true` ### [](#commit-manifest_merge_enabled)`commit.manifest_merge_enabled` Merge small manifest files during commits to reduce metadata overhead. **Type**: `bool` **Default**: `true` ### [](#commit-max_retries)`commit.max_retries` Maximum number of times to retry a failed transaction commit. **Type**: `int` **Default**: `3` ### [](#commit-max_snapshot_age)`commit.max_snapshot_age` Maximum age of snapshots to retain for time-travel queries. Set to zero to disable removing old snapshots. **Type**: `string` **Default**: `24h` ### [](#identifier_fields)`identifier_fields[]` The columns forming the row identity (the Iceberg identifier fields, or equality-delete key) used by `upsert` and `delete`. Required when `row_operation` can evaluate to `upsert` or `delete`, and must reference existing table columns of a primitive, non-floating-point type. See the [Row-level operations](#row-level-operations) section above for the full constraints, including the temporal-type and partitioning rules and when the requirement is enforced. **Type**: `array` **Default**: `[]` ```yaml # Examples: identifier_fields: - id # --- identifier_fields: - tenant_id - user_id ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `4` ### [](#merge_strategy)`merge_strategy` How `upsert` and `delete` are materialised on disk. - `merge-on-read` (the default) writes Iceberg v2 equality-delete files. Deletes are applied at read time, so writes stay cheap and streaming-friendly, but only catalog-native / Flink-world engines can read the result. Engine-backed catalogs such as Snowflake and the Databricks Unity Catalog cannot read equality deletes. - `copy-on-write` rewrites whole data files so the table only ever contains plain data files (no delete files), which every engine can read, including Snowflake and Databricks Unity Catalog. It works on version-1 or version-2 tables and never forces the irreversible v1→v2 upgrade. The trade-off is heavy write amplification: each mutating batch rewrites every data file that contains a touched key, so it is a batch / moderate-throughput mode. Sort the table by the identifier key and use large batches so each rewrite touches as few files as possible. See the [Merge strategies](#merge-strategies) section above for the full decision guide, copy-on-write support matrix (column and merge-key types, partitioning, table format), and maintenance guidance. **Type**: `string` **Default**: `merge-on-read` **Options**: `merge-on-read`, `copy-on-write` ### [](#namespace)`namespace` The Iceberg namespace for the table, dot delimiters are split as nested namespaces. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: namespace: analytics.events # --- namespace: production ``` ### [](#parquet)`parquet` Parquet writer configuration. **Type**: `object` ### [](#parquet-string_encoding)`parquet.string_encoding` The encoding to use for string and binary columns. Use `plain` for compatibility with readers that do not support `DELTA_LENGTH_BYTE_ARRAY` encoding, such as AWS Redshift Spectrum. **Type**: `string` **Default**: `delta_length_byte_array` **Options**: `plain`, `delta_length_byte_array` ### [](#row_operation)`row_operation` The row-level operation to apply for each message: `insert` (append), `upsert` (replace rows matching `identifier_fields`, then append), or `delete` (remove rows matching `identifier_fields`). Supports interpolation so the operation can be driven by the data, such as a change-data-capture stream’s operation field. Defaults to `insert`, preserving the original append-only behavior. See the [Row-level operations](#row-level-operations) section above for the full semantics, the format-version-2 upgrade, batching behavior, and important caveats. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `insert` ```yaml # Examples: row_operation: insert # --- row_operation: ${! metadata("op") } # --- row_operation: ${! this.op == "d" ? "delete" : "upsert" } ``` ### [](#schema_evolution)`schema_evolution` Schema evolution configuration. **Type**: `object` ### [](#schema_evolution-enabled)`schema_evolution.enabled` Enable automatic schema evolution. When enabled, new columns will be automatically added to the table. **Type**: `bool` **Default**: `false` ### [](#schema_evolution-new_column_type_mapping)`schema_evolution.new_column_type_mapping` An optional Bloblang mapping to customize column types during schema evolution. This mapping is executed for each new column and can override the inferred or schema-metadata-derived type. The mapping receives an object with fields `name` (column name), `path` (dot-separated path), `value` (sample value), `inferred_type` (the type that would be used without this mapping), `message` (the full message body), `namespace`, and `table`. It must return a string with a valid Iceberg type name: `boolean`, `int`, `long`, `float`, `double`, `string`, `binary`, `date`, `time`, `timestamp`, `timestamptz`, `uuid`, `decimal(p,s)`, or `fixed[n]`. **Type**: `string` ### [](#schema_evolution-partition_spec)`schema_evolution.partition_spec` A Bloblang expression to evaluate when a new table is created to determine the table’s partition spec. The result of the mapping should be an Iceberg partition spec in the same string format as the Redpanda Streaming Topic Property (see Redpanda Core documentation for Iceberg topics). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `()` ```yaml # Examples: partition_spec: (col1) # --- partition_spec: (nested.col) # --- partition_spec: (year(my_ts_col)) # --- partition_spec: (year(my_ts_col), col2) # --- partition_spec: (hour(my_ts_col), truncate(42, col2)) # --- partition_spec: (day(my_ts_col), bucket(4, nested.col)) # --- partition_spec: (day(my_ts_col), void(`non.nested column.with.dots`), identity(nested.column)) ``` ### [](#schema_evolution-require_schema_metadata)`schema_evolution.require_schema_metadata` When `true`, writing a numeric value into a `timestamp`, `timestamptz`, `date`, or `time` column without `schema_metadata` registered for that column is a hard error. The default `false` permits a fallback path that interprets bare numeric timestamps as Unix seconds and bare numeric times as already-microseconds — convenient, but silently wrong if upstream produced milliseconds. Enable this when you cannot guarantee the upstream attaches schema metadata and want to fail loudly rather than corrupt dates by ~50,000 years. No effect on time-typed columns receiving `time.Time`/`time.Duration` Go values, which carry their own unit unambiguously, and no effect on non-time columns. Requires `schema_metadata` to be set. **Type**: `bool` **Default**: `false` ### [](#schema_evolution-schema_metadata)`schema_evolution.schema_metadata` The name of a message metadata field containing a schema definition. When set, the schema is used to determine column types during schema evolution and table creation instead of inferring types from values. The schema must be in the standard common schema format (the same format used by the `parquet_encode` processor’s `schema_metadata` field). For batches of messages, the first message’s schema is used. Record presence drives schema shape: fields declared in the schema metadata that are absent from the record are not added to the table, while the metadata controls column ordering, naming, and types for fields that are present. In case-insensitive mode, top-level column names use the metadata’s casing — record keys are matched by case-folding and the metadata’s name is what lands in the table. **Type**: `string` **Default**: `""` ### [](#schema_evolution-table_location)`schema_evolution.table_location` A prefix used as the location for new tables when the catalog does not automatically assign one. For example, AWS Glue requires explicit table locations. When set, table locations are derived as `{prefix}{namespace}/{table}`. **Type**: `string` ```yaml # Examples: table_location: s3://my-iceberg-bucket/ ``` ### [](#storage)`storage` Storage backend configuration for data files. Exactly one of `aws_s3`, `gcp_cloud_storage`, or `azure_blob_storage` must be specified. **Type**: `object` ### [](#storage-aws_s3)`storage.aws_s3` S3 storage configuration. **Type**: `object` ### [](#storage-aws_s3-bucket)`storage.aws_s3.bucket` The S3 bucket name. **Type**: `string` ```yaml # Examples: bucket: my-iceberg-data ``` ### [](#storage-aws_s3-credentials)`storage.aws_s3.credentials` Static AWS credentials for S3 access. When not specified, credentials are loaded from the default AWS credential chain. **Type**: `object` ### [](#storage-aws_s3-credentials-id)`storage.aws_s3.credentials.id` The AWS access key ID. **Type**: `string` ### [](#storage-aws_s3-credentials-secret)`storage.aws_s3.credentials.secret` The AWS secret access key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#storage-aws_s3-credentials-token)`storage.aws_s3.credentials.token` The AWS session token, required when using short term credentials. **Type**: `string` ### [](#storage-aws_s3-endpoint)`storage.aws_s3.endpoint` Custom endpoint for S3-compatible storage (e.g., MinIO). **Type**: `string` ```yaml # Examples: endpoint: http://localhost:9000 ``` ### [](#storage-aws_s3-force_path_style_urls)`storage.aws_s3.force_path_style_urls` Forces the client API to use path style URLs, which is often required when connecting to custom endpoints. **Type**: `bool` **Default**: `false` ### [](#storage-aws_s3-region)`storage.aws_s3.region` The AWS region. **Type**: `string` ```yaml # Examples: region: us-west-2 ``` ### [](#storage-azure_blob_storage)`storage.azure_blob_storage` Azure Blob Storage (ADLS Gen2) configuration. **Type**: `object` ### [](#storage-azure_blob_storage-container)`storage.azure_blob_storage.container` The Azure blob container name. **Type**: `string` ```yaml # Examples: container: iceberg-data ``` ### [](#storage-azure_blob_storage-endpoint)`storage.azure_blob_storage.endpoint` Custom endpoint for Azure-compatible storage. **Type**: `string` ### [](#storage-azure_blob_storage-storage_access_key)`storage.azure_blob_storage.storage_access_key` Azure storage access key for shared key authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#storage-azure_blob_storage-storage_account)`storage.azure_blob_storage.storage_account` The Azure storage account name. **Type**: `string` ```yaml # Examples: storage_account: mystorageaccount ``` ### [](#storage-azure_blob_storage-storage_connection_string)`storage.azure_blob_storage.storage_connection_string` Azure storage connection string. Use this or other auth methods, not both. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#storage-azure_blob_storage-storage_sas_token)`storage.azure_blob_storage.storage_sas_token` SAS token for authentication. Prefix with the container name followed by a dot if container-specific. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#storage-gcp_cloud_storage)`storage.gcp_cloud_storage` Google Cloud Storage configuration. **Type**: `object` ### [](#storage-gcp_cloud_storage-bucket)`storage.gcp_cloud_storage.bucket` The GCS bucket name. **Type**: `string` ```yaml # Examples: bucket: my-iceberg-data ``` ### [](#storage-gcp_cloud_storage-credentials_file)`storage.gcp_cloud_storage.credentials_file` Path to a GCP credentials JSON file. **Type**: `string` ### [](#storage-gcp_cloud_storage-credentials_json)`storage.gcp_cloud_storage.credentials_json` GCP credentials JSON content. Use this or `credentials_file`, not both. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#storage-gcp_cloud_storage-credentials_type)`storage.gcp_cloud_storage.credentials_type` The type of credentials to use. Valid values: `service_account`, `authorized_user`, `impersonated_service_account`, `external_account`. **Type**: `string` ```yaml # Examples: credentials_type: service_account ``` ### [](#storage-gcp_cloud_storage-endpoint)`storage.gcp_cloud_storage.endpoint` Custom endpoint for GCS-compatible storage. **Type**: `string` ### [](#table)`table` The destination Iceberg table name. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: table: user_events # --- table: events_${!meta("topic")} ``` --- # Page 130: inproc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/inproc.md --- # inproc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: inproc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/inproc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/inproc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/inproc.adoc description: Sends data directly to Redpanda Connect inputs in the same process by connecting to a unique ID, to link isolated streams when running in streams mode. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- ```yml outputs: label: "" inproc: "" ``` Sends data directly to Redpanda Connect inputs by connecting to a unique ID. It is possible to connect multiple inputs to the same inproc ID, resulting in messages dispatching in a round-robin fashion to connected inputs. However, only one output can assume an inproc ID, and will replace existing outputs if a collision occurs. --- # Page 131: kafka_franz **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/kafka_franz.md --- # kafka_franz > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: kafka_franz latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/kafka_franz page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/kafka_franz.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/kafka_franz.adoc description: A Kafka output using the Franz Kafka client library. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > ⚠️ **WARNING: Deprecated in 4.68.0** > > Deprecated in 4.68.0 > > This component is deprecated and will be removed in the next major version release. Please consider moving onto the unified [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [`redpanda` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) components. The `kafka_franz` output writes a batch of messages to Kafka brokers and waits for acknowledgement before propagating any acknowledgments back to the input. This output often outperforms the traditional `kafka` output, as well as providing more useful logs and error messages. This output uses the [Franz Kafka client library](https://github.com/twmb/franz-go). #### Common ```yml outputs: label: "" kafka_franz: seed_brokers: [] # No default (required) topic: "" # No default (required) key: "" # No default (optional) partition: "" # No default (optional) metadata: include_prefixes: [] include_patterns: [] max_in_flight: 10 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" kafka_franz: seed_brokers: [] # No default (required) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s topic: "" # No default (required) key: "" # No default (optional) partition: "" # No default (optional) metadata: include_prefixes: [] include_patterns: [] timestamp_ms: "" # No default (optional) max_in_flight: 10 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) partitioner: "" # No default (optional) idempotent_write: true acks: all compression: "" # No default (optional) allow_auto_topic_creation: true timeout: 10s max_message_bytes: 1MiB broker_write_max_bytes: 100MiB max_buffered_records: 10000 max_buffered_bytes: 0 max_in_flight_requests: 1 record_retries: 0 record_delivery_timeout: 0s ``` ## [](#fields)Fields ### [](#acks)`acks` The number of acknowledgements the leader broker must receive from ISR brokers before responding to the produce request. When `idempotent_write` is enabled this must be set to `all`. **Type**: `string` **Default**: `all` | Option | Summary | | --- | --- | | all | Wait for all in-sync replicas to acknowledge (acks=-1). Required when idempotent_write is enabled. | | leader | Wait for the leader broker to acknowledge (acks=1). Messages are lost if the leader fails before replication. | | none | Do not wait for any acknowledgement (acks=0). Highest throughput but messages may be lost. | ### [](#allow_auto_topic_creation)`allow_auto_topic_creation` Enables topics to be auto created if they do not exist when fetching their metadata. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#broker_write_max_bytes)`broker_write_max_bytes` The maximum number of bytes this output can write to a broker connection in a single write. This field corresponds to Kafka’s `socket.request.max.bytes`. **Type**: `string` **Default**: `100MiB` ```yaml # Examples: broker_write_max_bytes: 128MB # --- broker_write_max_bytes: 50mib ``` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#compression)`compression` Set an explicit compression type (optional). The default preference is to use `snappy` when the broker supports it. Otherwise, use `none`. **Type**: `string` **Options**: `lz4`, `snappy`, `gzip`, `none`, `zstd` ### [](#conn_idle_timeout)`conn_idle_timeout` The maximum duration that connections can remain idle before they are automatically closed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `20s` ### [](#idempotent_write)`idempotent_write` Enables the idempotent write producer option. This requires the `IDEMPOTENT_WRITE` permission on `CLUSTER`. Disable this option if the `IDEMPOTENT_WRITE` permission is unavailable. **Type**: `bool` **Default**: `true` ### [](#key)`key` An optional key to populate for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_buffered_bytes)`max_buffered_bytes` The maximum number of bytes the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered. Set to `0` to disable the byte-level limit (only `max_buffered_records` applies). This limit is checked after `max_buffered_records`. **Type**: `string` **Default**: `0` ```yaml # Examples: max_buffered_bytes: 256MB # --- max_buffered_bytes: 50mib ``` ### [](#max_buffered_records)`max_buffered_records` The maximum number of records the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered and space frees up. Increase this value for high-throughput pipelines to avoid back-pressure stalls. **Type**: `int` **Default**: `10000` ### [](#max_in_flight)`max_in_flight` The maximum number of batches to send in parallel at any given time. **Type**: `int` **Default**: `10` ### [](#max_in_flight_requests)`max_in_flight_requests` The maximum number of produce requests in flight per broker connection. While `idempotent_write` is enabled (the default) this must be `1`, as the client relies on a single in-flight request per broker to guarantee ordering. To use higher values you must set `idempotent_write` to `false`, which allows requests to be pipelined for throughput but may cause duplicate and out-of-order delivery on retries. Note that this is distinct from the output’s `max_in_flight` field, which counts message batches being written in parallel rather than produce requests on the wire. A value of `1` is not a throughput ceiling: records from concurrent writes are coalesced into fewer, larger produce requests. **Type**: `int` **Default**: `1` ### [](#max_message_bytes)`max_message_bytes` The maximum space (in bytes) that an individual message may use. Messages larger than this value are rejected. This field corresponds to Kafka’s `max.message.bytes`. **Type**: `string` **Default**: `1MiB` ```yaml # Examples: max_message_bytes: 100MB # --- max_message_bytes: 50mib ``` ### [](#metadata)`metadata` Configure which metadata values are added to messages as headers. This allows you to pass additional context information along with your messages. **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#metadata_max_age)`metadata_max_age` The maximum period of time after which metadata is refreshed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. Lower values provide more responsive topic and partition discovery but may increase broker load. Higher values reduce broker queries but can delay detection of topology changes. **Type**: `string` **Default**: `1m` ### [](#partition)`partition` Set a partition for each message (optional). This field is only relevant when the `partitioner` is set to `manual`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). You must provide an interpolation string that is a valid integer. **Type**: `string` ```yaml # Examples: partition: ${! meta("partition") } ``` ### [](#partitioner)`partitioner` Override the default murmur2 hashing partitioner. **Type**: `string` | Option | Summary | | --- | --- | | least_backup | Chooses the least backed up partition (the partition with the fewest amount of buffered records). Partitions are selected per batch. | | manual | Manually select a partition for each message, requires the field partition to be specified. | | murmur2_hash | Kafka’s default hash algorithm that uses a 32-bit murmur2 hash of the key to compute which partition the record will be on. | | round_robin | Round-robin’s messages through all available partitions. This algorithm has lower throughput and causes higher CPU load on brokers, but can be useful if you want to ensure an even distribution of records to partitions. | ### [](#record_delivery_timeout)`record_delivery_timeout` The maximum time a record can sit in the producer buffer before it is failed, roughly equivalent to Kafka’s `delivery.timeout.ms`. This is evaluated before writing a request or after a produce response. When a record times out, all records in the same partition are also failed. Set to `0s` for no timeout (the default). With `idempotent_write` enabled, timeouts are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `string` **Default**: `0s` ### [](#record_retries)`record_retries` The maximum number of times a record produce is retried on failure before the record is failed. When a record fails, all records buffered in the same partition are also failed to preserve gapless ordering. Set to `0` for unlimited retries (the default). With `idempotent_write` enabled, retries are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `int` **Default**: `0` ### [](#request_timeout_overhead)`request_timeout_overhead` Grants an additional buffer or overhead to requests that have timeout fields defined. This field is based on the behavior of Apache Kafka’s `request.timeout.ms` parameter, but with the option to extend the timeout deadline. **Type**: `string` **Default**: `10s` ### [](#sasl)`sasl[]` Specify one or more methods or mechanisms of SASL authentication, which are attempted in order. If the broker supports the first SASL mechanism, all connections use it. If the first mechanism fails, the client picks the first supported mechanism. If the broker does not support any client mechanisms, all connections fail. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM based authentication as specified in RFC5802. | | none | Disable sasl authentication | ### [](#sasl-password)`sasl[].password` A password to provide for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` A username to provide for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to in order. Use commas to separate multiple addresses in a single list item. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` The maximum period of time to wait for message sends before abandoning the request and retrying. **Type**: `string` **Default**: `10s` ### [](#timestamp_ms)`timestamp_ms` Set a timestamp (in milliseconds) for each message (optional). When left empty, the current timestamp is used. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: timestamp_ms: ${! timestamp_unix_milli() } # --- timestamp_ms: ${! metadata("kafka_timestamp_ms") } ``` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic)`topic` A topic to write messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 132: kafka **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/kafka.md --- # kafka > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: kafka latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/kafka page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/kafka.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/kafka.adoc description: The kafka output type writes a batch of messages to Kafka brokers and waits for acknowledgement before propagating it back to the input. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > ⚠️ **WARNING: Deprecated in 4.68.0** > > Deprecated in 4.68.0 > > This component is deprecated and will be removed in the next major version release. Please consider moving onto the unified [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [`redpanda` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) components. The `kafka` output writes a batch of messages to Kafka brokers and waits for acknowledgement before propagating any acknowledgements back to the input. #### Common ```yml outputs: label: "" kafka: addresses: [] # No default (required) topic: "" # No default (required) target_version: "" # No default (optional) key: "" partitioner: fnv1a_hash compression: none static_headers: "" # No default (optional) metadata: exclude_prefixes: [] max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" kafka: addresses: [] # No default (required) tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: mechanism: none user: "" password: "" access_token: "" token_cache: "" token_key: "" topic: "" # No default (required) client_id: benthos target_version: "" # No default (optional) rack_id: "" key: "" partitioner: fnv1a_hash partition: "" custom_topic_creation: enabled: false partitions: -1 replication_factor: -1 compression: none static_headers: "" # No default (optional) metadata: exclude_prefixes: [] inject_tracing_map: "" # No default (optional) max_in_flight: 64 idempotent_write: false ack_replicas: false max_msg_bytes: 1000000 timeout: 5s retry_as_batch: false batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_retries: 0 backoff: initial_interval: 3s max_interval: 10s max_elapsed_time: 30s timestamp_ms: "" # No default (optional) ``` The configuration field `ack_replicas` determines whether Redpanda Connect waits for acknowledgement from all replicas or just a single broker. Both the `key` and `topic` fields can be dynamically set using function interpolations described in [Bloblang queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). [Metadata](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/metadata/) will be added to each message sent as headers (version 0.11+), but can be restricted using the field [`metadata`](#metadata). ## [](#strict-ordering-and-retries)Strict ordering and retries When strict ordering is required for messages written to topic partitions it is important to ensure that both the field `max_in_flight` is set to `1` and that the field `retry_as_batch` is set to `true`. You must also ensure that failed batches are never rerouted back to the same output. This can be done by setting the field `max_retries` to `0` and `backoff.max_elapsed_time` to empty, which will apply back pressure indefinitely until the batch is sent successfully. However, this also means that manual intervention will eventually be required in cases where the batch cannot be sent due to configuration problems such as an incorrect `max_msg_bytes` estimate. A less strict but automated alternative would be to route failed batches to a dead letter queue using a [`fallback` broker](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/), but this would allow subsequent batches to be delivered in the meantime whilst those failed batches are dealt with. ## [](#troubleshooting)Troubleshooting If you’re seeing issues writing to or reading from Kafka with this component then it’s worth trying out the newer [`kafka_franz` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/kafka_franz/). - I’m seeing logs that report `Failed to connect to kafka: kafka: client has run out of available brokers to talk to (Is your cluster reachable?)`, but the brokers are definitely reachable. Unfortunately this error message will appear for a wide range of connection problems even when the broker endpoint can be reached. Double check your authentication configuration and also ensure that you have [enabled TLS](#tlsenabled) if applicable. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#ack_replicas)`ack_replicas` Ensure that messages have been copied across all replicas before acknowledging receipt. **Type**: `bool` **Default**: `false` ### [](#addresses)`addresses[]` A list of broker addresses to connect to. If an item of the list contains commas it will be expanded into multiple addresses. **Type**: `array` ```yaml # Examples: addresses: - "localhost:9092" # --- addresses: - "localhost:9041,localhost:9042" # --- addresses: - "localhost:9041" - "localhost:9042" ``` ### [](#backoff)`backoff` Control time intervals between retry attempts. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` The initial period to wait between retry attempts. The retry interval increases for each failed attempt, up to the `backoff.max_interval` value. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `3s` ```yaml # Examples: initial_interval: 50ms # --- initial_interval: 1s ``` ### [](#backoff-max_elapsed_time)`backoff.max_elapsed_time` The maximum overall period of time to spend on retry attempts before the request is aborted. Setting this value to a zeroed duration (such as `0s`) will result in unbounded retries. **Type**: `string` **Default**: `30s` ```yaml # Examples: max_elapsed_time: 1m # --- max_elapsed_time: 1h ``` ### [](#backoff-max_interval)`backoff.max_interval` The maximum period to wait between retry attempts **Type**: `string` **Default**: `10s` ```yaml # Examples: max_interval: 5s # --- max_interval: 1m ``` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `benthos` ### [](#compression)`compression` The compression algorithm to use. **Type**: `string` **Default**: `none` **Options**: `none`, `snappy`, `lz4`, `gzip`, `zstd` ### [](#custom_topic_creation)`custom_topic_creation` If enabled, topics will be created with the specified number of partitions and replication factor if they do not already exist. **Type**: `object` ### [](#custom_topic_creation-enabled)`custom_topic_creation.enabled` Whether to enable custom topic creation. **Type**: `bool` **Default**: `false` ### [](#custom_topic_creation-partitions)`custom_topic_creation.partitions` The number of partitions to create for new topics. Leave at -1 to use the broker configured default. Must be >= 1. **Type**: `int` **Default**: `-1` ### [](#custom_topic_creation-replication_factor)`custom_topic_creation.replication_factor` The replication factor to use for new topics. Leave at -1 to use the broker configured default. Must be an odd number, and less then or equal to the number of brokers. **Type**: `int` **Default**: `-1` ### [](#idempotent_write)`idempotent_write` Enable the idempotent write producer option. This requires the `IDEMPOTENT_WRITE` permission on `CLUSTER` and can be disabled if this permission is not available. **Type**: `bool` **Default**: `false` ### [](#inject_tracing_map)`inject_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) used to inject an object containing tracing propagation information into outbound messages. The specification of the injected fields will match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: inject_tracing_map: meta = @.merge(this) # --- inject_tracing_map: root.meta.span = this ``` ### [](#key)`key` An optional key to populate for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#max_msg_bytes)`max_msg_bytes` The maximum size in bytes of messages sent to the target topic. **Type**: `int` **Default**: `1000000` ### [](#max_retries)`max_retries` The maximum number of retries before giving up on the request. If set to zero there is no discrete limit. **Type**: `int` **Default**: `0` ### [](#metadata)`metadata` Specify criteria for which metadata values are sent with messages as headers. **Type**: `object` ### [](#metadata-exclude_prefixes)`metadata.exclude_prefixes[]` Provide a list of explicit metadata key prefixes to be excluded when adding metadata to sent messages. **Type**: `array` **Default**: `[]` ### [](#partition)`partition` The manually-specified partition to publish messages to, relevant only when the field `partitioner` is set to `manual`. Must be able to parse as a 32-bit integer. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#partitioner)`partitioner` The partitioning algorithm to use. **Type**: `string` **Default**: `fnv1a_hash` **Options**: `fnv1a_hash`, `murmur2_hash`, `random`, `round_robin`, `manual` ### [](#rack_id)`rack_id` A rack identifier for this client. **Type**: `string` **Default**: `""` ### [](#retry_as_batch)`retry_as_batch` When enabled forces an entire batch of messages to be retried if any individual message fails on a send, otherwise only the individual messages that failed are retried. Disabling this helps to reduce message duplicates during intermittent errors, but also makes it impossible to guarantee strict ordering of messages. **Type**: `bool` **Default**: `false` ### [](#sasl)`sasl` Enables SASL authentication. **Type**: `object` ### [](#sasl-access_token)`sasl.access_token` A static OAUTHBEARER access token **Type**: `string` **Default**: `""` ### [](#sasl-mechanism)`sasl.mechanism` The SASL authentication mechanism, if left empty SASL authentication is not used. **Type**: `string` **Default**: `none` | Option | Summary | | --- | --- | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. NOTE: When using plain text auth it is extremely likely that you’ll also need to enable TLS. | | SCRAM-SHA-256 | Authentication using the SCRAM-SHA-256 mechanism. | | SCRAM-SHA-512 | Authentication using the SCRAM-SHA-512 mechanism. | | none | Default, no SASL authentication. | ### [](#sasl-password)`sasl.password` A PLAIN password. It is recommended that you use environment variables to populate this field. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: ${PASSWORD} ``` ### [](#sasl-token_cache)`sasl.token_cache` Instead of using a static `access_token` allows you to query a [`cache`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) resource to fetch OAUTHBEARER tokens from **Type**: `string` **Default**: `""` ### [](#sasl-token_key)`sasl.token_key` Required when using a `token_cache`, the key to query the cache with for tokens. **Type**: `string` **Default**: `""` ### [](#sasl-user)`sasl.user` A PLAIN username. It is recommended that you use environment variables to populate this field. **Type**: `string` **Default**: `""` ```yaml # Examples: user: ${USER} ``` ### [](#static_headers)`static_headers` An optional map of static headers that should be added to messages in addition to metadata. **Type**: `object` ```yaml # Examples: static_headers: first-static-header: value-1 second-static-header: value-2 ``` ### [](#target_version)`target_version` The version of the Kafka protocol to use. This limits the capabilities used by the client and should ideally match the version of your brokers. Defaults to the oldest supported stable version. **Type**: `string` ```yaml # Examples: target_version: 2.1.0 # --- target_version: 3.1.0 ``` ### [](#timeout)`timeout` The maximum period of time to wait for message sends before abandoning the request and retrying. **Type**: `string` **Default**: `5s` ### [](#timestamp_ms)`timestamp_ms` Set a timestamp (in milliseconds) for each message (optional). When left empty, the current timestamp is used. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: timestamp_ms: ${! timestamp_unix_milli() } # --- timestamp_ms: ${! metadata("kafka_timestamp_ms") } ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#topic)`topic` The topic to publish messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 133: mongodb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/mongodb.md --- # mongodb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mongodb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/mongodb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/mongodb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/mongodb.adoc description: Inserts items into a MongoDB collection. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- Inserts items into a MongoDB collection. #### Common ```yml outputs: label: "" mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" collection: "" # No default (required) operation: update-one write_concern: w: majority j: false w_timeout: "" document_map: "" filter_map: "" hint_map: "" upsert: false max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" app_name: benthos aws: enabled: false region: "" # No default (optional) session_duration: 1h id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) roles: [] # No default (optional) collection: "" # No default (required) operation: update-one write_concern: w: majority j: false w_timeout: "" document_map: "" filter_map: "" hint_map: "" upsert: false max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#performance)Performance This output benefits from sending multiple messages in flight, in parallel, for improved performance. You can tune the maximum number of in flight messages (or message batches) using the `max_in_flight` field. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. For more information, see [Message Batching](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#app_name)`app_name` The client application name. **Type**: `string` **Default**: `benthos` ### [](#aws)`aws` AWS IAM authentication using the `MONGODB-AWS` mechanism, for example against MongoDB Atlas. When enabled, IAM credentials are used instead of a static username and password. Role-derived session credentials are resolved when the component connects and are re-resolved whenever it reconnects. The `mongodb` processor and cache establish their client once at creation and cannot refresh expiring session credentials, so `role`, `roles` and session tokens are rejected for those components; use the ambient credential chain or long-lived access keys with them. For long-running pipelines, prefer the ambient credential chain (leave keys and roles unset), which the driver refreshes automatically. **Type**: `object` ### [](#aws-enabled)`aws.enabled` Enable AWS IAM authentication using the driver-native `MONGODB-AWS` mechanism. The MongoDB Atlas database user must be created with the AWS IAM authentication type, and connections require TLS. When no static credentials or roles are configured, the ambient AWS credential chain (environment variables, EC2 instance profile, EKS pod role) is used and expiring credentials are refreshed automatically. **Type**: `bool` **Default**: `false` ### [](#aws-id)`aws.id` The ID of credentials to use. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region used when assuming roles (for STS calls). Only used when `role` or `roles` are configured; the ambient and static-key paths ignore it. If no region is specified then the environment default is used. **Type**: `string` ### [](#aws-role)`aws.role` Optional AWS IAM role ARN to assume for authentication. Cannot be combined with `roles`; use the `roles` array instead when chaining multiple roles. **Type**: `string` ### [](#aws-role_external_id)`aws.role_external_id` Optional external ID for the role assumption. Only used with the `role` field, which cannot be combined with `roles`. **Type**: `string` ### [](#aws-roles)`aws.roles[]` Optional array of AWS IAM roles to assume for authentication. Roles can be assumed in sequence, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. Cannot be combined with `role`. **Type**: `array` ### [](#aws-roles-role)`aws.roles[].role` AWS IAM role ARN to assume. **Type**: `string` **Default**: `""` ### [](#aws-roles-role_external_id)`aws.roles[].role_external_id` Optional external ID for the role assumption. **Type**: `string` **Default**: `""` ### [](#aws-secret)`aws.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-session_duration)`aws.session_duration` The duration of the STS session requested when assuming roles. AWS requires at least 15 minutes and caps sessions created through role chaining at one hour. Only used when `role` or `roles` are configured. For long-running pipelines, prefer the ambient credential chain over a fixed session duration, since the driver refreshes ambient credentials automatically as they near expiry. **Type**: `string` **Default**: `1h` ### [](#aws-token)`aws.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#collection)`collection` The name of the target collection. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#database)`database` The name of the target MongoDB database. **Type**: `string` ### [](#document_map)`document_map` A Bloblang map that represents a document to store in MongoDB, expressed as [extended JSON in canonical form](https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/). The `document_map` parameter is required for the following database operations: `insert-one`, `replace-one`, and `update-one`. **Type**: `string` **Default**: `""` ```yaml # Examples: document_map: |- root.a = this.foo root.b = this.bar ``` ### [](#filter_map)`filter_map` A Bloblang map that represents a filter for a MongoDB command, expressed as [extended JSON in canonical form](https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/). The `filter_map` parameter is required for all database operations except `insert-one`. This output uses `filter_map` to find documents for the specified operation. For example, for a `delete-one` operation, the filter map should include the fields required to locate the document for deletion. **Type**: `string` **Default**: `""` ```yaml # Examples: filter_map: |- root.a = this.foo root.b = this.bar ``` ### [](#hint_map)`hint_map` A Bloblang map that represents a hint or index for a MongoDB command to use, expressed as [extended JSON in canonical form](https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/). This map is optional, and is used with all operations except `insert-one`. Define a `hint_map` to improve performance when finding documents in the MongoDB database. **Type**: `string` **Default**: `""` ```yaml # Examples: hint_map: |- root.a = this.foo root.b = this.bar ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this number to improve throughput. **Type**: `int` **Default**: `64` ### [](#operation)`operation` The MongoDB database operation to perform. **Type**: `string` **Default**: `update-one` **Options**: `insert-one`, `delete-one`, `delete-many`, `replace-one`, `update-one` ### [](#password)`password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#upsert)`upsert` The `upsert` parameter is optional, and only applies for `update-one` and `replace-one` operations. If the filter specified in `filter_map` matches an existing document, this operation updates or replaces the document, otherwise a new document is created. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target MongoDB server. **Type**: `string` ```yaml # Examples: url: mongodb://localhost:27017 ``` ### [](#username)`username` The username required to connect to the database. **Type**: `string` **Default**: `""` ### [](#write_concern)`write_concern` The [write concern settings](https://www.mongodb.com/docs/manual/reference/write-concern/) for the MongoDB connection. **Type**: `object` ### [](#write_concern-j)`write_concern.j` The `j` requests acknowledgement from MongoDB, which is created when write operations are written to the journal. **Type**: `bool` **Default**: `false` ### [](#write_concern-w)`write_concern.w` The `w` requests acknowledgement, which write operations propagate to the specified number of MongoDB instances. **Type**: `string` **Default**: `majority` ### [](#write_concern-w_timeout)`write_concern.w_timeout` The write concern timeout. **Type**: `string` **Default**: `""` --- # Page 134: mqtt **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/mqtt.md --- # mqtt > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mqtt latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/mqtt page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/mqtt.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/mqtt.adoc description: Pushes messages to an MQTT broker. page-git-created-date: "2024-11-07" page-git-modified-date: "2026-05-26" --- Pushes messages to an MQTT broker. #### Common ```yml outputs: label: "" mqtt: urls: [] # No default (required) client_id: "" connect_timeout: 30s topic: "" # No default (required) qos: 1 write_timeout: 3s retained: false max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" mqtt: urls: [] # No default (required) client_id: "" dynamic_client_id_suffix: "" # No default (optional) connect_timeout: 30s will: enabled: false qos: 0 retained: false topic: "" payload: "" user: "" password: "" keepalive: 30 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] topic: "" # No default (required) qos: 1 write_timeout: 3s retained: false retained_interpolated: "" # No default (optional) max_in_flight: 64 ``` The `topic` field can be dynamically set using function interpolations described [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). When sending batched messages these interpolations are performed per message part. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `""` ### [](#connect_timeout)`connect_timeout` The maximum amount of time to wait in order to establish a connection before the attempt is abandoned. **Type**: `string` **Default**: `30s` ```yaml # Examples: connect_timeout: 1s # --- connect_timeout: 500ms ``` ### [](#dynamic_client_id_suffix)`dynamic_client_id_suffix` Append a dynamically generated suffix to the specified `client_id` on each run of the pipeline. This can be useful when clustering Redpanda Connect producers. **Type**: `string` | Option | Summary | | --- | --- | | nanoid | append a nanoid of length 21 characters | ### [](#keepalive)`keepalive` Max seconds of inactivity before a keepalive message is sent. **Type**: `int` **Default**: `30` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#password)`password` A password to connect with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#qos)`qos` The QoS value to set for each message. Has options 0, 1, 2. **Type**: `int` **Default**: `1` ### [](#retained)`retained` Set message as retained on the topic. **Type**: `bool` **Default**: `false` ### [](#retained_interpolated)`retained_interpolated` Override the value of `retained` with an interpolable value, this allows it to be dynamically set based on message contents. The value must resolve to either `true` or `false`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#topic)`topic` The topic to publish messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#urls)`urls[]` A list of URLs to connect to. Use the format `scheme://host:port`, where: - `scheme` is one of the following: `tcp`, `ssl`, `ws` - `host` is the IP address or hostname - `port` is the port on which the MQTT broker accepts connections If an item in the list contains commas, it is expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "tcp://localhost:1883" ``` ### [](#user)`user` A username to connect with. **Type**: `string` **Default**: `""` ### [](#will)`will` Set last will message in case of Redpanda Connect failure **Type**: `object` ### [](#will-enabled)`will.enabled` Whether to enable last will messages. **Type**: `bool` **Default**: `false` ### [](#will-payload)`will.payload` Set payload for last will message. **Type**: `string` **Default**: `""` ### [](#will-qos)`will.qos` Set QoS for last will message. Valid values are: 0, 1, 2. **Type**: `int` **Default**: `0` ### [](#will-retained)`will.retained` Set retained for last will message. **Type**: `bool` **Default**: `false` ### [](#will-topic)`will.topic` Set topic for last will message. **Type**: `string` **Default**: `""` ### [](#write_timeout)`write_timeout` The maximum amount of time to wait to write data before the attempt is abandoned. **Type**: `string` **Default**: `3s` ```yaml # Examples: write_timeout: 1s # --- write_timeout: 500ms ``` --- # Page 135: nats_jetstream **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/nats_jetstream.md --- # nats_jetstream > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats_jetstream latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/nats_jetstream page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/nats_jetstream.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/nats_jetstream.adoc description: Write messages to a NATS JetStream subject. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Write messages to a NATS JetStream subject. #### Common ```yml outputs: label: "" nats_jetstream: urls: [] # No default (required) subject: "" # No default (required) headers: {} metadata: include_prefixes: [] include_patterns: [] max_in_flight: 1024 ``` #### Advanced ```yml outputs: label: "" nats_jetstream: urls: [] # No default (required) max_reconnects: "" # No default (optional) subject: "" # No default (required) headers: {} metadata: include_prefixes: [] include_patterns: [] max_in_flight: 1024 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) inject_tracing_map: "" # No default (optional) ``` ## [](#connection-name)Connection name When monitoring and managing a production [NATS system](https://docs.nats.io/nats-concepts/overview), it is often useful to know which connection a message was sent or received from. To achieve this, set the connection name option when creating a NATS connection. Redpanda Connect can then automatically set the connection name to the NATS component label, so that monitoring tools between NATS and Redpanda Connect can stay in sync. ## [](#authentication)Authentication A number of Redpanda Connect components use NATS services. Each of these components support optional, advanced authentication parameters for [NKeys](https://docs.nats.io/nats-server/configuration/securing_nats/auth_intro/nkey_auth) and [user credentials](https://docs.nats.io/using-nats/developer/connecting/creds). For an in-depth guide, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/nats_admin/security/jwt). ### [](#nkeys)NKeys NATS server can use NKeys in several ways for authentication. The simplest approach is to configure the server with a list of user’s public keys. The server can then generate a challenge for each connection request from a client, and the client must respond to the challenge by signing it with its private NKey, configured in the `nkey_file` or `nkey` field. For more details, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth). ### [](#user-credentials)User credentials NATS server also supports decentralized authentication based on JSON Web Tokens (JWTs). When a server is configured to use this authentication scheme, clients need a [user JWT](https://docs.nats.io/nats-server/configuration/securing_nats/jwt#json-web-tokens) and a corresponding [NKey secret](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth) to connect. You can use either of the following methods to supply the user JWT and NKey secret: - In the `user_credentials_file` field, enter the path to a file containing both the private key and the JWT. You can generate the file using the [nsc tool](https://docs.nats.io/nats-tools/nsc). - In the `user_jwt` field, enter a plain text JWT, and in the `user_nkey_seed` field, enter the plain text NKey seed or private key. For more details about authentication using JWTs, see the [NATS documentation](https://docs.nats.io/using-nats/developer/connecting/creds). ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the corresponding user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#headers)`headers` Explicit message headers to add to messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: Content-Type: application/json Timestamp: ${!meta("Timestamp")} ``` ### [](#inject_tracing_map)`inject_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) used to inject an object containing tracing propagation information into outbound messages. The specification of the injected fields will match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: inject_tracing_map: meta = @.merge(this) # --- inject_tracing_map: root.meta.span = this ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `1024` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#metadata)`metadata` Determine which (if any) metadata values should be added to messages as headers. **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#subject)`subject` A subject to write to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: subject: foo.bar.baz # --- subject: ${! meta("kafka_topic") } # --- subject: foo.${! json("meta.type") } ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 136: nats_kv **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/nats_kv.md --- # nats_kv > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats_kv latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/nats_kv page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/nats_kv.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/nats_kv.adoc description: Put messages in a NATS key-value bucket. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Put messages into a NATS key-value bucket. #### Common ```yml outputs: label: "" nats_kv: urls: [] # No default (required) bucket: "" # No default (required) key: "" # No default (required) max_in_flight: 1024 ``` #### Advanced ```yml outputs: label: "" nats_kv: urls: [] # No default (required) max_reconnects: "" # No default (optional) bucket: "" # No default (required) key: "" # No default (required) max_in_flight: 1024 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) ``` The `key` field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), which lets you create a unique key for each message. ## [](#connection-name)Connection name When monitoring and managing a production [NATS system](https://docs.nats.io/nats-concepts/overview), it is often useful to know which connection a message was sent or received from. To achieve this, set the connection name option when creating a NATS connection. Redpanda Connect can then automatically set the connection name to the NATS component label, so that monitoring tools between NATS and Redpanda Connect can stay in sync. ## [](#authentication)Authentication A number of Redpanda Connect components use NATS services. Each of these components support optional, advanced authentication parameters for [NKeys](https://docs.nats.io/nats-server/configuration/securing_nats/auth_intro/nkey_auth) and [user credentials](https://docs.nats.io/using-nats/developer/connecting/creds). For an in-depth guide, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/nats_admin/security/jwt). ### [](#nkeys)NKeys NATS server can use NKeys in several ways for authentication. The simplest approach is to configure the server with a list of user’s public keys. The server can then generate a challenge for each connection request from a client, and the client must respond to the challenge by signing it with its private NKey, configured in the `nkey_file` or `nkey` field. For more details, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth). ### [](#user-credentials)User credentials NATS server also supports decentralized authentication based on JSON Web Tokens (JWTs). When a server is configured to use this authentication scheme, clients need a [user JWT](https://docs.nats.io/nats-server/configuration/securing_nats/jwt#json-web-tokens) and a corresponding [NKey secret](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth) to connect. You can use either of the following methods to supply the user JWT and NKey secret: - In the `user_credentials_file` field, enter the path to a file containing both the private key and the JWT. You can generate the file using the [nsc tool](https://docs.nats.io/nats-tools/nsc). - In the `user_jwt` field, enter a plain text JWT, and in the `user_nkey_seed` field, enter the plain text NKey seed or private key. For more details about authentication using JWTs, see the [NATS documentation](https://docs.nats.io/using-nats/developer/connecting/creds). ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the corresponding user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#bucket)`bucket` The name of the KV bucket. **Type**: `string` ```yaml # Examples: bucket: my_kv_bucket ``` ### [](#key)`key` The key for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: key: foo # --- key: foo.bar.baz # --- key: foo.${! json("meta.type") } ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `1024` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 137: nats **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/nats.md --- # nats > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/nats page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/nats.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/nats.adoc description: Publish to an NATS subject. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Publish to an NATS subject. #### Common ```yml outputs: label: "" nats: urls: [] # No default (required) subject: "" # No default (required) headers: {} metadata: include_prefixes: [] include_patterns: [] max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" nats: urls: [] # No default (required) max_reconnects: "" # No default (optional) subject: "" # No default (required) headers: {} metadata: include_prefixes: [] include_patterns: [] max_in_flight: 64 tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) inject_tracing_map: "" # No default (optional) ``` This output interpolates functions within the subject field. For a full list of functions, see [configuration:interpolation.adoc#bloblang-queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#connection-name)Connection name When monitoring and managing a production [NATS system](https://docs.nats.io/nats-concepts/overview), it is often useful to know which connection a message was sent or received from. To achieve this, set the connection name option when creating a NATS connection. Redpanda Connect can then automatically set the connection name to the NATS component label, so that monitoring tools between NATS and Redpanda Connect can stay in sync. ## [](#authentication)Authentication A number of Redpanda Connect components use NATS services. Each of these components support optional, advanced authentication parameters for [NKeys](https://docs.nats.io/nats-server/configuration/securing_nats/auth_intro/nkey_auth) and [user credentials](https://docs.nats.io/using-nats/developer/connecting/creds). For an in-depth guide, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/nats_admin/security/jwt). ### [](#nkeys)NKeys NATS server can use NKeys in several ways for authentication. The simplest approach is to configure the server with a list of user’s public keys. The server can then generate a challenge for each connection request from a client, and the client must respond to the challenge by signing it with its private NKey, configured in the `nkey_file` or `nkey` field. For more details, see the [NATS documentation](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth). ### [](#user-credentials)User credentials NATS server also supports decentralized authentication based on JSON Web Tokens (JWTs). When a server is configured to use this authentication scheme, clients need a [user JWT](https://docs.nats.io/nats-server/configuration/securing_nats/jwt#json-web-tokens) and a corresponding [NKey secret](https://docs.nats.io/running-a-nats-service/configuration/securing_nats/auth_intro/nkey_auth) to connect. You can use either of the following methods to supply the user JWT and NKey secret: - In the `user_credentials_file` field, enter the path to a file containing both the private key and the JWT. You can generate the file using the [nsc tool](https://docs.nats.io/nats-tools/nsc). - In the `user_jwt` field, enter a plain text JWT, and in the `user_nkey_seed` field, enter the plain text NKey seed or private key. For more details about authentication using JWTs, see the [NATS documentation](https://docs.nats.io/using-nats/developer/connecting/creds). ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the corresponding user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#headers)`headers` Explicit message headers to add to messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: Content-Type: application/json Timestamp: ${!meta("Timestamp")} ``` ### [](#inject_tracing_map)`inject_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) used to inject an object containing tracing propagation information into outbound messages. The specification of the injected fields will match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: inject_tracing_map: meta = @.merge(this) # --- inject_tracing_map: root.meta.span = this ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#metadata)`metadata` Determine which (if any) metadata values should be added to messages as headers. **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#subject)`subject` The subject to publish to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: subject: foo.bar.baz ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 138: opensearch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/opensearch.md --- # opensearch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: opensearch latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/opensearch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/opensearch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/opensearch.adoc description: Publishes messages into an Elasticsearch index. If the index does not exist then it is created with a dynamic mapping. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Publishes messages into an Elasticsearch index. If the index does not exist then it is created with a dynamic mapping. #### Common ```yml outputs: label: "" opensearch: urls: [] # No default (required) index: "" # No default (required) action: "" # No default (required) id: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" opensearch: urls: [] # No default (required) index: "" # No default (required) action: "" # No default (required) id: "" # No default (required) pipeline: "" routing: "" tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] max_in_flight: 64 basic_auth: enabled: false username: "" password: "" batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) aws: enabled: false region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` Both the `id` and `index` fields can be dynamically set using function interpolations described [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). When sending batched messages these interpolations are performed per message part. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#examples)Examples ### [](#updating-documents)Updating Documents When [updating documents](https://opensearch.org/docs/latest/api-reference/document-apis/update-document/) the request body should contain a combination of a `doc`, `upsert`, and/or `script` fields at the top level, this should be done via mapping processors. ```yaml output: processors: - mapping: | meta id = this.id root.doc = this opensearch: urls: [ TODO ] index: foo id: ${! @id } action: update ``` ## [](#fields)Fields ### [](#action)`action` The action to take on the document. This field must resolve to one of the following action types: `index`, `update` or `delete`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#aws)`aws` Enables and customises connectivity to Amazon Elastic Service. **Type**: `object` ### [](#aws-credentials)`aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#aws-credentials-from_ec2_role)`aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#aws-credentials-id)`aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#aws-credentials-profile)`aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#aws-credentials-role)`aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#aws-credentials-role_external_id)`aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#aws-credentials-secret)`aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-credentials-token)`aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#aws-enabled)`aws.enabled` Whether to connect to Amazon Elastic Service. **Type**: `bool` **Default**: `false` ### [](#aws-endpoint)`aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region to target. **Type**: `string` ### [](#aws-tcp)`aws.tcp` TCP socket configuration. **Type**: `object` ### [](#aws-tcp-connect_timeout)`aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#aws-tcp-keep_alive)`aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#aws-tcp-keep_alive-count)`aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#aws-tcp-keep_alive-idle)`aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#aws-tcp-keep_alive-interval)`aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#aws-tcp-tcp_user_timeout)`aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#basic_auth)`basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#id)`id` The ID for indexed messages. Interpolation should be used in order to create a unique ID for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: id: ${!counter()}-${!timestamp_unix()} ``` ### [](#index)`index` The index to place messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#pipeline)`pipeline` An optional pipeline id to preprocess incoming documents. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#routing)`routing` The routing key to use for the document. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If an item of the list contains commas it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "http://localhost:9200" ``` --- # Page 139: otlp_grpc **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/otlp_grpc.md --- # otlp_grpc > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: otlp_grpc latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/otlp_grpc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/otlp_grpc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/otlp_grpc.adoc description: Send OpenTelemetry traces, logs, and metrics via OTLP/gRPC protocol. page-git-created-date: "2026-01-23" page-git-modified-date: "2026-08-11" --- Send OpenTelemetry traces, logs, and metrics via OTLP/gRPC protocol. Sends OpenTelemetry telemetry data to a remote collector via OTLP/gRPC protocol. Accepts batches of Redpanda OTEL v1 protobuf messages (spans, log records, or metrics) and converts them to OTLP format for transmission to OpenTelemetry collectors. #### Common ```yml outputs: label: "" otlp_grpc: endpoint: "" # No default (required) max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" otlp_grpc: endpoint: "" # No default (required) headers: {} timeout: 30s compression: gzip tls: enabled: false skip_cert_verify: false cert_file: "" key_file: "" tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} max_in_flight: 64 ``` ## [](#input-format)Input format Expects messages in Redpanda OTEL v1 protobuf format with metadata: - `signal_type`: "trace", "log", or "metric" Each batch must contain messages of the same signal type. The entire batch is converted to a single OTLP export request and sent via gRPC. ## [](#authentication)Authentication Supports multiple authentication methods: - Bearer token authentication (via `auth_token` field) - OAuth v2 (via `oauth2` configuration block) > 📝 **NOTE** > > OAuth2 requires TLS to be enabled. ## [](#fields)Fields ### [](#compression)`compression` Compression type for gRPC requests. Options: 'gzip' or 'none'. **Type**: `string` **Default**: `gzip` **Options**: `gzip`, `none` ### [](#endpoint)`endpoint` The gRPC endpoint of the remote OTLP collector. **Type**: `string` ### [](#headers)`headers` A map of headers to add to the gRPC request metadata. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: X-Custom-Header: value traceparent: ${! tracing_span().traceparent } ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#oauth2)`oauth2` Allows you to specify open authentication via OAuth version 2 using the client credentials token flow. **Type**: `object` ### [](#oauth2-client_key)`oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#oauth2-client_secret)`oauth2.client_secret` A secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth2-enabled)`oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2-endpoint_params)`oauth2.endpoint_params` A list of optional endpoint parameters, values should be arrays of strings. **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: audience: - https://example.com resource: - https://api.example.com ``` ### [](#oauth2-scopes)`oauth2.scopes[]` A list of optional requested permissions. **Type**: `array` **Default**: `[]` ### [](#oauth2-token_url)`oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` Timeout for gRPC requests. **Type**: `string` **Default**: `30s` ### [](#tls)`tls` TLS configuration for gRPC client. **Type**: `object` ### [](#tls-cert_file)`tls.cert_file` Path to the TLS certificate file for client authentication. **Type**: `string` **Default**: `""` ### [](#tls-enabled)`tls.enabled` Enable TLS connections. **Type**: `bool` **Default**: `false` ### [](#tls-key_file)`tls.key_file` Path to the TLS key file for client authentication. **Type**: `string` **Default**: `""` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Skip certificate verification (insecure). **Type**: `bool` **Default**: `false` --- # Page 140: otlp_http **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/otlp_http.md --- # otlp_http > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: otlp_http latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/otlp_http page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/otlp_http.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/otlp_http.adoc description: Send OpenTelemetry traces, logs, and metrics via OTLP/HTTP protocol. page-git-created-date: "2026-01-23" page-git-modified-date: "2026-08-11" --- Send OpenTelemetry traces, logs, and metrics via OTLP/HTTP protocol. Sends OpenTelemetry telemetry data to a remote collector via OTLP/HTTP protocol. Accepts batches of Redpanda OTEL v1 protobuf messages (spans, log records, or metrics) and converts them to OTLP format for transmission to OpenTelemetry collectors. #### Common ```yml outputs: label: "" otlp_http: endpoint: "" # No default (required) max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" otlp_http: endpoint: "" # No default (required) content_type: protobuf headers: {} timeout: 30s proxy_url: "" follow_redirects: false disable_http2: false tls: enabled: false skip_cert_verify: false cert_file: "" key_file: "" tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} max_in_flight: 64 ``` ## [](#input-format)Input format Expects messages in Redpanda OTEL v1 protobuf format with metadata: - `signal_type`: "trace", "log", or "metric" Each batch must contain messages of the same signal type. The entire batch is converted to a single OTLP export request and sent via HTTP POST. ## [](#endpoints)Endpoints The output automatically appends the signal type path to the base endpoint: - Traces: `{endpoint}/v1/traces` - Logs: `{endpoint}/v1/logs` - Metrics: `{endpoint}/v1/metrics` ## [](#content-types)Content types Supports two content types: - `protobuf` (default): `application/x-protobuf` - `json`: `application/json` ## [](#authentication)Authentication Supports multiple authentication methods: - Basic authentication - OAuth v1 - OAuth v2 - JWT ## [](#fields)Fields ### [](#basic_auth)`basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#content_type)`content_type` Content type for HTTP requests. Options: 'protobuf' or 'json'. **Type**: `string` **Default**: `protobuf` **Options**: `protobuf`, `json` ### [](#disable_http2)`disable_http2` Whether or not to disable HTTP/2. **Type**: `bool` **Default**: `false` ### [](#endpoint)`endpoint` The HTTP endpoint of the remote OTLP collector (without the signal path). **Type**: `string` ### [](#follow_redirects)`follow_redirects` Transparently follow redirects, i.e. responses with 300-399 status codes. If disabled, the response message will contain the body, status, and headers from the redirect response and the processor will not make a request to the URL set in the Location header of the response. **Type**: `bool` **Default**: `false` ### [](#headers)`headers` A map of headers to add to the request. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: X-Custom-Header: value traceparent: ${! tracing_span().traceparent } ``` ### [](#jwt)`jwt` (beta) Allows you to specify JWT authentication. **Type**: `object` ### [](#jwt-claims)`jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` Add optional key/value headers to the JWT. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` A file with the PEM encoded via PKCS1 or PKCS8 as private key. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` A method used to sign the token such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#oauth)`oauth` Allows you to specify open authentication via OAuth version 1. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2)`oauth2` Allows you to specify open authentication via OAuth version 2 using the client credentials token flow. **Type**: `object` ### [](#oauth2-client_key)`oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#oauth2-client_secret)`oauth2.client_secret` A secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth2-enabled)`oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2-endpoint_params)`oauth2.endpoint_params` A list of optional endpoint parameters, values should be arrays of strings. **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: audience: - https://example.com resource: - https://api.example.com ``` ### [](#oauth2-scopes)`oauth2.scopes[]` A list of optional requested permissions. **Type**: `array` **Default**: `[]` ### [](#oauth2-token_url)`oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#proxy_url)`proxy_url` An optional HTTP proxy URL. **Type**: `string` **Default**: `""` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` Timeout for HTTP requests. **Type**: `string` **Default**: `30s` ### [](#tls)`tls` TLS configuration for HTTP client. **Type**: `object` ### [](#tls-cert_file)`tls.cert_file` Path to the TLS certificate file for client authentication. **Type**: `string` **Default**: `""` ### [](#tls-enabled)`tls.enabled` Enable TLS connections. **Type**: `bool` **Default**: `false` ### [](#tls-key_file)`tls.key_file` Path to the TLS key file for client authentication. **Type**: `string` **Default**: `""` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Skip certificate verification (insecure). **Type**: `bool` **Default**: `false` --- # Page 141: pinecone **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/pinecone.md --- # pinecone > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: pinecone latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/pinecone page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/pinecone.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/pinecone.adoc description: Inserts items into a Pinecone index. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Inserts items into a Pinecone index. #### Common ```yml outputs: label: "" pinecone: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) host: "" # No default (required) api_key: "" # No default (required) operation: upsert-vectors id: "" # No default (required) vector_mapping: "" # No default (optional) metadata_mapping: "" # No default (optional) ``` #### Advanced ```yml outputs: label: "" pinecone: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) host: "" # No default (required) api_key: "" # No default (required) operation: upsert-vectors namespace: "" id: "" # No default (required) vector_mapping: "" # No default (optional) metadata_mapping: "" # No default (optional) ``` ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#api_key)`api_key` The Pinecone API key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#host)`host` The host for the Pinecone index. **Type**: `string` ### [](#id)`id` The ID for the index entry in Pinecone. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#metadata_mapping)`metadata_mapping` An optional mapping of message to metadata in the Pinecone index entry. **Type**: `string` ```yaml # Examples: metadata_mapping: root = @ # --- metadata_mapping: root = metadata() # --- metadata_mapping: root = {"summary": this.summary, "foo": this.other_field} ``` ### [](#namespace)`namespace` The namespace to write to - writes to the default namespace by default. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#operation)`operation` The operation to perform against the Pinecone index. **Type**: `string` **Default**: `upsert-vectors` **Options**: `update-vector`, `upsert-vectors`, `delete-vectors` ### [](#vector_mapping)`vector_mapping` The mapping to extract out the vector from the document. The result must be a floating point array. Required if not a delete operation. **Type**: `string` ```yaml # Examples: vector_mapping: root = this.embeddings_vector # --- vector_mapping: root = [1.2, 0.5, 0.76] ``` --- # Page 142: qdrant **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/qdrant.md --- # qdrant > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: qdrant latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/qdrant page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/qdrant.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/qdrant.adoc description: Adds items to a Qdrant collection. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Adds items to a [Qdrant](https://qdrant.tech/) collection #### Common ```yml outputs: label: "" qdrant: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) grpc_host: "" # No default (required) api_token: "" collection_name: "" # No default (required) id: "" # No default (required) vector_mapping: "" # No default (required) payload_mapping: root = {} ``` #### Advanced ```yml outputs: label: "" qdrant: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) grpc_host: "" # No default (required) api_token: "" tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] collection_name: "" # No default (required) id: "" # No default (required) vector_mapping: "" # No default (required) payload_mapping: root = {} ``` ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#api_token)`api_token` The Qdrant API token for authentication. Defaults to an empty string. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#collection_name)`collection_name` The name of the collection in Qdrant. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#grpc_host)`grpc_host` The gRPC host of the Qdrant server. **Type**: `string` ```yaml # Examples: grpc_host: localhost:6334 # --- grpc_host: xyz-example.eu-central.aws.cloud.qdrant.io:6334 ``` ### [](#id)`id` The ID of the point to insert. Can be a UUID string or positive integer. **Type**: `string` ```yaml # Examples: id: root = "dc88c126-679f-49f5-ab85-04b77e8c2791" # --- id: root = 832 ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#payload_mapping)`payload_mapping` An optional mapping of message to payload associated with the point. **Type**: `string` **Default**: `root = {}` ```yaml # Examples: payload_mapping: root = {"field": this.value, "field_2": 987} # --- payload_mapping: root = metadata() ``` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#vector_mapping)`vector_mapping` The mapping to extract the vector from the document. **Type**: `string` ```yaml # Examples: vector_mapping: root = {"dense_vector": [0.352,0.532,0.754],"sparse_vector": {"indices": [23,325,532],"values": [0.352,0.532,0.532]}, "multi_vector": [[0.352,0.532],[0.352,0.532]]} # --- vector_mapping: root = [1.2, 0.5, 0.76] # --- vector_mapping: root = this.vector # --- vector_mapping: root = [[0.352,0.532,0.532,0.234],[0.352,0.532,0.532,0.234]] # --- vector_mapping: root = {"some_sparse": {"indices":[23,325,532],"values":[0.352,0.532,0.532]}} # --- vector_mapping: root = {"some_multi": [[0.352,0.532,0.532,0.234],[0.352,0.532,0.532,0.234]]} # --- vector_mapping: root = {"some_dense": [0.352,0.532,0.532,0.234]} ``` --- # Page 143: questdb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/questdb.md --- # questdb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: questdb page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/questdb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/questdb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/questdb.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-11-07" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Pushes messages to a [QuestDB](https://questdb.io/docs/) table. #### Common ```yml outputs: label: "" questdb: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) address: "" # No default (required) username: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) table: "" # No default (required) designated_timestamp_field: "" # No default (optional) designated_timestamp_unit: auto timestamp_string_fields: [] # No default (optional) timestamp_string_format: Jan _2 15:04:05.000000Z0700 symbols: [] # No default (optional) doubles: [] # No default (optional) error_on_empty_messages: false ``` #### Advanced ```yml outputs: label: "" questdb: max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] address: "" # No default (required) username: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) retry_timeout: "" # No default (optional) request_timeout: "" # No default (optional) request_min_throughput: "" # No default (optional) table: "" # No default (required) designated_timestamp_field: "" # No default (optional) designated_timestamp_unit: auto timestamp_string_fields: [] # No default (optional) timestamp_string_format: Jan _2 15:04:05.000000Z0700 symbols: [] # No default (optional) doubles: [] # No default (optional) error_on_empty_messages: false ``` > ❗ **IMPORTANT** > > Redpanda Data recommends enabling the dedupe feature on the QuestDB server. For more information about deploying, configuring, and using QuestDB, see the [QuestDB documentation](https://questdb.io/docs/). ## [](#performance)Performance For improved performance, this output sends multiple messages in parallel. You can tune the maximum number of in-flight messages (or message batches), using the `max_in_flight` field. You can configure batches at both the input and output level. For more information, see [Message Batching](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#address)`address` The host and port of the QuestDB server. **Type**: `string` ```yaml # Examples: address: localhost:9000 ``` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#designated_timestamp_field)`designated_timestamp_field` The name of the designated timestamp field in QuestDB. **Type**: `string` ### [](#designated_timestamp_unit)`designated_timestamp_unit` Units used for the designated timestamp field in QuestDB. **Type**: `string` **Default**: `auto` ### [](#doubles)`doubles[]` Columns that must be the `double` type, with `int` as the default. **Type**: `array` ### [](#error_on_empty_messages)`error_on_empty_messages` Mark a message as an error if it is empty after field validation. **Type**: `bool` **Default**: `false` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this value to improve throughput. **Type**: `int` **Default**: `64` ### [](#password)`password` The password to use for basic authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#request_min_throughput)`request_min_throughput` The minimum expected throughput in bytes per second for HTTP requests. If the throughput is lower than this value, the connection times out. The `quest_db` output uses this value to calculate an additional timeout on top of the `request_timeout`. This setting is useful for large requests. Set it to `0` to disable this logic. **Type**: `int` ### [](#request_timeout)`request_timeout` The period of time to wait for a response from the QuestDB server in addition to any connection timeout calculated for the `request_min_throughput` field. **Type**: `string` ### [](#retry_timeout)`retry_timeout` The period of time to continue retrying after a failed HTTP request. The interval between retries is an exponential backoff starting at 10 ms, and doubling after each failed attempt up to a maximum of 1 second. **Type**: `string` ### [](#symbols)`symbols[]` Columns that must be the `symbol` type. String values default to `string` types. **Type**: `array` ### [](#table)`table` The destination table in QuestDB. **Type**: `string` ```yaml # Examples: table: trades ``` ### [](#timestamp_string_fields)`timestamp_string_fields[]` String fields with textual timestamps. **Type**: `array` ### [](#timestamp_string_format)`timestamp_string_format` The timestamp format, which is used when parsing timestamp string fields and uses Golang’s time formatting. **Type**: `string` **Default**: `Jan _2 15:04:05.000000Z0700` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#token)`token` The bearer token to use for authentication, which takes precedence over the basic authentication username and password. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#username)`username` The username to use for basic authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` --- # Page 144: redis_hash **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redis_hash.md --- # redis_hash > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_hash latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/redis_hash page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/redis_hash.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/redis_hash.adoc description: Sets Redis hash objects using the HSET command. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sets Redis hash objects using the HMSET command. #### Common ```yml outputs: label: "" redis_hash: url: "" # No default (required) key: "" # No default (required) walk_metadata: false walk_json_object: false fields: {} max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" redis_hash: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] key: "" # No default (required) walk_metadata: false walk_json_object: false fields: {} max_in_flight: 64 ``` The field `key` supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), allowing you to create a unique key for each message. The field `fields` allows you to specify an explicit map of field names to interpolated values, also evaluated per message of a batch: ```yaml output: redis_hash: url: tcp://localhost:6379 key: ${!json("id")} fields: topic: ${!meta("kafka_topic")} partition: ${!meta("kafka_partition")} content: ${!json("document.text")} ``` If the field `walk_metadata` is set to `true` then Redpanda Connect will walk all metadata fields of messages and add them to the list of hash fields to set. If the field `walk_json_object` is set to `true` then Redpanda Connect will walk each message as a JSON object, extracting keys and the string representation of their value and adds them to the list of hash fields to set. The order of hash field extraction is as follows: 1. Metadata (if enabled) 2. JSON object (if enabled) 3. Explicit fields Where latter stages will overwrite matching field names of a former stage. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#fields-2)`fields` A map of key/value pairs to set as hash fields. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ### [](#key)`key` The key for each message, function interpolations should be used to create a unique key per message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: key: ${! @.kafka_key } # --- key: ${! this.doc.id } # --- key: ${! counter() } ``` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` ### [](#walk_json_object)`walk_json_object` Whether to walk each message as a JSON object and add each key/value pair to the list of hash fields to set. **Type**: `bool` **Default**: `false` ### [](#walk_metadata)`walk_metadata` Whether all metadata fields of messages should be walked and added to the list of hash fields to set. **Type**: `bool` **Default**: `false` --- # Page 145: redis_list **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redis_list.md --- # redis_list > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_list latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/redis_list page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/redis_list.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/redis_list.adoc description: Pushes messages onto the end of a Redis list (which is created if it doesn't already exist) using the RPUSH command. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Pushes messages onto the end of a Redis list (which is created if it doesn’t already exist) using the RPUSH command. #### Common ```yml outputs: label: "" redis_list: url: "" # No default (required) key: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" redis_list: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] key: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) command: rpush ``` The field `key` supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries), allowing you to create a unique key for each message. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#command)`command` The command used to push elements to the Redis list **Type**: `string` **Default**: `rpush` **Options**: `rpush`, `lpush` ### [](#key)`key` The key for each message, function interpolations can be optionally used to create a unique key per message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: key: some_list # --- key: ${! @.kafka_key } # --- key: ${! this.doc.id } # --- key: ${! counter() } ``` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 146: redis_pubsub **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redis_pubsub.md --- # redis_pubsub > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_pubsub latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/redis_pubsub page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/redis_pubsub.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/redis_pubsub.adoc description: Publishes messages through the Redis PubSub model. It is not possible to guarantee that messages have been received. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Publishes messages through the Redis PubSub model. It is not possible to guarantee that messages have been received. #### Common ```yml outputs: label: "" redis_pubsub: url: "" # No default (required) channel: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" redis_pubsub: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] channel: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` This output will interpolate functions within the channel field, you can find a list of functions [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#channel)`channel` The channel to publish messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 147: redis_streams **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redis_streams.md --- # redis_streams > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_streams latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/redis_streams page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/redis_streams.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/redis_streams.adoc description: Pushes messages to a Redis (v5.0+) Stream (which is created if it doesn't already exist) using the XADD command. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Pushes messages to a Redis (v5.0+) Stream (which is created if it doesn’t already exist) using the XADD command. #### Common ```yml outputs: label: "" redis_streams: url: "" # No default (required) stream: "" # No default (required) id: * body_key: body max_length: 0 max_in_flight: 64 metadata: exclude_prefixes: [] batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" redis_streams: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] stream: "" # No default (required) id: * body_key: body max_length: 0 max_in_flight: 64 metadata: exclude_prefixes: [] batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` It’s possible to specify a maximum length of the target stream by setting it to a value greater than 0, in which case this cap is applied only when Redis is able to remove a whole macro node, for efficiency. Redis stream entries are key/value pairs, as such it is necessary to specify the key to be set to the body of the message. All metadata fields of the message will also be set as key/value pairs, if there is a key collision between a metadata item and the body then the body takes precedence. ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#body_key)`body_key` A key to set the raw body of the message to. **Type**: `string` **Default**: `body` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#id)`id` The entry ID for the stream message. Allows function interpolations. When set to `*` (the default), Redis auto-generates a unique ID based on the current time. Set a custom ID to control message ordering, for example to replay messages in upstream order. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `*` ```yaml # Examples: id: * # --- id: ${! @redis_stream } # --- id: ${! this.id } # --- id: ${! counter() }-0 ``` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#max_length)`max_length` When greater than zero enforces a rough cap on the length of the target stream. **Type**: `int` **Default**: `0` ### [](#metadata)`metadata` Specify criteria for which metadata values are included in the message body. **Type**: `object` ### [](#metadata-exclude_prefixes)`metadata.exclude_prefixes[]` Provide a list of explicit metadata key prefixes to be excluded when adding metadata to sent messages. **Type**: `array` **Default**: `[]` ### [](#stream)`stream` The stream to add messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 148: redpanda_common **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_common.md --- # redpanda_common > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda_common latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/redpanda_common page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/redpanda_common.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/redpanda_common.adoc description: Sends data to a Redpanda (Kafka) broker, using credentials defined in a common top-level redpanda config block. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- > ⚠️ **WARNING: Deprecated in 4.68.0** > > Deprecated in 4.68.0 > > This component is deprecated and will be removed in the next major version release. Please consider moving onto the unified [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [`redpanda` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) components. Sends data to a Redpanda (Kafka) broker, using credentials from a common `redpanda` configuration block. To avoid duplicating Redpanda cluster credentials in your `redpanda_common` input, output, or any other components in your data pipeline, you can use a single [`redpanda` configuration block](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/redpanda/about/). For more details, see the [Pipeline example](#pipeline-example). > 📝 **NOTE** > > If you need to move topic data between Redpanda clusters or other Apache Kafka clusters, consider using the [`redpanda` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda/) and [output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) instead. #### Common ```yml outputs: label: "" redpanda_common: topic: "" # No default (required) key: "" # No default (optional) partition: "" # No default (optional) metadata: include_prefixes: [] include_patterns: [] max_in_flight: 10 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" redpanda_common: topic: "" # No default (required) key: "" # No default (optional) partition: "" # No default (optional) metadata: include_prefixes: [] include_patterns: [] timestamp_ms: "" # No default (optional) max_in_flight: 10 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#pipeline-example)Pipeline example This data pipeline reads data from `topic_A` and `topic_B` on a Redpanda cluster, and then writes the data to `topic_C` on the same cluster. The cluster details are configured within the `redpanda` configuration block, so you only need to configure them once. This is a useful feature when you have multiple inputs and outputs in the same data pipeline that need to connect to the same cluster. ```none input: redpanda_common: topics: [ topic_A, topic_B ] output: redpanda_common: topic: topic_C key: ${! @id } redpanda: seed_brokers: [ "127.0.0.1:9092" ] tls: enabled: true sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#key)`key` A key to populate for each message (optional). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this number to improve throughput until performance plateaus. **Type**: `int` **Default**: `10` ### [](#metadata)`metadata` Configure which metadata values are added to messages as headers. This allows you to pass additional context information along with your messages. **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#partition)`partition` Set a partition for each message (optional). This field is only relevant when the `partitioner` is set to `manual`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). You must provide an interpolation string that is a valid integer. **Type**: `string` ```yaml # Examples: partition: ${! meta("partition") } ``` ### [](#timestamp_ms)`timestamp_ms` Set a timestamp (in milliseconds) for each message (optional). When left empty, the current timestamp is used. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: timestamp_ms: ${! timestamp_unix_milli() } # --- timestamp_ms: ${! metadata("kafka_timestamp_ms") } ``` ### [](#topic)`topic` A topic to write messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 149: redpanda_migrator **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_migrator.md --- # redpanda_migrator > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda_migrator latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/redpanda_migrator page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/redpanda_migrator.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/redpanda_migrator.adoc description: A specialised Kafka producer for comprehensive data migration between Apache Kafka and Redpanda clusters. page-git-created-date: "2024-10-02" page-git-modified-date: "2026-05-26" --- Migrates topics, schemas, and consumer groups between Kafka and Redpanda clusters. > ❗ **IMPORTANT** > > Pair this output with a [`redpanda_migrator` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda_migrator/) in the same pipeline. The following shows all available configuration fields and their defaults. #### Common ```yml outputs: label: "" redpanda_migrator: seed_brokers: [] # No default (required) schema_registry: url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} enabled: true interval: 5m include: [] # No default (optional) exclude: [] # No default (optional) subject: "" # No default (optional) versions: all include_deleted: false translate_ids: false normalize: false strict: false max_parallel_http_requests: 10 consumer_groups: enabled: true interval: 1m fetch_timeout: 10s include: [] # No default (optional) exclude: [] # No default (optional) only_empty: false topic: ${! @kafka_topic } topic_replication_factor: "" # No default (optional) sync_topic_acls: false headers: "" # No default (optional) max_in_flight: 10 ``` #### Advanced ```yml outputs: label: "" redpanda_migrator: seed_brokers: [] # No default (required) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s partitioner: "" # No default (optional) idempotent_write: true acks: all compression: "" # No default (optional) allow_auto_topic_creation: true timeout: 10s max_message_bytes: 1MiB broker_write_max_bytes: 100MiB max_buffered_records: 10000 max_buffered_bytes: 0 max_in_flight_requests: 1 record_retries: 0 record_delivery_timeout: 0s schema_registry: url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} enabled: true interval: 5m include: [] # No default (optional) exclude: [] # No default (optional) subject: "" # No default (optional) versions: all include_deleted: false translate_ids: false normalize: false strict: false max_parallel_http_requests: 10 consumer_groups: enabled: true interval: 1m fetch_timeout: 10s include: [] # No default (optional) exclude: [] # No default (optional) only_empty: false topic: ${! @kafka_topic } topic_replication_factor: "" # No default (optional) sync_topic_interval: 5m sync_topic_acls: false serverless: false headers: "" # No default (optional) provenance_header: redpanda-migrator-provenance offset_header: redpanda-migrator-offset max_in_flight: 10 ``` ## [](#requirements)Requirements When the destination cluster enforces ACLs, the destination principal needs permission to create topics and add partitions, not only to produce records. Grant these ACLs at minimum: - Topic `CREATE`, `WRITE`, `ALTER`, and `DESCRIBE_CONFIGS`. Cluster `CREATE` also authorizes topic creation. - When consumer group migration is enabled, consumer group `READ`. - When `sync_topic_acls` is enabled, cluster `ALTER`. For the source-principal ACLs and full details, see [Required permissions](https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/redpanda_migrator/#required-permissions). ## [](#multiple-migrator-pairs)Multiple migrator pairs Each migrator pair requires a unique `label`. Set the same label value on both the input and output within a pair. Labels must match exactly; mismatched labels prevent the input and output from coordinating. ## [](#performance-tuning)Performance tuning For high-throughput workloads, adjust the following settings: On this output: - `max_in_flight`: Set to the total number of partitions being migrated. Higher values provide no benefit beyond the partition count. On the paired [`redpanda_migrator` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda_migrator/#performance-tuning): - `partition_buffer_bytes`: Set to `2MB` to increase the per-partition buffer. - `max_yield_batch_bytes`: Set to `1MB` to yield larger batches. ## [](#synchronization-details)Synchronization details ### [](#topics)Topics - Topic names carry over from the source by default. Use the `topic` field with interpolation to rename topics at the destination. - The migrator creates each topic at the destination with the same partition count as the source. - Replication factor defaults to the source value. Set `topic_replication_factor` to override it. - The migrator copies a serverless-aware subset of topic configuration keys. - ACL replication is optional (`sync_topic_acls`): - The migrator does not copy `ALLOW WRITE` ACL entries. - The migrator downgrades `ALLOW ALL` ACL entries to `ALLOW READ`. - Resource pattern type and host filters carry over. ### [](#schema-registry)Schema Registry - Syncs once at startup, then periodically based on `schema_registry.interval` (default: every 5 minutes). Set `schema_registry.interval: 0s` for a one-time sync only. - Use include and exclude regex patterns to filter which subjects are migrated. - Use the `subject` field with interpolation to rename subjects at the destination. - Schema versions: `latest` (most recent version only) or `all` (full history). Defaults to `all`. - Include soft-deleted subjects with `schema_registry.include_deleted`. - The migrator can translate schema IDs to new destination IDs or preserve them as-is. See `schema_registry.translate_ids`. - Schema normalization is optional (`schema_registry.normalize`). - Compatibility settings carry over per-subject. Schema metadata and rules do not carry over in Serverless mode. ### [](#consumer-groups)Consumer groups - Sync periodically based on `consumer_groups.interval` (default: every 1 minute). - Use include and exclude regex patterns to filter which groups are migrated. - By default, all groups except those in `Dead` state migrate. Set `consumer_groups.only_empty: true` to migrate only `Empty` state groups. - The migrator translates consumer offsets using timestamps. This is approximate and may be imprecise when multiple records share the same timestamp. - Offsets only move forward and never rewind at the destination. - Source and destination must have matching partition counts for each migrated group. ## [](#how-it-works)How it works Each migration component synchronizes on a different schedule: - **Topics**: Sync from source on startup and every 5 minutes by default, including source topics that have no current data (for example, after retention cleanup). Configure the sync interval with [`sync_topic_interval`](#sync_topic_interval), or set it to `0s` to disable periodic sync. When periodic sync is disabled, topics are still created on demand when the first message arrives. - **Schema Registry**: Syncs at startup, then periodically as configured. - **Consumer groups**: Sync in the background, filtered to the topics being migrated. ## [](#guarantees)Guarantees The migrator upholds the following guarantees: - Creates each destination topic with the intended partition count and replication factor. - Never overwrites existing destination topics and logs any partition count mismatches. - Consumer group offsets never rewind. - ACL replication never grants write access at the destination. ## [](#limitations)Limitations - The destination cluster’s Schema Registry must be in `READWRITE` or `IMPORT` mode. - Offset translation is best-effort. - Consumer group migration requires identical partition counts at source and destination. ## [](#metrics)Metrics | Metric Name | Type | Labels | Description | | --- | --- | --- | --- | | Topic migration | | | | | redpanda_migrator_topics_created_total | counter | | Total topics created on destination | | redpanda_migrator_topic_create_errors_total | counter | | Topic creation errors | | redpanda_migrator_topic_create_latency_ns | timer | | Topic creation latency (ns) | | Schema Registry migration | | | | | redpanda_migrator_sr_schemas_created_total | counter | | Schemas created in destination registry | | redpanda_migrator_sr_schema_create_errors_total | counter | | Schema creation errors | | redpanda_migrator_sr_schema_create_latency_ns | timer | | Schema creation latency (ns) | | redpanda_migrator_sr_compatibility_updates_total | counter | | Compatibility level updates applied | | redpanda_migrator_sr_compatibility_update_errors_total | counter | | Compatibility update errors | | redpanda_migrator_sr_compatibility_update_latency_ns | timer | | Compatibility update latency (ns) | | Consumer group migration | | | | | redpanda_migrator_cg_offsets_translated_total | counter | group | Offsets translated per consumer group | | redpanda_migrator_cg_offset_translation_errors_total | counter | group | Offset translation errors per group | | redpanda_migrator_cg_offset_translation_latency_ns | timer | group | Offset translation latency per group (ns) | | redpanda_migrator_cg_offsets_committed_total | counter | group | Offsets committed per consumer group | | redpanda_migrator_cg_offset_commit_errors_total | counter | group | Offset commit errors per group | | redpanda_migrator_cg_offset_commit_latency_ns | timer | group | Offset commit latency per group (ns) | | Consumer lag | | | | | redpanda_lag | gauge | topic, partition | Current consumer lag in messages for each topic partition. Shows difference between high water mark and current consumer position. | ## [](#examples)Examples ### [](#basic-migration)Basic migration Migrate topics, schemas and consumer groups from source to destination. ```yaml input: redpanda_migrator: seed_brokers: ["source:9092"] topics: ["orders", "payments"] consumer_group: "migration" output: redpanda_migrator: seed_brokers: ["destination:9092"] # Write to the same topic name topic: ${! metadata("kafka_topic") } schema_registry: url: "http://dest-registry:8081" translate_ids: true consumer_groups: interval: 1m ``` ### [](#migration-to-redpanda-serverless)Migration to Redpanda Serverless Migrate from Confluent/Kafka to Redpanda Cloud serverless cluster with authentication. ```yaml input: redpanda_migrator: seed_brokers: ["source-kafka:9092"] regexp_topics_include: - '.' regexp_topics_exclude: - '^_' consumer_group: "migrator_cg" schema_registry: url: "http://source-registry:8081" output: redpanda_migrator: seed_brokers: ["serverless-cluster.redpanda.com:9092"] tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: "migrator" password: "migrator" schema_registry: url: "https://serverless-cluster.redpanda.com:8081" basic_auth: enabled: true username: "migrator" password: "migrator" translate_ids: true consumer_groups: exclude: - "migrator_cg" # Exclude the migration consumer group itself serverless: true # Enable serverless mode for restricted configurations ``` ## [](#fields)Fields ### [](#acks)`acks` The number of acknowledgements the leader broker must receive from ISR brokers before responding to the produce request. When `idempotent_write` is enabled this must be set to `all`. **Type**: `string` **Default**: `all` | Option | Summary | | --- | --- | | all | Wait for all in-sync replicas to acknowledge (acks=-1). Required when idempotent_write is enabled. | | leader | Wait for the leader broker to acknowledge (acks=1). Messages are lost if the leader fails before replication. | | none | Do not wait for any acknowledgement (acks=0). Highest throughput but messages may be lost. | ### [](#allow_auto_topic_creation)`allow_auto_topic_creation` Enables topics to be auto created if they do not exist when fetching their metadata. **Type**: `bool` **Default**: `true` ### [](#broker_write_max_bytes)`broker_write_max_bytes` The maximum number of bytes this output can write to a broker connection in a single write. This field corresponds to Kafka’s `socket.request.max.bytes`. **Type**: `string` **Default**: `100MiB` ```yaml # Examples: broker_write_max_bytes: 128MB # --- broker_write_max_bytes: 50mib ``` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#compression)`compression` Set an explicit compression type (optional). The default preference is to use `snappy` when the broker supports it. Otherwise, use `none`. **Type**: `string` **Options**: `lz4`, `snappy`, `gzip`, `none`, `zstd` ### [](#conn_idle_timeout)`conn_idle_timeout` The maximum duration that connections can remain idle before they are automatically closed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `20s` ### [](#consumer_groups)`consumer_groups` **Type**: `object` ### [](#consumer_groups-enabled)`consumer_groups.enabled` Whether consumer group offset migration is enabled. When disabled, no consumer group operations are performed. **Type**: `bool` **Default**: `true` ### [](#consumer_groups-exclude)`consumer_groups.exclude[]` Regular expressions for consumer groups to exclude from offset migration. Takes precedence over include patterns. Useful for excluding system or temporary groups. **Type**: `array` ```yaml # Examples: exclude: [".*-test", ".*-temp", "connect-.*"] # --- exclude: ["dev-.*", "local-.*"] ``` ### [](#consumer_groups-fetch_timeout)`consumer_groups.fetch_timeout` Maximum time to wait for data when fetching records for timestamp-based offset translation. Increase for clusters with low message throughput. **Type**: `string` **Default**: `10s` ```yaml # Examples: fetch_timeout: 1s # Fast clusters # --- fetch_timeout: 10s # Slower clusters ``` ### [](#consumer_groups-include)`consumer_groups.include[]` Regular expressions for consumer groups to include in offset migration. If empty, all groups are included (unless excluded). **Type**: `array` ```yaml # Examples: include: ["prod-.*", "staging-.*"] # --- include: ["app-.*", "service-.*"] ``` ### [](#consumer_groups-interval)`consumer_groups.interval` How often to synchronise consumer group offsets. Regular syncing helps maintain offset accuracy during ongoing migration. **Type**: `string` **Default**: `1m` ```yaml # Examples: interval: 0s # Disabled # --- interval: 30s # Sync every 30 seconds # --- interval: 5m # Sync every 5 minutes ``` ### [](#consumer_groups-only_empty)`consumer_groups.only_empty` Whether to only migrate Empty consumer groups. When false (default), all statuses except Dead are included; when true, only Empty groups are migrated. **Type**: `bool` **Default**: `false` ### [](#headers)`headers` Custom headers to add to migrated records, keyed by header name with interpolated string values. Useful for injecting metadata such as processing timestamps or latency measurements that should surface as header values on the destination cluster. A custom header name that collides with `provenance_header` or `offset_header` is ignored, so those migration-critical headers are always protected. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` ```yaml # Examples: headers: x-migration-latency-ms: ${! timestamp_unix_milli() - meta("kafka_timestamp_ms") } x-migration-processed-at: ${! timestamp_unix_milli() } ``` ### [](#idempotent_write)`idempotent_write` Enable the idempotent write producer option. This requires the `IDEMPOTENT_WRITE` permission on `CLUSTER`. Disable this option if the `IDEMPOTENT_WRITE` permission is unavailable. **Type**: `bool` **Default**: `true` ### [](#max_buffered_bytes)`max_buffered_bytes` The maximum number of bytes the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered. Set to `0` to disable the byte-level limit (only `max_buffered_records` applies). This limit is checked after `max_buffered_records`. **Type**: `string` **Default**: `0` ```yaml # Examples: max_buffered_bytes: 256MB # --- max_buffered_bytes: 50mib ``` ### [](#max_buffered_records)`max_buffered_records` The maximum number of records the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered and space frees up. Increase this value for high-throughput pipelines to avoid back-pressure stalls. **Type**: `int` **Default**: `10000` ### [](#max_in_flight)`max_in_flight` The maximum number of batches to send in parallel at any given time. Increase this value to improve throughput during migration. For optimal performance, set this to match the total number of partitions being migrated. Setting it higher than the partition count provides no additional benefit, as each partition can only have one in-flight batch at a time. Example: If migrating 100 partitions, set `max_in_flight: 100` for maximum throughput. **Type**: `int` **Default**: `10` ```yaml # Examples: max_in_flight: 64 # For a cluster with 64 partitions # --- max_in_flight: 128 # For multiple topics with combined 128 partitions ``` ### [](#max_in_flight_requests)`max_in_flight_requests` The maximum number of produce requests in flight per broker connection. While `idempotent_write` is enabled (the default) this must be `1`, as the client relies on a single in-flight request per broker to guarantee ordering. To use higher values you must set `idempotent_write` to `false`, which allows requests to be pipelined for throughput but may cause duplicate and out-of-order delivery on retries. Note that this is distinct from the output’s `max_in_flight` field, which counts message batches being written in parallel rather than produce requests on the wire. A value of `1` is not a throughput ceiling: records from concurrent writes are coalesced into fewer, larger produce requests. **Type**: `int` **Default**: `1` ### [](#max_message_bytes)`max_message_bytes` The maximum space in bytes that an individual message may use. Messages larger than this value are rejected. This field corresponds to Kafka’s `max.message.bytes`. **Type**: `string` **Default**: `1MiB` ```yaml # Examples: max_message_bytes: 100MB # --- max_message_bytes: 50mib ``` ### [](#metadata_max_age)`metadata_max_age` The maximum period of time after which metadata is refreshed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. Lower values provide more responsive topic and partition discovery but may increase broker load. Higher values reduce broker queries but can delay detection of topology changes. **Type**: `string` **Default**: `1m` ### [](#offset_header)`offset_header` The name of a message header to add to migrated records. This header contains the source offset, enabling exact consumer group offset translation during migration. When left empty (default), no offset header is added and consumer groups are migrated using timestamp-based positioning. This approach works well for most cases, but may be imprecise for consumer groups with no committed offsets when multiple records share the same timestamp (timestamps have millisecond resolution). Set this field to enable precise offset translation, especially when migrating consumer groups that are caught up or have minimal lag. Note: This header is only added when consumer group migration is enabled. **Type**: `string` **Default**: `redpanda-migrator-offset` ### [](#partitioner)`partitioner` Override the default murmur2 hashing partitioner. **Type**: `string` | Option | Summary | | --- | --- | | least_backup | Chooses the least backed up partition (the partition with the fewest amount of buffered records). Partitions are selected per batch. | | manual | Manually select a partition for each message, requires the field partition to be specified. | | murmur2_hash | Kafka’s default hash algorithm that uses a 32-bit murmur2 hash of the key to compute which partition the record will be on. | | round_robin | Round-robin’s messages through all available partitions. This algorithm has lower throughput and causes higher CPU load on brokers, but can be useful if you want to ensure an even distribution of records to partitions. | ### [](#provenance_header)`provenance_header` Header name to add to migrated records indicating their source cluster. When set, each migrated message receives a header with this name containing the source cluster’s seed broker addresses, enabling downstream systems to track message origins for auditing, debugging, or multi-cluster orchestration workflows. If empty, no provenance header is added to messages. The header value format is a comma-separated list of the source cluster’s `seed_brokers`. Example: Setting `provenance_header: "rp-source-cluster"` adds a header like `rp-source-cluster: "kafka-1:9092,kafka-2:9092"`. **Type**: `string` **Default**: `redpanda-migrator-provenance` ### [](#record_delivery_timeout)`record_delivery_timeout` The maximum time a record can sit in the producer buffer before it is failed, roughly equivalent to Kafka’s `delivery.timeout.ms`. This is evaluated before writing a request or after a produce response. When a record times out, all records in the same partition are also failed. Set to `0s` for no timeout (the default). With `idempotent_write` enabled, timeouts are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `string` **Default**: `0s` ### [](#record_retries)`record_retries` The maximum number of times a record produce is retried on failure before the record is failed. When a record fails, all records buffered in the same partition are also failed to preserve gapless ordering. Set to `0` for unlimited retries (the default). With `idempotent_write` enabled, retries are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `int` **Default**: `0` ### [](#request_timeout_overhead)`request_timeout_overhead` Grants an additional buffer or overhead to requests that have timeout fields defined. This field is based on the behavior of Apache Kafka’s `request.timeout.ms` parameter, but with the option to extend the timeout deadline. **Type**: `string` **Default**: `10s` ### [](#sasl)`sasl[]` Specify one or more methods of SASL authentication, which are tried in order. If the broker supports the first mechanism, all connections will use that mechanism. If the first mechanism fails, the client picks the first supported mechanism. Connections fail if the broker does not support any client mechanisms. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM based authentication as specified in RFC5802. | | none | Disable sasl authentication | ### [](#sasl-password)`sasl[].password` A password to provide for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` A username to provide for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#schema_registry)`schema_registry` Configuration for schema registry integration. Enables migration of schema subjects, versions, and compatibility settings between clusters. **Type**: `object` ### [](#schema_registry-basic_auth)`schema_registry.basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#schema_registry-basic_auth-enabled)`schema_registry.basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-basic_auth-password)`schema_registry.basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-basic_auth-username)`schema_registry.basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#schema_registry-enabled)`schema_registry.enabled` Whether schema registry migration is enabled. When disabled, no schema operations are performed. **Type**: `bool` **Default**: `true` ### [](#schema_registry-exclude)`schema_registry.exclude[]` Regular expressions for schema subjects to exclude from migration. Takes precedence over include patterns. Note: the migrator consumer group is always ignored. **Type**: `array` ```yaml # Examples: exclude: [".*-test", ".*-temp"] # --- exclude: ["dev-.*", "local-.*"] ``` ### [](#schema_registry-include)`schema_registry.include[]` Regular expressions for schema subjects to include in migration. If empty, all subjects are included (unless excluded). Note: the migrator consumer group is always ignored. **Type**: `array` ```yaml # Examples: include: ["prod-.*", "staging-.*"] # --- include: ["user-.*", "order-.*"] ``` ### [](#schema_registry-include_deleted)`schema_registry.include_deleted` Whether to include soft-deleted schemas in migration. Useful for complete migration but may not be supported by all schema registries. **Type**: `bool` **Default**: `false` ### [](#schema_registry-interval)`schema_registry.interval` How often to synchronise schema registry subjects. Set to 0s for one-time sync at startup only. **Type**: `string` **Default**: `5m` ```yaml # Examples: interval: 0s # One-time sync only # --- interval: 5m # Sync every 5 minutes # --- interval: 30m # Sync every 30 minutes ``` ### [](#schema_registry-jwt)`schema_registry.jwt` (beta) Allows you to specify JWT authentication. **Type**: `object` ### [](#schema_registry-jwt-claims)`schema_registry.jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-enabled)`schema_registry.jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-jwt-headers)`schema_registry.jwt.headers` Add optional key/value headers to the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-private_key_file)`schema_registry.jwt.private_key_file` A file with the PEM encoded via PKCS1 or PKCS8 as private key. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt-signing_method)`schema_registry.jwt.signing_method` A method used to sign the token such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#schema_registry-max_parallel_http_requests)`schema_registry.max_parallel_http_requests` Maximum number of parallel HTTP requests to the schema registry. Controls concurrency when syncing multiple schemas. **Type**: `int` **Default**: `10` ### [](#schema_registry-normalize)`schema_registry.normalize` Whether to normalize schemas when creating them in the destination registry. **Type**: `bool` **Default**: `false` ### [](#schema_registry-oauth)`schema_registry.oauth` Allows you to specify open authentication via OAuth version 1. **Type**: `object` ### [](#schema_registry-oauth-access_token)`schema_registry.oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-access_token_secret)`schema_registry.oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_key)`schema_registry.oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_secret)`schema_registry.oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-enabled)`schema_registry.oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-strict)`schema_registry.strict` Error on unknown schema IDs. Only relevant when translate\_ids is true. When false (default), unknown schema IDs are passed through unchanged, allowing migration of topics with mixed message formats. Note: messages with 0-byte prefixes (e.g., protobuf) cannot be distinguished from schema registry headers and may fail when strict is enabled. **Type**: `bool` **Default**: `false` ### [](#schema_registry-subject)`schema_registry.subject` Template for transforming subject names during migration. Use interpolation to rename subjects systematically. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: subject: prod_${! metadata("schema_registry_subject") } # --- subject: ${! metadata("schema_registry_subject") | replace("dev_", "prod_") } ``` ### [](#schema_registry-timeout)`schema_registry.timeout` HTTP client timeout for schema registry requests. **Type**: `string` **Default**: `5s` ### [](#schema_registry-tls)`schema_registry.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#schema_registry-tls-client_certs)`schema_registry.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#schema_registry-tls-client_certs-cert)`schema_registry.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-cert_file)`schema_registry.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key)`schema_registry.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key_file)`schema_registry.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-password)`schema_registry.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#schema_registry-tls-enable_renegotiation)`schema_registry.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-enabled)`schema_registry.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-root_cas)`schema_registry.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#schema_registry-tls-root_cas_file)`schema_registry.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#schema_registry-tls-skip_cert_verify)`schema_registry.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#schema_registry-translate_ids)`schema_registry.translate_ids` Whether to translate schema IDs during migration. **Type**: `bool` **Default**: `false` ### [](#schema_registry-url)`schema_registry.url` The base URL of the schema registry service. Required for schema migration functionality. **Type**: `string` ```yaml # Examples: url: http://localhost:8081 # --- url: https://schema-registry.example.com:8081 ``` ### [](#schema_registry-versions)`schema_registry.versions` Which schema versions to migrate. 'latest' migrates only the current version, 'all' migrates complete version history for better compatibility. **Type**: `string` **Default**: `all` **Options**: `latest`, `all` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to. Use commas to separate multiple addresses in a single list item. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#serverless)`serverless` Enable serverless mode for Redpanda Cloud serverless clusters. This restricts topic configurations and schema features to those supported by serverless environments. **Type**: `bool` **Default**: `false` ### [](#sync_topic_acls)`sync_topic_acls` Whether to synchronise topic ACLs from source to destination cluster. ACLs are transformed safely: ALLOW WRITE permissions are excluded, and ALLOW ALL is downgraded to ALLOW READ to prevent conflicts. **Type**: `bool` **Default**: `false` ### [](#sync_topic_interval)`sync_topic_interval` How often to synchronize topics from the source cluster to the destination. This creates destination topics for any new source topics, including empty topics with no message flow. Set to 0s to disable periodic sync (topics are still created on first message). **Type**: `string` **Default**: `5m` ```yaml # Examples: sync_topic_interval: 0s # Disable periodic sync # --- sync_topic_interval: 1m # Sync every minute # --- sync_topic_interval: 5m # Sync every 5 minutes ``` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` The maximum period of time to wait for message sends before abandoning the request and retrying. **Type**: `string` **Default**: `10s` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic)`topic` A topic to write messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `${! @kafka_topic }` ```yaml # Examples: topic: prod_${! @kafka_topic } ``` ### [](#topic_replication_factor)`topic_replication_factor` The replication factor for created topics. If not specified, inherits the replication factor from source topics. Useful when migrating to clusters with different sizes. **Type**: `int` ```yaml # Examples: topic_replication_factor: 3 # --- topic_replication_factor: 1 # For single-node clusters ``` --- # Page 150: redpanda **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda.md --- # redpanda > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/redpanda page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/redpanda.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/redpanda.adoc page-git-created-date: "2024-11-19" page-git-modified-date: "2026-05-26" --- Sends message data to Kafka brokers and waits for acknowledgement before propagating any acknowledgements back to the input. #### Common ```yml outputs: label: "" redpanda: seed_brokers: [] # No default (optional) topic: "" # No default (required) key: "" # No default (optional) partition: "" # No default (optional) metadata: include_prefixes: [] include_patterns: [] max_in_flight: 256 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" redpanda: seed_brokers: [] # No default (optional) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s topic: "" # No default (required) key: "" # No default (optional) partition: "" # No default (optional) metadata: include_prefixes: [] include_patterns: [] timestamp_ms: "" # No default (optional) max_in_flight: 256 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) inject_tracing_map: "" # No default (optional) partitioner: "" # No default (optional) idempotent_write: true acks: all compression: "" # No default (optional) allow_auto_topic_creation: true timeout: 10s max_message_bytes: 1MiB broker_write_max_bytes: 100MiB max_buffered_records: 10000 max_buffered_bytes: 0 max_in_flight_requests: 1 record_retries: 0 record_delivery_timeout: 0s ``` ## [](#fields)Fields ### [](#acks)`acks` The number of acknowledgements the leader broker must receive from ISR brokers before responding to the produce request. When `idempotent_write` is enabled this must be set to `all`. **Type**: `string` **Default**: `all` | Option | Summary | | --- | --- | | all | Wait for all in-sync replicas to acknowledge (acks=-1). Required when idempotent_write is enabled. | | leader | Wait for the leader broker to acknowledge (acks=1). Messages are lost if the leader fails before replication. | | none | Do not wait for any acknowledgement (acks=0). Highest throughput but messages may be lost. | ### [](#allow_auto_topic_creation)`allow_auto_topic_creation` Enables topics to be auto created if they do not exist when fetching their metadata. **Type**: `bool` **Default**: `true` ### [](#batching)`batching` Optional explicit batching policy for the output. Note that when batches are formed at the input level they can be expanded by this policy, but not contracted. When consuming data from a Redpanda input it is recommended to tune batches from the input config via the `max_yield_batch_bytes` field, or the `unordered_processing.batching` field if appropriate. **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#broker_write_max_bytes)`broker_write_max_bytes` The maximum number of bytes this output can write to a broker connection in a single write. This field corresponds to Kafka’s `socket.request.max.bytes`. **Type**: `string` **Default**: `100MiB` ```yaml # Examples: broker_write_max_bytes: 128MB # --- broker_write_max_bytes: 50mib ``` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#compression)`compression` Set an explicit compression type (optional). The default preference is to use `snappy` when the broker supports it. Otherwise, use `none`. **Type**: `string` **Options**: `lz4`, `snappy`, `gzip`, `none`, `zstd` ### [](#conn_idle_timeout)`conn_idle_timeout` The maximum duration that connections can remain idle before they are automatically closed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `20s` ### [](#idempotent_write)`idempotent_write` Enable the idempotent write producer option. This requires the `IDEMPOTENT_WRITE` permission on `CLUSTER`. Disable this option if the `IDEMPOTENT_WRITE` permission is not available. **Type**: `bool` **Default**: `true` ### [](#inject_tracing_map)`inject_tracing_map` EXPERIMENTAL: A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) used to inject an object containing tracing propagation information into outbound messages. The specification of the injected fields will match the format used by the service wide tracer. **Type**: `string` ```yaml # Examples: inject_tracing_map: meta = @.merge(this) # --- inject_tracing_map: root.meta.span = this ``` ### [](#key)`key` An optional key to populate for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_buffered_bytes)`max_buffered_bytes` The maximum number of bytes the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered. Set to `0` to disable the byte-level limit (only `max_buffered_records` applies). This limit is checked after `max_buffered_records`. **Type**: `string` **Default**: `0` ```yaml # Examples: max_buffered_bytes: 256MB # --- max_buffered_bytes: 50mib ``` ### [](#max_buffered_records)`max_buffered_records` The maximum number of records the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered and space frees up. Increase this value for high-throughput pipelines to avoid back-pressure stalls. **Type**: `int` **Default**: `10000` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this number to improve throughput until performance plateaus. **Type**: `int` **Default**: `256` ### [](#max_in_flight_requests)`max_in_flight_requests` The maximum number of produce requests in flight per broker connection. While `idempotent_write` is enabled (the default) this must be `1`, as the client relies on a single in-flight request per broker to guarantee ordering. To use higher values you must set `idempotent_write` to `false`, which allows requests to be pipelined for throughput but may cause duplicate and out-of-order delivery on retries. Note that this is distinct from the output’s `max_in_flight` field, which counts message batches being written in parallel rather than produce requests on the wire. A value of `1` is not a throughput ceiling: records from concurrent writes are coalesced into fewer, larger produce requests. **Type**: `int` **Default**: `1` ### [](#max_message_bytes)`max_message_bytes` The maximum space (in bytes) that an individual message may use. Messages larger than this value are rejected. This field corresponds to Kafka’s `max.message.bytes`. **Type**: `string` **Default**: `1MiB` ```yaml # Examples: max_message_bytes: 100MB # --- max_message_bytes: 50mib ``` ### [](#metadata)`metadata` Configure which metadata values are added to messages as headers. This allows you to pass additional context information along with your messages. **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#metadata_max_age)`metadata_max_age` The maximum period of time after which metadata is refreshed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. Lower values provide more responsive topic and partition discovery but may increase broker load. Higher values reduce broker queries but can delay detection of topology changes. **Type**: `string` **Default**: `1m` ### [](#partition)`partition` Set a partition for each message (optional). This field is only relevant when the `partitioner` is set to `manual`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). You must provide an interpolation string that is a valid integer. **Type**: `string` ```yaml # Examples: partition: ${! meta("partition") } ``` ### [](#partitioner)`partitioner` Override the default murmur2 hashing partitioner. **Type**: `string` | Option | Summary | | --- | --- | | least_backup | Chooses the least backed up partition (the partition with the fewest amount of buffered records). Partitions are selected per batch. | | manual | Manually select a partition for each message, requires the field partition to be specified. | | murmur2_hash | Kafka’s default hash algorithm that uses a 32-bit murmur2 hash of the key to compute which partition the record will be on. | | round_robin | Round-robin’s messages through all available partitions. This algorithm has lower throughput and causes higher CPU load on brokers, but can be useful if you want to ensure an even distribution of records to partitions. | ### [](#record_delivery_timeout)`record_delivery_timeout` The maximum time a record can sit in the producer buffer before it is failed, roughly equivalent to Kafka’s `delivery.timeout.ms`. This is evaluated before writing a request or after a produce response. When a record times out, all records in the same partition are also failed. Set to `0s` for no timeout (the default). With `idempotent_write` enabled, timeouts are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `string` **Default**: `0s` ### [](#record_retries)`record_retries` The maximum number of times a record produce is retried on failure before the record is failed. When a record fails, all records buffered in the same partition are also failed to preserve gapless ordering. Set to `0` for unlimited retries (the default). With `idempotent_write` enabled, retries are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `int` **Default**: `0` ### [](#request_timeout_overhead)`request_timeout_overhead` Grants an additional buffer or overhead to requests that have timeout fields defined. This field is based on the behavior of Apache Kafka’s `request.timeout.ms` parameter, but with the option to extend the timeout deadline. **Type**: `string` **Default**: `10s` ### [](#sasl)`sasl[]` Specify one or more methods or mechanisms of SASL authentication, which are attempted in order. If the broker supports the first SASL mechanism, all connections use it. If the first mechanism fails, the client picks the first supported mechanism. If the broker does not support any client mechanisms, all connections fail. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM based authentication as specified in RFC5802. | | none | Disable sasl authentication | ### [](#sasl-password)`sasl[].password` A password to provide for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` A username to provide for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to in order. Use commas to separate multiple addresses in a single list item. Optional when `seed_brokers` is configured in a top-level `redpanda` block. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` The maximum period of time to wait for message sends before abandoning the request and retrying. **Type**: `string` **Default**: `10s` ### [](#timestamp_ms)`timestamp_ms` Set a timestamp (in milliseconds) for each message (optional). When left empty, the current timestamp is used. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: timestamp_ms: ${! timestamp_unix_milli() } # --- timestamp_ms: ${! metadata("kafka_timestamp_ms") } ``` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic)`topic` A topic to write messages to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 151: reject_errored **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/reject_errored.md --- # reject_errored > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: reject_errored latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/reject_errored page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/reject_errored.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/reject_errored.adoc description: Rejects messages that have failed their processing steps, resulting in nack behavior at the input level, otherwise sends them to a child output. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Rejects messages that have failed their processing steps, resulting in nack behavior at the input level, otherwise sends them to a child output. ```yml # Config fields, showing default values output: label: "" reject_errored: null # No default (required) ``` The routing of messages rejected by this output depends on the type of input it came from. For inputs that support propagating nacks upstream such as AMQP or NATS the message will be nacked. However, for inputs that are sequential such as files or Kafka the messages will simply be reprocessed from scratch. ## [](#examples)Examples ### [](#rejecting-failed-messages)Rejecting Failed Messages The most straight forward use case for this output type is to nack messages that have failed their processing steps. In this example our mapping might fail, in which case the messages that failed are rejected and will be nacked by our input: ```yaml input: nats_jetstream: urls: [ nats://127.0.0.1:4222 ] subject: foos.pending pipeline: processors: - mutation: 'root.age = this.fuzzy.age.int64()' output: reject_errored: nats_jetstream: urls: [ nats://127.0.0.1:4222 ] subject: foos.processed ``` ### [](#dlqing-failed-messages)DLQing Failed Messages Another use case for this output is to send failed messages straight into a dead-letter queue. You use it within a [fallback output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/) that allows you to specify where these failed messages should go to next. ```yaml pipeline: processors: - mutation: 'root.age = this.fuzzy.age.int64()' output: fallback: - reject_errored: http_client: url: http://foo:4195/post/might/become/unreachable retries: 3 retry_period: 1s - http_client: url: http://bar:4196/somewhere/else retries: 3 retry_period: 1s ``` --- # Page 152: reject **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/reject.md --- # reject > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: reject latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/reject page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/reject.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/reject.adoc description: Rejects all messages, treating them as though the output destination failed to publish them. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Rejects all messages, treating them as though the output destination failed to publish them. ```yml # Config fields, showing default values output: label: "" reject: "" ``` The routing of messages after this output depends on the type of input it came from. For inputs that support propagating nacks upstream such as AMQP or NATS the message will be nacked. However, for inputs that are sequential such as files or Kafka the messages will simply be reprocessed from scratch. To learn when this output could be useful, see \[the [Examples](#examples). ## [](#examples)Examples ### [](#rejecting-failed-messages)Rejecting Failed Messages This input is particularly useful for routing messages that have failed during processing, where instead of routing them to some sort of dead letter queue we wish to push the error upstream. We can do this with a switch broker: ```yaml output: switch: retry_until_success: false cases: - check: '!errored()' output: amqp_1: urls: [ amqps://guest:guest@localhost:5672/ ] target_address: queue:/the_foos - output: reject: "processing failed due to: ${! error() }" ``` --- # Page 153: resource **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/resource.md --- # resource > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: resource latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/resource page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/resource.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/resource.adoc description: Resource is an output type that channels messages to a resource output, identified by its name. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Resource is an output type that channels messages to a resource output, identified by its name. ```yml # Config fields, showing default values output: resource: "" ``` Resources allow you to tidy up deeply nested configs. For example, the config: ```yaml output: broker: pattern: fan_out outputs: - kafka: addresses: [ TODO ] topic: foo - gcp_pubsub: project: bar topic: baz ``` Could also be expressed as: ```yaml output: broker: pattern: fan_out outputs: - resource: foo - resource: bar output_resources: - label: foo kafka: addresses: [ TODO ] topic: foo - label: bar gcp_pubsub: project: bar topic: baz ``` --- # Page 154: retry **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/retry.md --- # retry > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: retry latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/retry page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/retry.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/retry.adoc description: Attempts to write messages to a child output and if the write fails for any reason the message is retried either until success or, if the retries or max elapsed time fields are non-zero, either is reached. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Attempts to write messages to a child output and if the write fails for any reason the message is retried either until success or, if the retries or max elapsed time fields are non-zero, either is reached. #### Common ```yml outputs: label: "" retry: output: "" # No default (required) ``` #### Advanced ```yml outputs: label: "" retry: max_retries: 0 backoff: initial_interval: 500ms max_interval: 3s max_elapsed_time: 0s output: "" # No default (required) ``` All messages in Redpanda Connect are always retried on an output error, but this would usually involve propagating the error back to the source of the message, whereby it would be reprocessed before reaching the output layer once again. This output type is useful whenever we wish to avoid reprocessing a message on the event of a failed send. We might, for example, have a deduplication processor that we want to avoid reapplying to the same message more than once in the pipeline. Rather than retrying the same output you may wish to retry the send using a different output target (a dead letter queue). In which case you should instead use the [`fallback`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/) output type. ## [](#fields)Fields ### [](#backoff)`backoff` Control time intervals between retry attempts. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` The initial period to wait between retry attempts. The retry interval increases for each failed attempt, up to the `backoff.max_interval` value. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `500ms` ### [](#backoff-max_elapsed_time)`backoff.max_elapsed_time` The maximum period to wait before retry attempts are abandoned. If zero then no limit is used. **Type**: `string` **Default**: `0s` ### [](#backoff-max_interval)`backoff.max_interval` The maximum period to wait between retry attempts. **Type**: `string` **Default**: `3s` ### [](#max_retries)`max_retries` The maximum number of retries before giving up on the request. If set to zero there is no discrete limit. **Type**: `int` **Default**: `0` ### [](#output)`output` A child output. **Type**: `output` --- # Page 155: salesforce_sink **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/salesforce_sink.md --- # salesforce_sink > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: salesforce_sink latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/salesforce_sink page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/salesforce_sink.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/salesforce_sink.adoc description: Writes messages to Salesforce, routing each Kafka topic to its own sObject configuration. page-git-created-date: "2026-05-01" page-git-modified-date: "2026-08-11" --- Writes messages to Salesforce, routing each Kafka topic to its own sObject configuration. Consumes batches of messages and writes them to Salesforce. Each message must have a `topic` field (set by the per-topic processor) and a `data` field containing the Salesforce record fields. The `topic` is used to look up the correct `topic_mappings` entry which defines the sObject, operation, and write mode. **Realtime mode** uses the sObject Collections REST API (synchronous, up to 200 records/call). **Bulk mode** uses the Bulk API 2.0 (asynchronous, polls until complete). #### Common ```yml outputs: label: "" salesforce_sink: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 bulk_batch_size: 1000 max_concurrent_bulk_jobs: 10 bulk_poll_interval: 5s batch_period: 5s max_in_flight: 1 topic_mappings: [] # No default (required) ``` #### Advanced ```yml outputs: label: "" salesforce_sink: org_url: "" # No default (required) client_id: "" # No default (required) client_secret: "" # No default (required) api_version: v65.0 bulk_batch_size: 1000 max_concurrent_bulk_jobs: 10 bulk_poll_interval: 5s batch_period: 5s max_in_flight: 1 topic_mappings: [] # No default (required) http: timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] proxy_url: "" disable_http2: false tps_limit: 0 tps_burst: 1 backoff: initial_interval: 1s max_interval: 30s max_retries: 3 tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s http: max_idle_conns: 100 max_idle_conns_per_host: 0 max_conns_per_host: 64 idle_conn_timeout: 1m30s tls_handshake_timeout: 10s expect_continue_timeout: 1s response_header_timeout: 0s disable_keep_alives: false disable_compression: false max_response_header_bytes: 1048576 max_response_body_bytes: 10485760 write_buffer_size: 4096 read_buffer_size: 4096 h2: strict_max_concurrent_requests: false max_decoder_header_table_size: 4096 max_encoder_header_table_size: 4096 max_read_frame_size: 16384 max_receive_buffer_per_connection: 1048576 max_receive_buffer_per_stream: 1048576 send_ping_timeout: 0s ping_timeout: 15s write_byte_timeout: 0s access_log_level: "" access_log_body_limit: 0 ``` ## [](#fields)Fields ### [](#api_version)`api_version` Salesforce REST API version to target, prefixed with `v`. Affects endpoint paths (`/services/data/{api_version}/…​`) and available fields/objects. Must be supported by your org — check Setup → Company Information. Older versions may lack recent fields. **Type**: `string` **Default**: `v65.0` ```yaml # Examples: api_version: v65.0 # --- api_version: v62.0 ``` ### [](#batch_period)`batch_period` Maximum period to wait before flushing an incomplete batch. **Type**: `string` **Default**: `5s` ### [](#bulk_batch_size)`bulk_batch_size` Number of records per bulk job. Also controls the output batch size. **Type**: `int` **Default**: `1000` ### [](#bulk_poll_interval)`bulk_poll_interval` How often to poll Salesforce for bulk job completion status. **Type**: `string` **Default**: `5s` ### [](#client_id)`client_id` Client ID for the Salesforce Connected App. **Type**: `string` ### [](#client_secret)`client_secret` Client secret for the Salesforce Connected App. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#http)`http` HTTP client configuration for Salesforce REST calls (OAuth token endpoint and, where applicable, data queries). **Type**: `object` ### [](#http-access_log_body_limit)`http.access_log_body_limit` Maximum bytes of request/response body to include in logs. 0 to skip body logging. **Type**: `int` **Default**: `0` ### [](#http-access_log_level)`http.access_log_level` Log level for HTTP request/response logging. Empty disables logging. **Type**: `string` **Default**: `""` **Options**: `` `, `TRACE ``, `DEBUG`, `INFO`, `WARN`, `ERROR` ### [](#http-backoff)`http.backoff` Adaptive backoff configuration for 429 (Too Many Requests) responses. Always active. **Type**: `object` ### [](#http-backoff-initial_interval)`http.backoff.initial_interval` Initial interval between retries on 429 responses. **Type**: `string` **Default**: `1s` ### [](#http-backoff-max_interval)`http.backoff.max_interval` Maximum interval between retries on 429 responses. **Type**: `string` **Default**: `30s` ### [](#http-backoff-max_retries)`http.backoff.max_retries` Maximum number of retries on 429 responses. **Type**: `int` **Default**: `3` ### [](#http-disable_http2)`http.disable_http2` Disable HTTP/2 and force HTTP/1.1. **Type**: `bool` **Default**: `false` ### [](#http-http)`http.http` HTTP transport settings controlling connection pooling, timeouts, and HTTP/2. **Type**: `object` ### [](#http-http-disable_compression)`http.http.disable_compression` Disable automatic decompression of gzip responses. **Type**: `bool` **Default**: `false` ### [](#http-http-disable_keep_alives)`http.http.disable_keep_alives` Disable HTTP keep-alive connections; each request uses a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-expect_continue_timeout)`http.http.expect_continue_timeout` Maximum time to wait for a server’s 100-continue response before sending the body. 0 means the body is sent immediately. **Type**: `string` **Default**: `1s` ### [](#http-http-h2)`http.http.h2` HTTP/2-specific transport settings. Only applied when HTTP/2 is enabled. **Type**: `object` ### [](#http-http-h2-max_decoder_header_table_size)`http.http.h2.max_decoder_header_table_size` Upper limit in bytes for the HPACK header table used to decode headers from the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_encoder_header_table_size)`http.http.h2.max_encoder_header_table_size` Upper limit in bytes for the HPACK header table used to encode headers sent to the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-http-h2-max_read_frame_size)`http.http.h2.max_read_frame_size` Largest HTTP/2 frame this endpoint will read. Valid range: 16 KiB to 16 MiB. **Type**: `int` **Default**: `16384` ### [](#http-http-h2-max_receive_buffer_per_connection)`http.http.h2.max_receive_buffer_per_connection` Maximum flow-control window size in bytes for data received on a connection. Must be at least 64 KiB and less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-max_receive_buffer_per_stream)`http.http.h2.max_receive_buffer_per_stream` Maximum flow-control window size in bytes for data received on a single stream. Must be less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-http-h2-ping_timeout)`http.http.h2.ping_timeout` Timeout waiting for a PING response before closing the connection. **Type**: `string` **Default**: `15s` ### [](#http-http-h2-send_ping_timeout)`http.http.h2.send_ping_timeout` Idle timeout after which a PING frame is sent to verify connection health. 0 disables health checks. **Type**: `string` **Default**: `0s` ### [](#http-http-h2-strict_max_concurrent_requests)`http.http.h2.strict_max_concurrent_requests` When true, new requests block when a connection’s concurrency limit is reached instead of opening a new connection. **Type**: `bool` **Default**: `false` ### [](#http-http-h2-write_byte_timeout)`http.http.h2.write_byte_timeout` Timeout for writing data to a connection. The timer resets whenever bytes are written. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-idle_conn_timeout)`http.http.idle_conn_timeout` How long an idle connection remains in the pool before being closed. 0 disables the timeout. **Type**: `string` **Default**: `1m30s` ### [](#http-http-max_conns_per_host)`http.http.max_conns_per_host` Maximum total connections (active + idle) per host. 0 means unlimited. **Type**: `int` **Default**: `64` ### [](#http-http-max_idle_conns)`http.http.max_idle_conns` Maximum total number of idle (keep-alive) connections across all hosts. 0 means unlimited. **Type**: `int` **Default**: `100` ### [](#http-http-max_idle_conns_per_host)`http.http.max_idle_conns_per_host` Maximum idle connections to keep per host. 0 (the default) uses GOMAXPROCS+1. **Type**: `int` **Default**: `0` ### [](#http-http-max_response_body_bytes)`http.http.max_response_body_bytes` Maximum bytes of response body the client will read. The response body is wrapped with a limit reader; reads beyond this cap return EOF. 0 disables the limit. **Type**: `int` **Default**: `10485760` ### [](#http-http-max_response_header_bytes)`http.http.max_response_header_bytes` Maximum bytes of response headers to allow. **Type**: `int` **Default**: `1048576` ### [](#http-http-read_buffer_size)`http.http.read_buffer_size` Size in bytes of the per-connection read buffer. **Type**: `int` **Default**: `4096` ### [](#http-http-response_header_timeout)`http.http.response_header_timeout` Maximum time to wait for response headers after writing the full request. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-http-tls_handshake_timeout)`http.http.tls_handshake_timeout` Maximum time to wait for a TLS handshake to complete. 0 disables the timeout. **Type**: `string` **Default**: `10s` ### [](#http-http-write_buffer_size)`http.http.write_buffer_size` Size in bytes of the per-connection write buffer. **Type**: `int` **Default**: `4096` ### [](#http-proxy_url)`http.proxy_url` HTTP proxy URL. Empty string disables proxying. **Type**: `string` **Default**: `""` ### [](#http-tcp)`http.tcp` TCP socket configuration. **Type**: `object` ### [](#http-tcp-connect_timeout)`http.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-tcp-keep_alive)`http.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#http-tcp-keep_alive-count)`http.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#http-tcp-keep_alive-idle)`http.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#http-tcp-keep_alive-interval)`http.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#http-tcp-tcp_user_timeout)`http.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#http-timeout)`http.timeout` HTTP request timeout. **Type**: `string` **Default**: `5s` ### [](#http-tls)`http.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#http-tls-client_certs)`http.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#http-tls-client_certs-cert)`http.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-cert_file)`http.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key)`http.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-key_file)`http.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#http-tls-client_certs-password)`http.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#http-tls-enable_renegotiation)`http.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#http-tls-enabled)`http.tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#http-tls-root_cas)`http.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#http-tls-root_cas_file)`http.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#http-tls-skip_cert_verify)`http.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#http-tps_burst)`http.tps_burst` Maximum burst size for rate limiting. **Type**: `int` **Default**: `1` ### [](#http-tps_limit)`http.tps_limit` Rate limit in requests per second. 0 disables rate limiting. **Type**: `float` **Default**: `0` ### [](#max_concurrent_bulk_jobs)`max_concurrent_bulk_jobs` Maximum number of bulk jobs polling concurrently in the background. Each in-flight job buffers its CSV payload in memory. Lower this value if memory usage is a concern. **Type**: `int` **Default**: `10` ### [](#max_in_flight)`max_in_flight` Maximum number of batches to send concurrently. Increasing this value improves real-time write throughput. **Type**: `int` **Default**: `1` ### [](#org_url)`org_url` Salesforce instance base URL (for example, [https://your-domain.salesforce.com](https://your-domain.salesforce.com)). **Type**: `string` ```yaml # Examples: org_url: https://acme.my.salesforce.com # --- org_url: https://acme--staging.sandbox.my.salesforce.com ``` ### [](#topic_mappings)`topic_mappings[]` Per-topic Salesforce write configuration. Each entry maps a Kafka topic to an sObject and write settings. **Type**: `array` ### [](#topic_mappings-all_or_none)`topic_mappings[].all_or_none` Real-time only: rolls back the entire batch if any record fails. **Type**: `bool` **Default**: `false` ### [](#topic_mappings-external_id_field)`topic_mappings[].external_id_field` External ID field name. Required for upsert operations. **Type**: `string` **Default**: `""` ### [](#topic_mappings-mode)`topic_mappings[].mode` Write mode: `realtime` (sObject Collections API) or `bulk` (Bulk API 2.0). **Type**: `string` **Default**: `realtime` ### [](#topic_mappings-operation)`topic_mappings[].operation` Write operation: insert, update, upsert, or delete. **Type**: `string` **Default**: `upsert` ### [](#topic_mappings-sobject)`topic_mappings[].sobject` Salesforce sObject API name (for example, Account, Contact, MyObject\_\_c). **Type**: `string` ### [](#topic_mappings-topic)`topic_mappings[].topic` Kafka topic name to match against the message’s `topic` field. **Type**: `string` --- # Page 156: schema_registry **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/schema_registry.md --- # schema_registry > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: schema_registry latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/schema_registry page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/schema_registry.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/schema_registry.adoc description: Publishes schemas to SchemaRegistry. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Publishes schemas to a schema registry. This output uses the [Franz Kafka Schema Registry client](https://github.com/twmb/franz-go/tree/master/pkg/sr). #### Common ```yml outputs: label: "" schema_registry: url: "" # No default (required) subject: "" # No default (required) max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" schema_registry: url: "" # No default (required) subject: "" # No default (required) subject_compatibility_level: "" # No default (optional) backfill_dependencies: true translate_ids: false normalize: true remove_metadata: true remove_rule_set: true input_resource: schema_registry_input tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] max_in_flight: 64 oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} ``` ## [](#performance)Performance The `schema_registry` output sends multiple messages in parallel for improved performance. You can use the `max_in_flight` field to tune the maximum number of in-flight messages, or message batches. ## [](#example)Example This example writes schemas to a schema registry instance and logs errors for existing schemas. ```yaml output: fallback: - schema_registry: url: http://localhost:8082 subject: ${! @schema_registry_subject } - switch: cases: - check: '@fallback_error == "request returned status: 422"' output: drop: {} processors: - log: message: | Subject '${! @schema_registry_subject }' version ${! @schema_registry_version } already has schema: ${! content() } - output: reject: ${! @fallback_error } ``` ## [](#fields)Fields ### [](#backfill_dependencies)`backfill_dependencies` Backfill missing schema references and previous schema versions. If set to `true`, you must also configure a [`schema_registry`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/schema_registry/) input to read source schemas. **Type**: `bool` **Default**: `true` ### [](#basic_auth)`basic_auth` Configure basic authentication for requests from this component to your schema registry. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` The username of the account credentials to authenticate as. Used together with `password` for basic authentication. **Type**: `string` **Default**: `""` ### [](#input_resource)`input_resource` The label of the [`schema_registry` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/schema_registry/) from which to read source schemas. **Type**: `string` **Default**: `schema_registry_input` ### [](#jwt)`jwt` (beta) Configure JSON Web Token (JWT) authentication for secure data transmission from this component to your schema registry. This feature is in beta and may change in future releases. **Type**: `object` ### [](#jwt-claims)`jwt.claims` Values used to pass the identity of the authenticated entity to the service provider. In this case, between this component and the schema registry. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` The key/value pairs that identify the type of token and signing algorithm. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` A PEM-encoded file containing a private key that is formatted using either PKCS1 or PKCS8 standards. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` The method used to sign the token, such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this number to improve throughput. **Type**: `int` **Default**: `64` ### [](#normalize)`normalize` Normalize schemas. **Type**: `bool` **Default**: `true` ### [](#oauth)`oauth` Configure OAuth version 1.0 to give this component authorized access to your schema registry. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` The value this component can use to gain access to the schema registry. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` The secret that establishes ownership of the `oauth.access_token` in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` The value used to identify this component or client to your schema registry. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` The secret that establishes ownership of the consumer key in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#remove_metadata)`remove_metadata` Removes metadata fields from schema output. Use this to produce leaner schema definitions for downstream consumers or when metadata is not required. **Type**: `bool` **Default**: `true` ### [](#remove_rule_set)`remove_rule_set` Removes rule set definitions from schema output. Useful for simplifying schemas when rule sets are not required by consumers or applications. **Type**: `bool` **Default**: `true` ### [](#subject)`subject` The subject name. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#subject_compatibility_level)`subject_compatibility_level` The compatibility level for the subject. Can be one of `BACKWARD`, `BACKWARD_TRANSITIVE`, `FORWARD`, `FORWARD_TRANSITIVE`, `FULL`, `FULL_TRANSITIVE`, `NONE`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#translate_ids)`translate_ids` When set to `true`, this field automatically translates the schema ID in each message to match the corresponding schema in the destination schema registry. The updated message is then written to the destination schema registry. **Type**: `bool` **Default**: `false` ### [](#url)`url` The base URL of the schema registry service. **Type**: `string` --- # Page 157: sftp **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/sftp.md --- # sftp > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sftp latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/sftp page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/sftp.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/sftp.adoc description: Writes files to an SFTP server. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Writes files to an SFTP server. #### Common ```yml outputs: label: "" sftp: address: "" # No default (required) credentials: username: "" password: "" host_public_key_file: "" # No default (optional) host_public_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key: "" # No default (optional) private_key_pass: "" path: "" # No default (required) codec: all-bytes max_in_flight: 64 ``` #### Advanced ```yml outputs: label: "" sftp: address: "" # No default (required) connection_timeout: 30s credentials: username: "" password: "" host_public_key_file: "" # No default (optional) host_public_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key: "" # No default (optional) private_key_pass: "" path: "" # No default (required) codec: all-bytes max_in_flight: 64 ``` In order to have a different path for each object you should use function interpolations described [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. ## [](#fields)Fields ### [](#address)`address` The address (hostname or IP address) of the SFTP server to connect to. **Type**: `string` ### [](#codec)`codec` The way in which the bytes of messages should be written out into the output data stream. It’s possible to write lines using a custom delimiter with the `delim:x` codec, where x is the character sequence custom delimiter. **Type**: `string` **Default**: `all-bytes` | Option | Summary | | --- | --- | | all-bytes | Only applicable to file based outputs. Writes each message to a file in full, if the file already exists the old content is deleted. | | append | Append each message to the output stream without any delimiter or special encoding. | | delim:x | Append each message to the output stream followed by a custom delimiter. | | lines | Append each message to the output stream followed by a line break. | ```yaml # Examples: codec: lines # --- codec: delim: # --- codec: delim:foobar ``` ### [](#connection_timeout)`connection_timeout` The connection timeout to use when connecting to the target server. **Type**: `string` **Default**: `30s` ### [](#credentials)`credentials` The credentials required to log in to the SFTP server. This can include a username and password, or a private key for secure access. **Type**: `object` ### [](#credentials-host_public_key)`credentials.host_public_key` The raw contents of the SFTP server’s public key, used for host key verification. **Type**: `string` ### [](#credentials-host_public_key_file)`credentials.host_public_key_file` The path to the SFTP server’s public key file, used for host key verification. **Type**: `string` ### [](#credentials-password)`credentials.password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#credentials-private_key)`credentials.private_key` The private key used to authenticate with the SFTP server. This field provides an alternative to the [`private_key_file`](#credentials-private_key_file). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-private_key_file)`credentials.private_key_file` The path to a private key file used to authenticate with the SFTP server. You can also provide a private key using the [`private_key`](#credentials-private_key) field. **Type**: `string` ### [](#credentials-private_key_pass)`credentials.private_key_pass` A passphrase for private key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#credentials-username)`credentials.username` The username required to authenticate with the SFTP server. **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#path)`path` The file to save the messages to on the SFTP server. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 158: slack_post **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/slack_post.md --- # slack_post > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: slack_post latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/slack_post page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/slack_post.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/slack_post.adoc page-git-created-date: "2025-05-02" page-git-modified-date: "2026-05-26" --- Posts a new message to a Slack channel using the Slack API method [chat.postMessage](https://api.slack.com/methods/chat.postMessage). ```yml # Common configuration fields, showing default values output: label: "" slack_post: bot_token: "" # No default (required) channel_id: "" # No default (required) thread_ts: "" # No default (optional) text: "" # No default (optional) blocks: "" # No default (optional) markdown: true unfurl_links: false unfurl_media: true link_names: 0 ``` See also: [Examples](#examples) ## [](#fields)Fields ### [](#blocks)`blocks` A Bloblang query that should return a JSON array of [Slack blocks](https://api.slack.com/reference/block-kit/blocks). You can either specify message content in the `text` or `blocks` fields, but not both. **Type**: `string` ### [](#bot_token)`bot_token` Your Slack bot user’s OAuth token, which must have the correct permissions to post messages to the target Slack channel. **Type**: `string` ### [](#channel_id)`channel_id` The encoded ID of the target Slack channel. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#link_names)`link_names` When set to `1`, this output finds and links to [user groups](https://api.slack.com/reference/surfaces/formatting#mentioning-groups) mentioned in Slack messages. **Type**: `bool` **Default**: `false` ### [](#markdown)`markdown` When set to `true`, this output accepts message content in Markdown format. **Type**: `bool` **Default**: `true` ### [](#text)`text` The text content of the message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). You can either specify message content in the `text` or `blocks` fields, but not both. **Type**: `string` **Default**: `""` ### [](#thread_ts)`thread_ts` Specify the thread timestamp (`ts` value) of another message to post a reply within the same thread. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#unfurl_links)`unfurl_links` When set to `true`, this output provides previews of linked content in Slack messages. For more information about unfurling links, see the [Slack documentation](https://api.slack.com/reference/messaging/link-unfurling). **Type**: `bool` **Default**: `false` ### [](#unfurl_media)`unfurl_media` When set to `true`, this output provides previews of rich content in Slack messages, such as videos or embedded tweets. **Type**: `bool` **Default**: `true` ## [](#examples)Examples ### [](#echo-slackbot)Echo Slackbot A slackbot that echo messages from other users ```yaml input: slack: app_token: "${APP_TOKEN:xapp-demo}" bot_token: "${BOT_TOKEN:xoxb-demo}" pipeline: processors: - mutation: | # ignore hidden or non message events if this.event.type != "message" || (this.event.hidden | false) { root = deleted() } # Don't respond to our own messages if this.authorizations.any(auth -> auth.user_id == this.event.user) { root = deleted() } output: slack_post: bot_token: "${BOT_TOKEN:xoxb-demo}" channel_id: "${!this.event.channel}" thread_ts: "${!this.event.ts}" text: "ECHO: ${!this.event.text}" ``` --- # Page 159: slack_reaction **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/slack_reaction.md --- # slack_reaction > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: slack_reaction latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/slack_reaction page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/slack_reaction.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/slack_reaction.adoc description: Add or remove an emoji reaction to a Slack message. page-git-created-date: "2025-07-08" page-git-modified-date: "2026-08-11" --- Add or remove an emoji reaction to a Slack message using [`reactions.add`](https://api.slack.com/methods/reactions.add) and [`reactions.remove`](https://api.slack.com/methods/reactions.remove). ```yaml output: label: "" slack_reaction: bot_token: "" # No default (required) channel_id: "" # No default (required) timestamp: "" # No default (required) emoji: "" # No default (required) action: add max_in_flight: 64 ``` ## [](#fields)Fields ### [](#action)`action` Whether to add or remove the reaction. When set to `add`, the specified emoji reaction is applied to the target message. When set to `remove`, the emoji reaction is removed from the target message. **Type**: `string` **Default**: `add` **Options**: `add`, `remove` ### [](#bot_token)`bot_token` Your Slack Bot User OAuth token used to authenticate the API request. This token must have the necessary `reactions:write` and `channels:read` (or related) scopes. It typically begins with `xoxb-`. **Type**: `string` ### [](#channel_id)`channel_id` The unique Slack channel ID where the target message resides. Channel IDs usually start with `C` for public channels or `G` for private channels. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#emoji)`emoji` The name of the emoji to be added or removed, without surrounding colons. Use the plain emoji name, such as `thumbsup` or `tada`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increasing this value can improve throughput in high-volume scenarios, but be cautious not to exceed Slack’s API rate limits. **Type**: `int` **Default**: `64` ### [](#timestamp)`timestamp` The timestamp of the message to react to. This is a unique identifier for the message, usually obtained from a previous Slack API call (such as `chat.postMessage` or `conversations.history`). It typically looks like a Unix timestamp with a decimal. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 160: snowflake_put **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/snowflake_put.md --- # snowflake_put > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: snowflake_put latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/snowflake_put page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/snowflake_put.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/snowflake_put.adoc description: Sends messages to Snowflake stages and, optionally, calls Snowpipe to load this data into one or more tables. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > 💡 **TIP** > > Use the [`snowflake_streaming` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/snowflake_streaming/) for improved performance, cost-effectiveness, and ease of use. Sends messages to Snowflake stages and, optionally, calls Snowpipe to load this data into one or more tables. #### Common ```yml outputs: label: "" snowflake_put: account: "" # No default (required) region: "" # No default (optional) cloud: "" # No default (optional) user: "" # No default (required) password: "" # No default (optional) private_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key_pass: "" # No default (optional) role: "" # No default (required) database: "" # No default (required) warehouse: "" # No default (required) schema: "" # No default (required) stage: "" # No default (required) path: "" file_name: "" file_extension: "" compression: AUTO request_id: "" snowpipe: "" # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 1 ``` #### Advanced ```yml outputs: label: "" snowflake_put: account: "" # No default (required) region: "" # No default (optional) cloud: "" # No default (optional) user: "" # No default (required) password: "" # No default (optional) private_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key_pass: "" # No default (optional) role: "" # No default (required) database: "" # No default (required) warehouse: "" # No default (required) schema: "" # No default (required) stage: "" # No default (required) path: "" file_name: "" file_extension: "" upload_parallel_threads: 4 compression: AUTO request_id: "" snowpipe: "" # No default (optional) client_session_keep_alive: false batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 1 ``` In order to use a different stage and / or Snowpipe for each message, you can use function interpolations as described in [Bloblang queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). When using batching, messages are grouped by the calculated stage and Snowpipe and are streamed to individual files in their corresponding stage and, optionally, a Snowpipe `insertFiles` REST API call will be made for each individual file. ## [](#credentials)Credentials Two authentication mechanisms are supported: - User/password - Key Pair Authentication ### [](#userpassword)User/password This is a basic authentication mechanism which allows you to PUT data into a stage. However, it is not compatible with Snowpipe. ### [](#key-pair-authentication)Key pair authentication This authentication mechanism allows Snowpipe functionality, but it does require configuring an SSH Private Key beforehand. Please consult the [documentation](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication) for details on how to set it up and assign the Public Key to your user. Note that the Snowflake documentation [used to suggest](https://twitter.com/felipehoffa/status/1560811785606684672) using this command: ```bash openssl genrsa 2048 | openssl pkcs8 -topk8 -inform PEM -out rsa_key.p8 ``` to generate an encrypted SSH private key. However, in this case, it uses an encryption algorithm called `pbeWithMD5AndDES-CBC`, which is part of the PKCS#5 v1.5 and is considered insecure. Due to this, Redpanda Connect does not support it and, if you wish to use password-protected keys directly, you must use PKCS#5 v2.0 to encrypt them by using the following command (as the current Snowflake docs suggest): ```bash openssl genrsa 2048 | openssl pkcs8 -topk8 -v2 des3 -inform PEM -out rsa_key.p8 ``` If you have an existing key encrypted with PKCS#5 v1.5, you can re-encrypt it with PKCS#5 v2.0 using this command: ```bash openssl pkcs8 -in rsa_key_original.p8 -topk8 -v2 des3 -out rsa_key.p8 ``` Please consult the [pkcs8 command documentation](https://linux.die.net/man/1/pkcs8) for details on PKCS#5 algorithms. ## [](#batching)Batching It’s common to want to upload messages to Snowflake as batched archives. The easiest way to do this is to batch your messages at the output level and join the batch of messages with an [`archive`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive/) and/or [`compress`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/compress/) processor. For the optimal batch size, please consult the Snowflake [documentation](https://docs.snowflake.com/en/user-guide/data-load-considerations-prepare.html). ## [](#snowpipe)Snowpipe Given a table called `BENTHOS_TBL` with one column of type `variant`: ```sql CREATE OR REPLACE TABLE BENTHOS_DB.PUBLIC.BENTHOS_TBL(RECORD variant) ``` and the following `BENTHOS_PIPE` Snowpipe: ```sql CREATE OR REPLACE PIPE BENTHOS_DB.PUBLIC.BENTHOS_PIPE AUTO_INGEST = FALSE AS COPY INTO BENTHOS_DB.PUBLIC.BENTHOS_TBL FROM (SELECT * FROM @%BENTHOS_TBL) FILE_FORMAT = (TYPE = JSON COMPRESSION = AUTO) ``` you can configure Redpanda Connect to use the implicit table stage `@%BENTHOS_TBL` as the `stage` and `BENTHOS_PIPE` as the `snowpipe`. In this case, you must set `compression` to `AUTO` and, if using message batching, you’ll need to configure an [`archive`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive/) processor with the `concatenate` format. Since the `compression` is set to `AUTO`, the [gosnowflake](https://github.com/snowflakedb/gosnowflake) client library will compress the messages automatically so you don’t need to add a [`compress`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/compress/) processor for message batches. If you add `STRIP_OUTER_ARRAY = TRUE` in your Snowpipe `FILE_FORMAT` definition, then you must use `json_array` instead of `concatenate` as the archive processor format. > 📝 **NOTE** > > Only Snowpipes with `FILE_FORMAT` `TYPE` `JSON` are currently supported. ## [](#snowpipe-troubleshooting)Snowpipe troubleshooting Snowpipe [provides](https://docs.snowflake.com/en/user-guide/data-load-snowpipe-rest-apis.html) the `insertReport` and `loadHistoryScan` REST API endpoints which can be used to get information about recent Snowpipe calls. In order to query them, you’ll first need to generate a valid JWT token for your Snowflake account. There are two methods for doing so: - Using the `snowsql` [utility](https://docs.snowflake.com/en/user-guide/snowsql.html): ```bash snowsql --private-key-path rsa_key.p8 --generate-jwt -a -u ``` - Using the Python `sql-api-generate-jwt` [utility](https://docs.snowflake.com/en/developer-guide/sql-api/authenticating.html#generating-a-jwt-in-python): ```bash python3 sql-api-generate-jwt.py --private_key_file_path=rsa_key.p8 --account= --user= ``` Once you successfully generate a JWT token and store it into the `JWT_TOKEN` environment variable, then you can, for example, query the `insertReport` endpoint using `curl`: ```bash curl -H "Authorization: Bearer ${JWT_TOKEN}" "https://.snowflakecomputing.com/v1/data/pipes/../insertReport" ``` If you need to pass in a valid `requestId` to any of these Snowpipe REST API endpoints, you can set a [uuid\_v4()](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#uuid_v4) string in a metadata field called `request_id`, log it via the [`log`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/log/) processor and then configure `request_id: ${ @request_id }` ). Alternatively, you can [enable debug logging](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/logger/about/) and Redpanda Connect will print the Request IDs that it sends to Snowpipe. ## [](#general-troubleshooting)General troubleshooting The underlying [`gosnowflake` driver](https://github.com/snowflakedb/gosnowflake) requires write access to the default directory to use for temporary files. Please consult the [`os.TempDir`](https://pkg.go.dev/os#TempDir) docs for details on how to change this directory via environment variables. A silent failure can occur due to [this issue](https://github.com/snowflakedb/gosnowflake/issues/701), where the underlying [`gosnowflake` driver](https://github.com/snowflakedb/gosnowflake) doesn’t return an error and doesn’t log a failure if it can’t figure out the current username. One way to trigger this behavior is by running Redpanda Connect in a Docker container with a non-existent user ID (such as `--user 1000:1000`). ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#examples)Examples ### [](#kafka-realtime-brokers)Kafka / realtime brokers Upload message batches from realtime brokers such as Kafka persisting the batch partition and offsets in the stage path and filename similarly to the [Kafka Connector scheme](https://docs.snowflake.com/en/user-guide/kafka-connector-ts.html#step-1-view-the-copy-history-for-the-table) and call Snowpipe to load them into a table. When batching is configured at the input level, it is done per-partition. ```yaml input: redpanda: seed_brokers: - localhost:9092 topics: - foo consumer_group: rpcn max_yield_batch_bytes: 8MB processors: - mapping: | meta kafka_start_offset = meta("kafka_offset").from(0) meta kafka_end_offset = meta("kafka_offset").from(-1) meta batch_timestamp = if batch_index() == 0 { now() } - mapping: | meta batch_timestamp = if batch_index() != 0 { meta("batch_timestamp").from(0) } output: snowflake_put: account: benthos user: test@benthos.dev private_key_file: path_to_ssh_key.pem role: ACCOUNTADMIN database: BENTHOS_DB warehouse: COMPUTE_WH schema: PUBLIC stage: "@%BENTHOS_TBL" path: benthos/BENTHOS_TBL/${! @kafka_partition } file_name: ${! @kafka_start_offset }_${! @kafka_end_offset }_${! meta("batch_timestamp") } upload_parallel_threads: 4 compression: NONE snowpipe: BENTHOS_PIPE ``` ### [](#no-compression)No compression Upload concatenated messages into a `.json` file to a table stage without calling Snowpipe. ```yaml output: snowflake_put: account: benthos user: test@benthos.dev private_key_file: path_to_ssh_key.pem role: ACCOUNTADMIN database: BENTHOS_DB warehouse: COMPUTE_WH schema: PUBLIC stage: "@%BENTHOS_TBL" path: benthos upload_parallel_threads: 4 compression: NONE batching: count: 10 period: 3s processors: - archive: format: concatenate ``` ### [](#parquet-format-with-snappy-compression)Parquet format with snappy compression Upload concatenated messages into a `.parquet` file to a table stage without calling Snowpipe. ```yaml output: snowflake_put: account: benthos user: test@benthos.dev private_key_file: path_to_ssh_key.pem role: ACCOUNTADMIN database: BENTHOS_DB warehouse: COMPUTE_WH schema: PUBLIC stage: "@%BENTHOS_TBL" path: benthos file_extension: parquet upload_parallel_threads: 4 compression: NONE batching: count: 10 period: 3s processors: - parquet_encode: schema: - name: ID type: INT64 - name: CONTENT type: BYTE_ARRAY default_compression: snappy ``` ### [](#automatic-compression)Automatic compression Upload concatenated messages compressed automatically into a `.gz` archive file to a table stage without calling Snowpipe. ```yaml output: snowflake_put: account: benthos user: test@benthos.dev private_key_file: path_to_ssh_key.pem role: ACCOUNTADMIN database: BENTHOS_DB warehouse: COMPUTE_WH schema: PUBLIC stage: "@%BENTHOS_TBL" path: benthos upload_parallel_threads: 4 compression: AUTO batching: count: 10 period: 3s processors: - archive: format: concatenate ``` ### [](#deflate-compression)DEFLATE compression Upload concatenated messages compressed into a `.deflate` archive file to a table stage and call Snowpipe to load them into a table. ```yaml output: snowflake_put: account: benthos user: test@benthos.dev private_key_file: path_to_ssh_key.pem role: ACCOUNTADMIN database: BENTHOS_DB warehouse: COMPUTE_WH schema: PUBLIC stage: "@%BENTHOS_TBL" path: benthos upload_parallel_threads: 4 compression: DEFLATE snowpipe: BENTHOS_PIPE batching: count: 10 period: 3s processors: - archive: format: concatenate - mapping: | root = content().compress("zlib") ``` ### [](#raw_deflate-compression)RAW_DEFLATE compression Upload concatenated messages compressed into a `.raw_deflate` archive file to a table stage and call Snowpipe to load them into a table. ```yaml output: snowflake_put: account: benthos user: test@benthos.dev private_key_file: path_to_ssh_key.pem role: ACCOUNTADMIN database: BENTHOS_DB warehouse: COMPUTE_WH schema: PUBLIC stage: "@%BENTHOS_TBL" path: benthos upload_parallel_threads: 4 compression: RAW_DEFLATE snowpipe: BENTHOS_PIPE batching: count: 10 period: 3s processors: - archive: format: concatenate - mapping: | root = content().compress("flate") ``` ## [](#fields)Fields ### [](#account)`account` Account name, which is the same as the [Account Identifier](https://docs.snowflake.com/en/user-guide/admin-account-identifier.html#where-are-account-identifiers-used). However, when using an [Account Locator](https://docs.snowflake.com/en/user-guide/admin-account-identifier.html#using-an-account-locator-as-an-identifier), the Account Identifier is formatted as `..` and this field needs to be populated using the `` part. **Type**: `string` ### [](#batching-2)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#client_session_keep_alive)`client_session_keep_alive` Enable Snowflake keepalive mechanism to prevent the client session from expiring after 4 hours (error 390114). **Type**: `bool` **Default**: `false` ### [](#cloud)`cloud` Optional cloud platform field which needs to be populated when using an [Account Locator](https://docs.snowflake.com/en/user-guide/admin-account-identifier.html#using-an-account-locator-as-an-identifier) and it must be set to the `` part of the Account Identifier (`..`). **Type**: `string` ```yaml # Examples: cloud: aws # --- cloud: gcp # --- cloud: azure ``` ### [](#compression)`compression` Compression type. **Type**: `string` **Default**: `AUTO` | Option | Summary | | --- | --- | | AUTO | Compression (gzip) is applied automatically by the output and messages must contain plain-text JSON. Default file_extension: gz. | | DEFLATE | Messages must be pre-compressed using the zlib algorithm (with zlib header, RFC1950). Default file_extension: deflate. | | GZIP | Messages must be pre-compressed using the gzip algorithm. Default file_extension: gz. | | NONE | No compression is applied and messages must contain plain-text JSON. Default file_extension: json. | | RAW_DEFLATE | Messages must be pre-compressed using the flate algorithm (without header, RFC1951). Default file_extension: raw_deflate. | | ZSTD | Messages must be pre-compressed using the Zstandard algorithm. Default file_extension: zst. | ### [](#database)`database` Database. **Type**: `string` ### [](#file_extension)`file_extension` Stage file extension. Will be derived from the configured `compression` if not set or empty. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: file_extension: csv # --- file_extension: parquet ``` ### [](#file_name)`file_name` Stage file name. Will be equal to the Request ID if not set or empty. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#max_in_flight)`max_in_flight` The maximum number of parallel message batches to have in flight at any given time. **Type**: `int` **Default**: `1` ### [](#password)`password` An optional password. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#path)`path` Stage path. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#private_key)`private_key` Your private SSH key. When using encrypted keys, you must also set a value for [`private_key_pass`](#private_key_pass). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#private_key_file)`private_key_file` The path to a file containing your private SSH key. When using encrypted keys, you must also set a value for [`private_key_pass`](#private_key_pass). **Type**: `string` ### [](#private_key_pass)`private_key_pass` The passphrase for your private SSH key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#region)`region` Optional region field which needs to be populated when using an [Account Locator](https://docs.snowflake.com/en/user-guide/admin-account-identifier.html#using-an-account-locator-as-an-identifier) and it must be set to the `` part of the Account Identifier (`..`). **Type**: `string` ```yaml # Examples: region: us-west-2 ``` ### [](#request_id)`request_id` Request ID. Will be assigned a random UUID (v4) string if not set or empty. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#role)`role` Role. **Type**: `string` ### [](#schema)`schema` Schema. **Type**: `string` ### [](#snowpipe-2)`snowpipe` An optional Snowpipe name. Use the `` part from `..`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#stage)`stage` Stage name. Use either one of the [supported](https://docs.snowflake.com/en/user-guide/data-load-local-file-system-create-stage.html) stage types. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#upload_parallel_threads)`upload_parallel_threads` Specifies the number of threads to use for uploading files. **Type**: `int` **Default**: `4` ### [](#user)`user` Username. **Type**: `string` ### [](#warehouse)`warehouse` Warehouse. **Type**: `string` --- # Page 161: snowflake_streaming **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/snowflake_streaming.md --- # snowflake_streaming > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: snowflake_streaming latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/snowflake_streaming page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/snowflake_streaming.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/snowflake_streaming.adoc page-git-created-date: "2024-11-19" page-git-modified-date: "2026-05-26" --- Allows Snowflake to ingest data from your data pipeline using [Snowpipe Streaming](https://docs.snowflake.com/en/user-guide/data-load-snowpipe-streaming-overview). To help you configure your own `snowflake_streaming` output, this page includes [example data pipelines](#example-pipelines). ### Common ```yml outputs: label: "" snowflake_streaming: account: "" # No default (required) user: "" # No default (required) role: "" # No default (required) database: "" # No default (required) schema: "" # No default (required) table: "" # No default (required) private_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key_pass: "" # No default (optional) mapping: "" # No default (optional) init_statement: "" # No default (optional) schema_evolution: enabled: "" # No default (required) ignore_nulls: true processors: [] # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 4 ``` ### Advanced ```yml outputs: label: "" snowflake_streaming: account: "" # No default (required) url: "" # No default (optional) user: "" # No default (required) role: "" # No default (required) database: "" # No default (required) schema: "" # No default (required) table: "" # No default (required) private_key: "" # No default (optional) private_key_file: "" # No default (optional) private_key_pass: "" # No default (optional) mapping: "" # No default (optional) init_statement: "" # No default (optional) schema_evolution: enabled: "" # No default (required) ignore_nulls: true processors: [] # No default (optional) build_options: parallelism: 1 chunk_size: 50000 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) max_in_flight: 4 channel_prefix: "" # No default (optional) channel_name: "" # No default (optional) offset_token: "" # No default (optional) commit_backoff: initial_interval: 32ms max_interval: 512ms max_elapsed_time: 60s multiplier: 2 message_format: object timestamp_format: 2006-01-02T15:04:05.999999999Z07:00 ``` ## [](#conversion-of-message-data-into-snowflake-table-rows)Conversion of message data into Snowflake table rows Message data conversion to Snowflake table rows is determined by the: - Output message contents. - [Schema evolution settings](#schema_evolution). - Schema of the [target Snowflake table](#table). The following scenarios highlight how these three factors affect data written to the target table. > 📝 **NOTE** > > For reduced complexity, consider [turning on schema evolution](#schema_evolution), which automatically creates and updates the Snowflake table schema based on message contents. ### [](#scenario-data-and-table-schema-match-schema-evolution-turned-on-or-off)Scenario: Data and table schema match (schema evolution turned on or off) An output message matches the existing table schema, and the `schema_evolution.enabled` field is set to `true` or `false`. The target Snowflake table has two columns: - `product_id` (NUMBER) - `product_code` (STRING) A pipeline generates the following message: ```json {"product_id": 521, "product_code": “EST-PR”} ``` In this scenario: - The JSON keys in the message (`"product_id"` and `"product_code"`) match column names in the target Snowflake table. - The message values match the column data types. (If there was a data mismatch, the message would be rejected.) - Redpanda Connect inserts the message values into a new row in the target Snowflake table. | product_id | product_code | | --- | --- | | 521 | EST-PR | ### [](#scenario-data-and-table-schema-mismatch-schema-evolution-turned-on)Scenario: Data and table schema mismatch (schema evolution turned on) An output message includes schema updates, and the `schema_evolution.enabled` field is set to `true`. The target Snowflake table has the same two columns as the [previous scenario](#scenario-data-and-table-schema-match-schema-evolution-turned-on-or-off): - `product_id` (NUMBER) - `product_code` (STRING) This time, the pipeline generates the following message: ```json {"product_batch": 11111, "product_color": “yellow”} ``` In this scenario: - The JSON keys (`"product_batch"` and `"product_color"`) do not match column names in the target Snowflake table. - As schema evolution is enabled, Redpanda Connect adds two new columns to the target table with data types derived from the output message values. For more information about the mapping of data types, see [Supported data formats for Snowflake columns](#supported-data-formats-for-snowflake-columns). - Redpanda Connect inserts the message values into a new table row. | product_id | product_code | product_batch | product_color | | --- | --- | --- | --- | | (null) | (null) | 11111 | yellow | > 📝 **NOTE** > > You can [configure processors](#schema_evolution-processors) to override the schema updates derived from the message values. ### [](#scenario-data-and-table-schema-mismatch-schema-evolution-turned-off)Scenario: Data and table schema mismatch (schema evolution turned off) An output message includes schema updates, and the `schema_evolution.enabled` field is set to `false`. The target Snowflake table has the same two columns: - `product_id` (NUMBER) - `product_code` (STRING) The pipeline generates the same message as the [previous scenario](#scenario-data-and-table-schema-mismatch-schema-evolution-turned-on): ```json {"product_batch": 11111, "product_color": “yellow”} ``` In this scenario: - The JSON keys (`"product_batch"` and `"product_color"`) do not match any existing column names. - Because schema evolution is turned off, Redpanda Connect ignores the extra column names and values and inserts a row of null values. | product_id | product_code | | --- | --- | | (null) | (null) | ## [](#supported-data-formats-for-snowflake-columns)Supported data formats for Snowflake columns The message data from your output must match the columns in the Snowflake table that you want to write data to. The following table shows you the [column data types supported by Snowflake](https://docs.snowflake.com/en/sql-reference/intro-summary-data-types) and how they correspond to the [Bloblang data types](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#type) in Redpanda Connect. | Snowflake column data type | Bloblang data types | | --- | --- | | CHAR, VARCHAR | string | | BINARY | string or bytes | | NUMBER | number, or string where the string is parsed into a number | | FLOAT, including special values, such as NaN (Not a Number), -inf (negative infinity), and inf (positive infinity) | number | | BOOLEAN | bool, or number where a non-zero number is true | | TIME, DATE, TIMESTAMP | timestamp, or number where the number is a converted to a Unix timestamp, or string where the string is parsed using RFC 3339 format | | VARIANT, ARRAY, OBJECT | Any data type converted into JSON | | GEOGRAPHY,GEOMETRY | Not supported | ## [](#authentication)Authentication You can authenticate with Snowflake using an [RSA key pair](https://docs.snowflake.com/en/user-guide/key-pair-auth). Either specify: - A PEM-encoded private key, in the [`private_key` field](#private_key). - The path to a file from which the output can load the private RSA key, in the [`private_key_file` field](#private_key_file). ## [](#performance)Performance For improved performance, this output: - Sends multiple messages in parallel. You can tune the maximum number of in-flight messages (or message batches) with the field `max_in_flight`. - Sends messages as a batch. You can configure batches at both the input and output level. For more information, see [Message Batching](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ### [](#batch-sizes)Batch sizes Redpanda recommends that every message batch writes at least 16 MiB of compressed output to Snowflake. You can monitor batch sizes using the `snowflake_compressed_output_size_bytes` metric. ### [](#metrics)Metrics This output emits the following metrics. | Metric name | Description | | --- | --- | | snowflake_compressed_output_size_bytes | The size in bytes of each message batch uploaded to Snowflake. | | snowflake_convert_latency_ns | The time taken to convert messages into the Snowflake column data types. | | snowflake_serialize_latency_ns | The time taken to serialize the converted columnar data into a file for upload to Snowflake. | | snowflake_build_output_latency_ns | The time taken to build the file that is uploaded to Snowflake. This metric is the sum of snowflake_convert_latency_ns + snowflake_serialize_latency_ns. | | snowflake_upload_latency_ns | The time taken to upload the output file to Snowflake. | | snowflake_register_latency_ns | The time taken to register the uploaded output file with Snowflake. | | snowflake_commit_latency_ns | The time taken to commit the uploaded data updates to the target Snowflake table. | ## [](#fields)Fields ### [](#account)`account` The [Snowflake account name to use](https://docs.snowflake.com/en/user-guide/admin-account-identifier#account-name). Use the format `-` where: - The `` is the name of your Snowflake organization. - The `` is the unique name of your account with your Snowflake organization. To find the correct value for this field, run the following query in Snowflake: ```sql WITH HOSTLIST AS (SELECT * FROM TABLE(FLATTEN(INPUT => PARSE_JSON(SYSTEM$allowlist())))) SELECT REPLACE(VALUE:host,'.snowflakecomputing.com','') AS ACCOUNT_IDENTIFIER FROM HOSTLIST WHERE VALUE:type = 'SNOWFLAKE_DEPLOYMENT_REGIONLESS'; ``` **Type**: `string` ```yaml # Examples: account: ORG-ACCOUNT ``` ### [](#batching)`batching` Lets you configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). Type\*: `object` ```yml # Examples batching: byte_size: 5000 count: 0 period: 1s batching: count: 10 period: 1s batching: check: this.contains("END BATCH") count: 0 period: 1m ``` **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` The number of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. All resulting messages are flushed as a single batch, and therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#build_options)`build_options` Options for optimizing the build of the output data that is sent to Snowflake. Monitor the `snowflake_build_output_latency_ns` metric to assess whether you need to update these options. **Type**: `object` ### [](#build_options-chunk_size)`build_options.chunk_size` The number of table rows to submit in each chunk for processing. **Type**: `int` **Default**: `50000` ### [](#build_options-parallelism)`build_options.parallelism` The maximum amount of parallel processing to use when building the output for Snowflake. **Type**: `int` **Default**: `1` ### [](#channel_name)`channel_name` The channel name to use when connecting to a Snowflake table. Duplicate channel names cause errors and prevent multiple instances of Redpanda Connect from writing at the same time, and so this field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). Redpanda Connect assumes that a message batch contains messages for a single channel, which means that interpolation is only executed on the first message in each batch. If your pipeline uses an input that is partitioned, such as an Apache Kafka topic, batch messages at the input level to make sure all messages are processed by the same channel. You can specify either the `channel_name` or `channel_prefix`, but not both. If neither field is populated, this output creates a channel name based on a table’s fully-qualified name, which results in a single stream per table. > 📝 **NOTE** > > Snowflake limits the number of streams per table to 10,000. If you need to use more than 10,000 streams, contact [Snowflake support](https://www.snowflake.com/en/support/). **Type**: `string` ```yaml # Examples: channel_name: partition-${!@kafka_partition} ``` ### [](#channel_prefix)`channel_prefix` The prefix to use when creating a channel name for connecting to a Snowflake table. Adding a `channel_prefix` avoids the creation of duplicate channel names, which result in errors and prevent multiple instances of Redpanda Connect from writing at the same time. You can specify either the `channel_prefix` or `channel_name`, but not both. If neither field is populated, this output creates a channel name based on a table’s fully-qualified name, which results in a single stream per table. The maximum number of channels open at any time is determined by the value in the `max_in_flight` field. > 📝 **NOTE** > > Snowflake limits the number of streams per table to 10,000. If you need to use more than 10,000 streams, contact [Snowflake support](https://www.snowflake.com/en/support/). **Type**: `string` ```yaml # Examples: channel_prefix: channel-${HOST} ``` ### [](#commit_backoff)`commit_backoff` Control how frequently Snowflake is polled to check if data has been committed. **Type**: `object` ### [](#commit_backoff-initial_interval)`commit_backoff.initial_interval` The initial period to wait between status polls. **Type**: `string` **Default**: `32ms` ### [](#commit_backoff-max_elapsed_time)`commit_backoff.max_elapsed_time` The maximum total time to wait for data to be committed. If zero then no limit is used. **Type**: `string` **Default**: `60s` ### [](#commit_backoff-max_interval)`commit_backoff.max_interval` The maximum period to wait between status polls. **Type**: `string` **Default**: `512ms` ### [](#commit_backoff-multiplier)`commit_backoff.multiplier` The factor by which the poll interval grows on each attempt. **Type**: `float` **Default**: `2` ### [](#database)`database` The Snowflake database you want to write data to. **Type**: `string` ```yaml # Examples: database: MY_DATABASE ``` ### [](#init_statement)`init_statement` Optional SQL statements to execute immediately after this output connects to Snowflake for the first time. This is a useful way to initialize tables before processing data. > 📝 **NOTE** > > Make sure your SQL statements are idempotent, so they do not cause issues when run multiple times after service restarts. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS mytable (amount NUMBER); # --- init_statement: |- ALTER TABLE t1 ALTER COLUMN c1 DROP NOT NULL; ALTER TABLE t1 ADD COLUMN a2 NUMBER; ``` ### [](#mapping)`mapping` The [Bloblang `mapping`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) to execute on each message. **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this number to improve throughput until performance plateaus. **Type**: `int` **Default**: `4` ### [](#message_format)`message_format` The format to expect incoming messages from the rest of the pipeline. **Type**: `string` **Default**: `object` | Option | Summary | | --- | --- | | array | Messages are an array of values where each position matches the ordinal of the column in Snowflake. | | object | Messages are JSON or Bloblang objects where each key is the Snowflake column name and the value is the column value. | ```yaml # Examples: message_format: array ``` ### [](#offset_token)`offset_token` The offset token to use for exactly-once delivery of data to a Snowflake table. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). This output assumes that messages within a batch are in increasing order by offset token. When data is sent on a channel, the offset token of each message in the batch is compared to the latest token processed by the channel. If the offset token is lexicographically less than the latest token, it’s assumed the message is a duplicate and is dropped. Messages must be delivered to the output in order, otherwise they are processed as duplicates and dropped. To avoid dropping retried messages if later messages have succeeded in the meantime, use a dead-letter queue to process failed messages. See the [Ingesting data exactly once from Redpanda](#example-pipelines) example. > 📝 **NOTE** > > If you’re using a numeric value as an offset token, pad the value so that it’s lexicographically ordered in its string representation because offset tokens are compared in string form. For more details, see the [Ingesting data exactly once from Redpanda](#example-pipelines) example. For more information about offset tokens, see [Snowflake Documentation](https://docs.snowflake.com/en/user-guide/data-load-snowpipe-streaming-overview#offset-tokens). **Type**: `string` ```yaml # Examples: offset_token: offset-${!"%016X".format(@kafka_offset)} # --- offset_token: postgres-${!@lsn} ``` ### [](#private_key)`private_key` The PEM-encoded private RSA key to use for authentication with Snowflake. You must specify a value for this field or the `private_key_file` field. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#private_key_file)`private_key_file` A `.p8`, PEM-encoded file to load the private RSA key from. You must specify a value for this field or the `private_key` field. **Type**: `string` ### [](#private_key_pass)`private_key_pass` If the RSA key is encrypted, specify the RSA key passphrase. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#role)`role` The role of the user specified in the `user` field. The user’s role must have the [required privileges](https://docs.snowflake.com/en/user-guide/data-load-snowpipe-streaming-overview#required-access-privileges) to call the Snowpipe Streaming APIs. For more information about user roles, see the [Snowflake documentation](https://docs.snowflake.com/en/user-guide/admin-user-management#user-roles). **Type**: `string` ```yaml # Examples: role: ACCOUNTADMIN ``` ### [](#schema)`schema` The schema of the Snowflake database you want to write data to. **Type**: `string` ```yaml # Examples: schema: PUBLIC ``` ### [](#schema_evolution)`schema_evolution` Options to control schema updates when messages are written to the Snowflake table. **Type**: `object` ### [](#schema_evolution-enabled)`schema_evolution.enabled` Whether schema evolution is enabled. When set to `true`, the Snowflake table is automatically created based on the schema of the first message written to it, if the table does not already exist. As new fields are added to subsequent messages in the pipeline, new columns are created in the Snowflake table. Any required columns are marked as `nullable` if new messages do not include data for them. **Type**: `bool` ### [](#schema_evolution-ignore_nulls)`schema_evolution.ignore_nulls` When set to `true` and schema evolution is enabled, new columns that have `null` values _are not_ added to the Snowflake table. This behavior: - Prevents unnecessary schema changes caused by placeholder or incomplete data. - Avoids creating table columns with incorrect data types. > 📝 **NOTE** > > Redpanda does not recommend updating the default setting unless you are confident about the data type of `null` columns in advance. **Type**: `bool` **Default**: `true` ### [](#schema_evolution-processors)`schema_evolution.processors[]` A series of processors to execute when new columns are added to the Snowflake table. You can use these processors to: - Run side effects when the schema evolves. - Enrich the message with additional information to guide the schema changes. For example, a processor could read the schema from the schema registry that a message was produced with and use that schema to determine the data type of the new column in Snowflake. The input to these processors is an object with the value and name of the new column, the original message, and details of the Snowflake table the output writes to. For example: `{"value": 42.3, "name":"new_data_field", "message": {"existing_data_field": 42, "new_data_field": "db_field_name"}, "db": MY_DATABASE", "schema": "MY_SCHEMA", "table": "MY_TABLE"}` The output from the processors must be a valid message, which contains a string that specifies the column type for the new column in Snowflake. The metadata remains the same as in the original message that triggered the schema update. **Type**: `array` ```yaml # Examples: processors: - mapping: |- root = match this.value.type() { this == "string" => "STRING" this == "bytes" => "BINARY" this == "number" => "DOUBLE" this == "bool" => "BOOLEAN" this == "timestamp" => "TIMESTAMP" _ => "VARIANT" } ``` ### [](#table)`table` The Snowflake table you want to write data to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: table: MY_TABLE ``` ### [](#timestamp_format)`timestamp_format` The format to parse string values for `TIMESTAMP`, `TIMESTAMP_LTZ` and `TIMESTAMP_NTZ` columns. Should be a layout for [time.Parse](https://pkg.go.dev/time#Parse) in Go. **Type**: `string` **Default**: `2006-01-02T15:04:05.999999999Z07:00` ### [](#url)`url` Specify a custom URL to connect to Snowflake. This parameter overrides the default URL, which is automatically generated from the value of `output.snowflake_streaming.account`. By default, the URL is constructed as follows: `[https://.snowflakecomputing.com](https://\.snowflakecomputing.com)`. **Type**: `string` ```yaml # Examples: url: https://org-account.privatelink.snowflakecomputing.com ``` ### [](#user)`user` Specify a user to run the Snowpipe Stream. To learn how to create a user, see the [Snowflake documentation](https://docs.snowflake.com/en/user-guide/admin-user-management). **Type**: `string` ## [](#example-pipelines)Example pipelines The following examples show you how to ingest, process, and write data to Snowflake from: - A PostgreSQL table using change data capture (CDC) - A Redpanda cluster - A REST API that posts JSON payloads to a HTTP server See also: [Ingest data into Snowflake cookbook](https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/snowflake_ingestion/) ### Write data exactly once to a Snowflake table using CDC Send data from a PostgreSQL table and write it to Snowflake exactly once using PostgreSQL logical replication. This example includes some important features: - To make sure that a Snowflake streaming channel does not assume that older data is already committed, the configuration sets a 45-second interval between message batches. This interval prevents a message batch from being sent while another batch is retried. - The log sequence number of each data update from the Write-Ahead Log (WAL) in PostgreSQL makes sure that data is only uploaded once to the `snowflake_streaming` output, and that messages sent to the output are already lexicographically ordered. > 📝 **NOTE** > > To do exactly-once data delivery, it’s important that records are delivered in order to the output, and are correctly partitioned. Before you start, read the [`offset_token`](#offset_token) field description. Alternatively, remove the `offset_token` field to use Redpanda Connect’s default at-least-once delivery model. ```yaml input: postgres_cdc: dsn: postgres://foouser:foopass@localhost:5432/foodb schema: "public" tables: ["my_pg_table"] # Use very large batches. Each batch is sent to Snowflake individually, # so to optimize query performance, use the largest file size # your memory allows batching: count: 50000 period: 45s # Set an interval between message batches to prevent multiple batches # from being in flight at once checkpoint_limit: 1 output: snowflake_streaming: # Using the log sequence number makes sure data is only updated exactly once offset_token: "${!@lsn}" # Sending a single ordered log means you can only send one update # at a time and properly increment the offset_token # and use only a single channel. max_in_flight: 1 account: "MYSNOW-ACCOUNT" user: MYUSER role: ACCOUNTADMIN database: "MYDATABASE" schema: "PUBLIC" table: "MY_PG_TABLE" private_key_file: "my/private/key.p8" ``` ### Ingest data exactly once from Redpanda Ingest data from Redpanda using consumer groups, decode the schema using the schema registry, then write the corresponding data into Snowflake. This example includes some important features: - To create multiple Redpanda Connect streams to write to each output table, you need a unique channel prefix per stream. The `channel_prefix` field constructs a unique prefix for each stream using the host name. - To prevent message failures from being retried and changing the order of delivered messages, a dead-letter queue processes them. > 📝 **NOTE** > > To do exactly-once data delivery, it’s important that records are delivered in order to the output, and are correctly partitioned. Before you start, read the [`channel_name`](#channel_name) and [`offset_token`](#offset_token) field descriptions. Alternatively, remove the `offset_token` field to use Redpanda Connect’s default at-least-once delivery model. ```yaml input: redpanda_common: topics: ["my_topic_going_to_snow"] consumer_group: "redpanda_connect_to_snowflake" # Use very large batches. Each batch is sent to Snowflake individually, # so to optimize query performance, use the largest file size # your memory allows fetch_max_bytes: 100MiB fetch_min_bytes: 50MiB partition_buffer_bytes: 100MiB pipeline: processors: - schema_registry_decode: url: "redpanda.example.com:8081" basic_auth: enabled: true username: MY_USER_NAME password: "${TODO}" output: fallback: - snowflake_streaming: # To write an ordered stream of messages, each partition in # Apache Kafka gets its own channel. channel_name: "partition-${!@kafka_partition}" # Offsets are lexicographically sorted in string form by padding with # leading zeros offset_token: offset-${!"%016X".format(@kafka_offset)} account: "MYSNOW-ACCOUNT" user: MYUSER role: ACCOUNTADMIN database: "MYDATABASE" schema: "PUBLIC" table: "MYTABLE" private_key_file: "my/private/key.p8" schema_evolution: enabled: true # To prevent delivery failures from changing the order of # delivered records, it's important that they are immediately # sent to a dead-letter queue. - retry: output: redpanda_common: topic: "dead_letter_queue" ``` ### HTTP server to push data to Snowflake Create a HTTP server input that receives HTTP PUT requests with JSON payloads. The payloads are buffered locally then written to Snowflake in batches. To create multiple Redpanda Connect streams to write to each output table, you need a unique channel prefix per stream. In this example, the `channel_prefix` field constructs a unique prefix for each stream using the host name. > 📝 **NOTE** > > Using a buffer to immediately respond to the HTTP requests may result in data loss if there are delivery failures between the output and Snowflake. For more information about the configuration of buffers, see [buffers](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/buffers/memory/). Alternatively, remove the buffer entirely to respond to the HTTP request only once the data is written to Snowflake. ```yaml input: http_server: path: /snowflake buffer: memory: # Max inflight data before applying backpressure limit: 524288000 # 50MiB # Batching policy the size of the files sent to Snowflake batch_policy: enabled: true byte_size: 33554432 # 32MiB period: "10s" output: snowflake_streaming: account: "MYSNOW-ACCOUNT" user: MYUSER role: ACCOUNTADMIN database: "MYDATABASE" schema: "PUBLIC" table: "MYTABLE" private_key_file: "my/private/key.p8" channel_prefix: "snowflake-channel-for-${HOST}" schema_evolution: enabled: true ``` --- # Page 162: splunk_hec **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/splunk_hec.md --- # splunk_hec > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: splunk_hec latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/splunk_hec page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/splunk_hec.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/splunk_hec.adoc description: Publishes messages to a Splunk HTTP Endpoint Collector (HEC). page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Publishes messages to a Splunk HTTP Endpoint Collector (HEC). #### Common ```yml outputs: label: "" splunk_hec: url: "" # No default (required) token: "" # No default (required) gzip: false event_host: "" # No default (optional) event_source: "" # No default (optional) event_sourcetype: "" # No default (optional) event_index: "" # No default (optional) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" splunk_hec: url: "" # No default (required) token: "" # No default (required) gzip: false event_host: "" # No default (optional) event_source: "" # No default (optional) event_sourcetype: "" # No default (optional) event_index: "" # No default (optional) tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#performance)Performance This output benefits from sending multiple messages in flight in parallel for improved performance. You can tune the max number of in flight messages (or message batches) with the field `max_in_flight`. This output benefits from sending messages as a batch for improved performance. Batches can be formed at both the input and output level. You can find out more [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#event_host)`event_host` Set the host value to assign to the event data. Overrides existing host field if present. **Type**: `string` ### [](#event_index)`event_index` Set the index value to assign to the event data. Overrides existing index field if present. **Type**: `string` ### [](#event_source)`event_source` Set the source value to assign to the event data. Overrides existing source field if present. **Type**: `string` ### [](#event_sourcetype)`event_sourcetype` Set the sourcetype value to assign to the event data. Overrides existing sourcetype field if present. **Type**: `string` ### [](#gzip)`gzip` Enable gzip compression **Type**: `bool` **Default**: `false` ### [](#max_in_flight)`max_in_flight` The maximum number of messages to have in flight at a given time. Increase this to improve throughput. **Type**: `int` **Default**: `64` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#token)`token` A bot token used for authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#url)`url` Full HTTP Endpoint Collector (HEC) URL. **Type**: `string` ```yaml # Examples: url: https://foobar.splunkcloud.com/services/collector/event ``` --- # Page 163: sql_insert **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/sql_insert.md --- # sql_insert > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql_insert latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/sql_insert page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/sql_insert.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/sql_insert.adoc description: Inserts a row into an SQL database for each message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Inserts a row into an SQL database for each message. #### Common ```yml outputs: label: "" sql_insert: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) args_mapping: "" # No default (required) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" sql_insert: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) args_mapping: "" # No default (required) prefix: "" # No default (optional) suffix: "" # No default (optional) options: [] # No default (optional) max_in_flight: 64 init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` ## [](#examples)Examples ### [](#table-insert-mysql)Table Insert (MySQL) Here we insert rows into a database by populating the columns id, name and topic with values extracted from messages and metadata: ```yaml output: sql_insert: driver: mysql dsn: foouser:foopassword@tcp(localhost:3306)/foodb table: footable columns: [ id, name, topic ] args_mapping: | root = [ this.user.id, this.user.name, meta("kafka_topic"), ] ``` ## [](#dynamic-sql-operations)Dynamic SQL operations The `table` and `columns` fields are static strings that do not support Bloblang interpolation. For dynamic table names, dynamic column lists, DELETE operations, or any other SQL that `sql_insert` cannot express, use the [`sql_raw` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/sql_raw/) instead. There is no dedicated `sql_delete` output. To delete rows, use `sql_raw` with a DELETE statement: ```yaml output: sql_raw: driver: postgres dsn: postgres://user:pass@localhost:5432/mydb?sslmode=disable query: "DELETE FROM my_table WHERE id = $1" args_mapping: root = [ this.id ] ``` To insert into a table determined at runtime, use `sql_raw` with `unsafe_dynamic_query: true`, which enables Bloblang interpolation in the `query` field. > ⚠️ **CAUTION** > > Interpolating user-supplied values into a query can introduce SQL injection risks. Always validate or sanitize the interpolated value beforehand. ```yaml output: sql_raw: driver: postgres dsn: postgres://user:pass@localhost:5432/mydb?sslmode=disable unsafe_dynamic_query: true query: 'INSERT INTO ${! this.table_name } (id, value) VALUES ($1, $2)' args_mapping: root = [ this.id, this.value ] ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of columns specified. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#columns)`columns[]` A list of columns to insert. **Type**: `array` ```yaml # Examples: columns: - foo - bar - baz ``` ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#max_in_flight)`max_in_flight` The maximum number of inserts to run in parallel. **Type**: `int` **Default**: `64` ### [](#options)`options[]` A list of keyword options to add before the INTO clause of the query. **Type**: `array` ```yaml # Examples: options: - DELAYED - IGNORE ``` ### [](#prefix)`prefix` An optional prefix to prepend to the insert query (before INSERT). **Type**: `string` ### [](#suffix)`suffix` An optional suffix to append to the insert query. **Type**: `string` ```yaml # Examples: suffix: ON CONFLICT (name) DO NOTHING ``` ### [](#table)`table` The table to insert to. **Type**: `string` ```yaml # Examples: table: foo ``` --- # Page 164: sql_raw **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/sql_raw.md --- # sql_raw > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql_raw latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/sql_raw page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/sql_raw.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/sql_raw.adoc description: Executes an arbitrary SQL query for each message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes an arbitrary SQL query for each message. #### Common ```yml outputs: label: "" sql_raw: driver: "" # No default (required) dsn: "" # No default (required) query: "" # No default (optional) args_mapping: "" # No default (optional) queries: [] # No default (optional) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` #### Advanced ```yml outputs: label: "" sql_raw: driver: "" # No default (required) dsn: "" # No default (required) query: "" # No default (optional) unsafe_dynamic_query: false args_mapping: "" # No default (optional) queries: [] # No default (optional) max_in_flight: 64 init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` For some scenarios where you might use this output, see [Examples](#examples). ## [](#batch-execution-and-ordered-writes)Batch execution and ordered writes Configure a `batching` policy to accumulate messages and write them together. When you set more than one query (see [Conditional queries](#conditional-queries)), all messages in a batch run in a single database transaction, which reduces round-trips to the database. A single-query configuration runs each message individually without a transaction, preserving per-message error granularity. When consuming from Redpanda or Kafka, this output orders messages by partition within the transaction. This lets you raise `max_in_flight` above `1` to parallelize writes across partitions while preserving consume order within each partition, so you can scale throughput without breaking change-data-capture ordering. Messages without a `kafka_partition` metadata field are treated as partition `0`. This makes `sql_raw` suitable for high-throughput sink pipelines, including CDC, that previously required `max_in_flight: 1` to stay ordered. ## [](#conditional-queries)Conditional queries Use the `queries` field to route each message to a different SQL statement based on its content or metadata, without a preprocessing `mapping` step or `unsafe_dynamic_query`. Set a `when` condition (a Bloblang expression) on each entry: the first query whose `when` evaluates to `true`, or the first query with no `when`, runs for that message. If none of your queries have a `when` condition, every query runs for each message within the same transaction. Because an unconditioned query always matches, Redpanda Connect lints against configuring one before later queries in the list. For example, route change-data-capture tombstones to a `DELETE` and everything else to an upsert: ```yaml output: sql_raw: driver: postgres dsn: "${DSN}" max_in_flight: 8 batching: count: 100 period: 100ms queries: - when: 'root = meta("kafka_tombstone_message") == "true"' query: 'DELETE FROM orders WHERE id = $1' args_mapping: 'root = [ meta("kafka_key").parse_json().id ]' - query: | INSERT INTO orders (id, name, updated_at) VALUES ($1, $2, $3) ON CONFLICT (id) DO UPDATE SET name = EXCLUDED.name, updated_at = EXCLUDED.updated_at args_mapping: 'root = [ this.id, this.name, this.updated_at ]' ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that includes the same number of values in an array as the placeholder arguments in the [`query`](#query) field. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#batching)`batching` Allows you to configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yaml # Examples: batching: byte_size: 5000 count: 0 period: 1s # --- batching: count: 10 period: 1s # --- batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-byte_size)`batching.byte_size` An amount of bytes at which the batch should be flushed. If `0` disables size based batching. **Type**: `int` **Default**: `0` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "end_of_transaction" ``` ### [](#batching-count)`batching.count` A number of messages at which the batch should be flushed. If `0` disables count based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` A period in which an incomplete batch should be flushed regardless of its size. **Type**: `string` **Default**: `""` ```yaml # Examples: period: 1s # --- period: 1m # --- period: 500ms ``` ### [](#batching-processors)`batching.processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. This allows you to aggregate and archive the batch however you see fit. Please note that all resulting messages are flushed as a single batch, therefore splitting the batch into smaller batches using these processors is a no-op. **Type**: `array` ```yaml # Examples: processors: - archive: format: concatenate # --- processors: - archive: format: lines # --- processors: - archive: format: json_array ``` ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#max_in_flight)`max_in_flight` The maximum number of database statements to run in parallel. When consuming from Redpanda or Kafka, messages are ordered by partition within each transaction, so you can raise this above `1` to parallelize writes across partitions while preserving consume order within each partition. This lets high-throughput and change-data-capture pipelines scale without setting `max_in_flight` to `1`. Messages without a `kafka_partition` metadata field are treated as partition `0`. **Type**: `int` **Default**: `64` ### [](#queries)`queries[]` A list of database statements to run in addition to your main [`query`](#query). If you specify multiple queries, they are executed within a single transaction. For more information, see [Examples](#examples). **Type**: `array` ### [](#queries-args_mapping)`queries[].args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of placeholder arguments in the field `query`. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#queries-query)`queries[].query` The query to execute. The style of placeholder to use depends on the driver, some drivers require question marks (`?`) whereas others expect incrementing dollar signs (`$1`, `$2`, and so on) or colons (`:1`, `:2` and so on). The style to use is outlined in this table: | Driver | Placeholder Style | |---|---| | `clickhouse` | Dollar sign | | `mysql` | Question mark | | `postgres` | Dollar sign | | `pgx` | Dollar sign | | `mssql` | Question mark | | `sqlite` | Question mark | | `oracle` | Colon | | `snowflake` | Question mark | | `trino` | Question mark | | `gocosmos` | Colon | **Type**: `string` ### [](#queries-when)`queries[].when` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that, when set, is evaluated for each message to determine whether to execute this query. The mapping should return a boolean value. The first query in the list whose `when` condition evaluates to `true` (or that has no `when` condition) is executed. This enables conditional query routing based on message content or metadata without requiring `unsafe_dynamic_query`. **Type**: `string` ```yaml # Examples: when: root = meta("kafka_tombstone_message") == "true" # --- when: root = this.operation == "delete" ``` ### [](#query)`query` The query to execute. You must include the correct placeholders for the specified database driver. Some drivers use question marks (`?`), whereas others expect incrementing dollar signs (`$1`, `$2`, and so on) or colons (`:1`, `:2`, and so on). | Driver | Placeholder Style | | --- | --- | | clickhouse | Dollar sign ($) | | gocosmos | Colon (:) | | mysql | Question mark (?) | | mssql | Question mark (?) | | oracle | Colon (:) | | postgres | Dollar sign ($) | | snowflake | Question mark (?) | | spanner | Question mark (?) | | sqlite | Question mark (?) | | trino | Question mark (?) | **Type**: `string` ```yaml # Examples: query: INSERT INTO footable (foo, bar, baz) VALUES (?, ?, ?); ``` ### [](#unsafe_dynamic_query)`unsafe_dynamic_query` Whether to enable [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) in the query. Great care should be made to ensure your queries are defended against injection attacks. **Type**: `bool` **Default**: `false` ## [](#examples)Examples ### [](#table-insert-mysql)Table Insert (MySQL) Here we insert rows into a database by populating the columns id, name and topic with values extracted from messages and metadata: ```yaml output: sql_raw: driver: mysql dsn: foouser:foopassword@tcp(localhost:3306)/foodb query: "INSERT INTO footable (id, name, topic) VALUES (?, ?, ?);" args_mapping: | root = [ this.user.id, this.user.name, meta("kafka_topic"), ] ``` ### [](#dynamically-creating-tables-postgresql)Dynamically Creating Tables (PostgreSQL) Here we dynamically create output tables transactionally with inserting a record into the newly created table. ```yaml output: processors: - mapping: | root = this # Prevent SQL injection when using unsafe_dynamic_query meta table_name = "\"" + metadata("table_name").replace_all("\"", "\"\"") + "\"" sql_raw: driver: postgres dsn: postgres://localhost/postgres unsafe_dynamic_query: true queries: - query: | CREATE TABLE IF NOT EXISTS ${!metadata("table_name")} (id varchar primary key, document jsonb); - query: | INSERT INTO ${!metadata("table_name")} (id, document) VALUES ($1, $2) ON CONFLICT (id) DO UPDATE SET document = EXCLUDED.document; args_mapping: | root = [ this.id, this.document.string() ] ``` ### [](#conditional-cdc-queries-postgresql)Conditional CDC Queries (PostgreSQL) Route messages to different SQL operations based on message metadata. Tombstone messages trigger a DELETE, while all other messages perform an upsert. All operations within a batch execute in a single transaction, ordered by Kafka partition. ```yaml output: sql_raw: driver: postgres dsn: postgres://localhost/postgres max_in_flight: 8 batching: count: 100 period: 100ms queries: - when: 'root = meta("kafka_tombstone_message") == "true"' query: 'DELETE FROM users WHERE id = $1' args_mapping: 'root = [this.id]' - query: | INSERT INTO users (id, name, updated_at) VALUES ($1, $2, $3) ON CONFLICT (id) DO UPDATE SET name = EXCLUDED.name, updated_at = EXCLUDED.updated_at args_mapping: 'root = [this.id, this.name, this.updated_at]' ``` --- # Page 165: switch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/switch.md --- # switch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: switch latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/switch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/switch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/switch.adoc description: The switch output type allows you to route messages to different outputs based on their contents. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- The switch output type allows you to route messages to different outputs based on their contents. #### Common ```yml outputs: label: "" switch: retry_until_success: false cases: [] # No default (required) ``` #### Advanced ```yml outputs: label: "" switch: retry_until_success: false strict_mode: false cases: [] # No default (required) ``` Messages that do not pass the check of a single output case are effectively dropped. In order to prevent this outcome set the field [`strict_mode`](#strict_mode) to `true`, in which case messages that do not pass at least one case are considered failed and will be nacked and/or reprocessed depending on your input. ## [](#examples)Examples ### [](#basic-multiplexing)Basic Multiplexing The most common use for a switch output is to multiplex messages across a range of output destinations. The following config checks the contents of the field `type` of messages and sends `foo` type messages to an `amqp_1` output, `bar` type messages to a `gcp_pubsub` output, and everything else to a `redis_streams` output. Outputs can have their own processors associated with them, and in this example the `redis_streams` output has a processor that enforces the presence of a type field before sending it. ```yaml output: switch: cases: - check: this.type == "foo" output: amqp_1: urls: [ amqps://guest:guest@localhost:5672/ ] target_address: queue:/the_foos - check: this.type == "bar" output: gcp_pubsub: project: dealing_with_mike topic: mikes_bars - output: redis_streams: url: tcp://localhost:6379 stream: everything_else processors: - mapping: | root = this root.type = this.type | "unknown" ``` ### [](#control-flow)Control Flow The `continue` field allows messages that have passed a case to be tested against the next one also. This can be useful when combining non-mutually-exclusive case checks. In the following example a message that passes both the check of the first case as well as the second will be routed to both. ```yaml output: switch: cases: - check: 'this.user.interests.contains("walks").catch(false)' output: amqp_1: urls: [ amqps://guest:guest@localhost:5672/ ] target_address: queue:/people_what_think_good continue: true - check: 'this.user.dislikes.contains("videogames").catch(false)' output: gcp_pubsub: project: people topic: that_i_dont_want_to_hang_with ``` ## [](#fields)Fields ### [](#cases)`cases[]` A list of switch cases, outlining outputs that can be routed to. **Type**: `array` ```yaml # Examples: cases: - check: this.urls.contains("http://benthos.dev") continue: true output: cache: key: ${!json("id")} target: foo - output: s3: bucket: bar path: ${!json("id")} ``` ### [](#cases-check)`cases[].check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should be routed to the case output. If left empty the case always passes. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "foo" # --- check: this.contents.urls.contains("https://benthos.dev/") ``` ### [](#cases-continue)`cases[].continue` Indicates whether, if this case passes for a message, the next case should also be tested. **Type**: `bool` **Default**: `false` ### [](#cases-output)`cases[].output` An [output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about/) for messages that pass the check to be routed to. **Type**: `output` ### [](#retry_until_success)`retry_until_success` If a selected output fails to send a message this field determines whether it is reattempted indefinitely. If set to false the error is instead propagated back to the input level. If a message can be routed to >1 outputs it is usually best to set this to true in order to avoid duplicate messages being routed to an output. **Type**: `bool` **Default**: `false` ### [](#strict_mode)`strict_mode` This field determines whether an error should be reported if no condition is met. If set to true, an error is propagated back to the input level. The default behavior is false, which will drop the message. **Type**: `bool` **Default**: `false` --- # Page 166: sync_response **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/sync_response.md --- # sync_response > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sync_response latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/sync_response page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/sync_response.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/sync_response.adoc description: Returns the final message payload back to the input origin of the message, where it is dealt with according to that specific input type. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Returns the final message payload back to the input origin of the message, where it is dealt with according to that specific input type. ```yml # Config fields, showing default values output: label: "" sync_response: {} ``` --- # Page 167: timeplus **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/timeplus.md --- # timeplus > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: timeplus page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/outputs/timeplus page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/outputs/timeplus.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/outputs/timeplus.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-11-05" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Sends messages to a data stream on [Timeplus Enterprise (Cloud or Self-Hosted)](https://docs.timeplus.com/) using the [Ingest API](https://docs.timeplus.com/ingest-api), or directly to the `timeplusd` component in Timeplus Enterprise. #### Common ```yml # Common configuration fields, showing default values output: label: "" timeplus: target: timeplus url: https://us-west-2.timeplus.cloud workspace: "" # No default (optional) stream: "" # No default (required) apikey: "" # No default (optional) username: "" # No default (optional) password: "" # No default (optional) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" ``` #### Advanced ```yml # All configuration fields, showing default values output: label: "" timeplus: target: timeplus url: https://us-west-2.timeplus.cloud workspace: "" # No default (optional) stream: "" # No default (required) apikey: "" # No default (optional) username: "" # No default (optional) password: "" # No default (optional) max_in_flight: 64 batching: count: 0 byte_size: 0 period: "" check: "" processors: [] # No default (optional) ``` This output only accepts structured messages. All messages must: - Contain the same keys. - Use a structure that matches the schema of the destination data stream. If your upstream data source or pipeline returns unstructured messages, such as strings, you can configure an output processor to transform the messages. See the [Unstructured messages](#unstructured-messages) section for examples. ## [](#examples)Examples #### Timeplus Enterprise (Cloud) You must [generate an API key](https://docs.timeplus.com/apikey) using the web console of Timeplus Enterprise (Cloud). ```yaml output: timeplus: workspace: stream: apikey: ``` Replace the following placeholders with your own values: - ``: The ID of the workspace you want to send messages to. - ``: The name of the destination data stream. - ``: The API key for the Ingest API. #### Timeplus Enterprise (Self-Hosted) You must specify the username, password, and URL of the application server. ```yaml output: timeplus: url: http://localhost:8000 workspace: stream: username: password: ``` Replace the following placeholders with your own values: - ``: The ID of the workspace you want to send messages to. - ``: The name of the destination data stream. - ``: The username for the Timeplus application server. - ``: The password for the Timeplus application server. #### timeplusd You must specify the HTTP port for `timeplusd`. ```yaml output: timeplus: url: http://localhost:3218 stream: username: password: ``` Replace the following placeholders with your own values: - ``: The name of the destination data stream. - ``: The username for the Timeplus application server. - ``: The password for the Timeplus application server. ### [](#unstructured-messages)Unstructured messages If your upstream data source or pipeline returns unstructured messages, such as strings, you can configure an output processor to transform them into structured messages and then pass them to the output. In the following example, the `mapping` processor creates a field called `raw`, and uses the functions `content().string()` to store the original string content into it, thereby creating structured messages. If you use this example, you must also add the `raw` field name to the destination data stream, so that your message structure matches the schema of your destination data stream. ```yaml output: timeplus: workspace: stream: apikey: processors: - mapping: | root = {} root.raw = content().string() ``` ## [](#fields)Fields ### [](#target)`target` The destination platform. For Timeplus Enterprise (Cloud or Self-Hosted), enter `timeplus`, or `timeplusd` for the `timeplusd` component. **Type**: `string` **Default**: `timeplus` **Options**: `timeplus`, `timeplusd` ### [](#url)`url` The URL of your Timeplus instance, which should always include the schema and host. **Type**: `string` **Default**: `[https://us-west-2.timeplus.cloud](https://us-west-2.timeplus.cloud)` ```yml # Examples url: http://localhost:8000 url: http://127.0.0.1:3218 ``` ### [](#workspace)`workspace` The ID of the workspace you want to send messages to. This field is required if the `target` field is set to `timeplus`. **Type**: `string` ### [](#stream)`stream` The name of the destination data stream. Make sure the schema of the data stream matches this output. **Type**: `string` ### [](#apikey)`apikey` The API key for the Ingest API. You need to generate this in the web console of Timeplus Enterprise (Cloud). This field is required if you are sending messages to Timeplus Enterprise (Cloud). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#username)`username` The username for the Timeplus application server. This field is required if you are sending messages to Timeplus Enterprise (Self-Hosted) or `timeplusd`. **Type**: `string` ### [](#password)`password` The password for the Timeplus application server. This field is required if you are sending messages to Timeplus Enterprise (Self-Hosted) or `timeplusd`. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#max_in_flight)`max_in_flight` The maximum number of message batches to have in flight at a given time. Increase this number to improve throughput. **Type**: `int` **Default**: `64` ### [](#batching)`batching` Configure a [batching policy](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). **Type**: `object` ```yml # Examples batching: byte_size: 5000 count: 0 period: 1s batching: count: 10 period: 1s batching: check: this.contains("END BATCH") count: 0 period: 1m ``` ### [](#batching-count)`batching.count` The number of messages after which the batch is flushed. Set to `0` to disable count-based batching. **Type**: `int` **Default**: `0` ### [](#batching-byte_size)`batching.byte_size` The amount of bytes at which the batch is flushed. Set to `0` to disable size-based batching. **Type**: `int` **Default**: `0` ### [](#batching-period)`batching.period` The period of time after which an incomplete batch is flushed regardless of its size. **Type**: `string` **Default**: `""` ```yml # Examples period: 1s period: 1m period: 500ms ``` ### [](#batching-check)`batching.check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that returns a boolean value indicating whether a message should end a batch. **Type**: `string` **Default**: `""` ```yml # Examples check: this.type == "end_of_transaction" ``` ### [](#batching-processors)`batching.processors` For aggregating and archiving message batches, you can add a list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to a batch as it is flushed. All resulting messages are flushed as a single batch even when you configure processors to split the batch into smaller batches. **Type**: `array` ```yml # Examples processors: - archive: format: concatenate processors: - archive: format: lines processors: - archive: format: json_array ``` --- # Page 168: a2a_message **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/a2a_message.md --- # a2a_message > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: a2a_message latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/a2a_message page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/a2a_message.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/a2a_message.adoc description: Sends messages to an A2A (Agent-to-Agent) protocol agent and returns the response. page-git-created-date: "2026-02-18" page-git-modified-date: "2026-08-11" --- Sends messages to an A2A (Agent-to-Agent) protocol agent and returns the response. This processor enables Redpanda Connect pipelines to communicate with A2A protocol agents. Currently only JSON-RPC transport is supported. The processor sends a message to the agent and polls for task completion. The agent’s response is returned as the processor output. For more information about the A2A protocol, see [https://a2a-protocol.org/latest/specification](https://a2a-protocol.org/latest/specification) #### Common ```yml processors: label: "" a2a_message: agent_card_url: "" # No default (required) prompt: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" a2a_message: agent_card_url: "" # No default (required) prompt: "" # No default (optional) final_message_only: true ``` ## [](#fields)Fields ### [](#agent_card_url)`agent_card_url` URL for the A2A agent card. Can be either a base URL (e.g., `[https://example.com](https://example.com)`) or a full path to the agent card (e.g., `[https://example.com/.well-known/agent.json](https://example.com/.well-known/agent.json)`). If no path is provided, defaults to `/.well-known/agent.json`. Authentication uses OAuth2 from environment variables. **Type**: `string` ### [](#final_message_only)`final_message_only` If true, returns only the text from the final agent message (concatenated from all text parts). If false, returns the complete Message or Task object as structured data with full history, artifacts, and metadata. Example with final\_message\_only: true (default): ```none Here is the answer to your question... ``` Example with final\_message\_only: false: ```json { "id": "task-123", "contextId": "ctx-456", "status": { "state": "completed" }, "history": [ {"role": "user", "parts": [{"text": "Your question"}]}, {"role": "agent", "parts": [{"text": "Here is the answer to your question..."}]} ], "artifacts": [] } ``` **Type**: `bool` **Default**: `true` ### [](#prompt)`prompt` The user prompt to send to the agent. By default, the processor submits the entire payload as a string. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 169: Processors **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about.md --- # Processors > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Processors latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Redpanda Connect processors are functions applied to messages passing through a pipeline. The function signature allows a processor to mutate or drop messages depending on the content of the message. There are many types on offer but the most powerful are the [`mapping`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/) and [`mutation`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mutation/) processors. Processors are set via config, and depending on where in the config they are placed they will be run either immediately after a specific input (set in the input section), on all messages (set in the pipeline section) or before a specific output (set in the output section). Most processors apply to all messages and can be placed in the pipeline section: ```yaml pipeline: threads: 1 processors: - label: my_cool_mapping mapping: | root.message = this root.meta.link_count = this.links.length() ``` The `threads` field in the pipeline section determines how many parallel processing threads are created. You can read more about parallel processing in the [pipeline guide](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/processing_pipelines/). ## [](#labels)Labels Processors have an optional field `label` that can uniquely identify them in observability data such as metrics and logs. This can be useful when running configs with multiple nested processors, otherwise their metrics labels will be generated based on their composition. For more information check out the [metrics documentation](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/). ## [](#error-handling)Error handling Some processors have conditions whereby they might fail. Rather than throw these messages into the abyss Redpanda Connect still attempts to send these messages onwards, and has mechanisms for filtering, recovering or dead-letter queuing messages that have failed which can be read about [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ### [](#error-logs)Error logs Errors that occur during processing can be roughly separated into two groups; those that are unexpected intermittent errors such as connectivity problems, and those that are logical errors such as bad input data or unmatched schemas. All processing errors result in the messages being flagged as failed, [error metrics](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/) increasing for the given errored processor, and debug level logs being emitted that describe the error. Only errors that are known to be intermittent are also logged at the error level. The reason for this behavior is to prevent noisy logging in cases where logical errors are expected and will likely be [handled in config](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). However, this can also sometimes make it easy to miss logical errors in your configs when they lack error handling. If you suspect you are experiencing processing errors and do not wish to add error handling yet then a quick and easy way to expose those errors is to enable debug level logs with the cli flag `--log.level=debug` or by setting the level in config: ```yaml logger: level: DEBUG ``` ## [](#using-processors-as-outputs)Using processors as outputs It might be the case that a processor that results in a side effect, such as the [`sql_insert`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sql_insert/) or [`redis`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/redis/) processors, is the only side effect of a pipeline, and therefore could be considered the output. In such cases it’s possible to place these processors within a [`reject` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/reject/) so that they behave the same as regular outputs, where success results in dropping the message with an acknowledgement and failure results in a nack (or retry): ```yaml output: reject: 'failed to send data: ${! error() }' processors: - try: - redis: url: tcp://localhost:6379 command: sadd args_mapping: 'root = [ this.key, this.value ]' - mapping: root = deleted() ``` The way this works is that if your processor with the side effect (`redis` in this case) succeeds then the final `mapping` processor deletes the message which results in an acknowledgement. If the processor fails then the `try` block exits early without executing the `mapping` processor and instead the message is routed to the `reject` output, which nacks the message with an error message containing the error obtained from the `redis` processor. ## [](#batching-and-multiple-part-messages)Batching and multiple-part messages All Redpanda Connect processors support multiple-part messages, which are synonymous with batches. This enables [windowed processing](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/windowed_processing/) capabilities. Many processors are able to perform their behaviors on specific parts of a message batch, or on all parts, and have a field `parts` for specifying an array of part indexes they should apply to. If the list of target parts is empty these processors will be applied to all message parts. Part indexes can be negative, and if so the part will be selected from the end counting backwards starting from -1. E.g. if part = -1 then the selected part will be the last part of the message, if part = -2 then the part before the last element will be selected, and so on. Some processors such as [`dedupe`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/dedupe/) act across an entire batch, when instead we might like to perform them on individual messages of a batch. In this case the [`for_each`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/for_each/) processor can be used. You can read more about batching [in this document](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). --- # Page 170: archive **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive.md --- # archive > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: archive latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/archive page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/archive.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/archive.adoc description: Archives all the messages of a batch into a single message according to the selected archive format. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Archives all the messages of a batch into a single message according to the selected archive format. ```yml # Config fields, showing default values label: "" archive: format: "" # No default (required) path: "" ``` Some archive formats (such as tar, zip) treat each archive item (message part) as a file with a path. Since message parts only contain raw data a unique path must be generated for each part. This can be done by using function interpolations on the 'path' field as described in [Bloblang queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). For types that aren’t file based (such as binary) the file field is ignored. The resulting archived message adopts the metadata of the _first_ message part of the batch. The functionality of this processor depends on being applied across messages that are batched. You can find out more about batching [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#format)`format` The archiving format to apply. **Type**: `string` | Option | Summary | | --- | --- | | binary | Archive messages to a binary blob format. | | concatenate | Join the raw contents of each message into a single binary message. | | json_array | Attempt to parse each message as a JSON document and append the result to an array, which becomes the contents of the resulting message. | | lines | Join the raw contents of each message and insert a line break between each one. | | tar | Archive messages to a unix standard tape archive. | | zip | Archive messages to a zip file. | ### [](#path)`path` The path to set for each message in the archive (when applicable). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ```yaml # Examples: path: ${!count("files")}-${!timestamp_unix_nano()}.txt # --- path: ${!meta("kafka_key")}-${!json("id")}.json ``` ## [](#examples)Examples ### [](#tar-archive)Tar Archive If we had JSON messages in a batch each of the form: ```json {"doc":{"id":"foo","body":"hello world 1"}} ``` And we wished to tar archive them, setting their filenames to their respective unique IDs (with the extension `.json`), our config might look like this: ```yaml pipeline: processors: - archive: format: tar path: ${!json("doc.id")}.json ``` --- # Page 171: avro **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/avro.md --- # avro > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: avro latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/avro page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/avro.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/avro.adoc description: Performs Avro based operations on messages based on a schema. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Performs Avro based operations on messages based on a schema. ```yml # Config fields, showing default values label: "" avro: operator: "" # No default (required) encoding: textual schema: "" schema_path: "" ``` > ⚠️ **WARNING** > > If you are consuming or generating messages using a schema registry service then it is likely this processor will fail as those services require messages to be prefixed with the identifier of the schema version being used. Instead, try the [`schema_registry_encode`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/schema_registry_encode/) and [`schema_registry_decode`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/schema_registry_decode/) processors. ## [](#operators)Operators ### [](#to_json)`to_json` Converts Avro documents into a JSON structure. This makes it easier to manipulate the contents of the document within Benthos. The encoding field specifies how the source documents are encoded. ### [](#from_json)`from_json` Attempts to convert JSON documents into Avro documents according to the specified encoding. ## [](#fields)Fields ### [](#encoding)`encoding` An Avro encoding format to use for conversions to and from a schema. **Type**: `string` **Default**: `textual` **Options**: `textual`, `binary`, `single` ### [](#operator)`operator` The [operator](#operators) to execute **Type**: `string` **Options**: `to_json`, `from_json` ### [](#schema)`schema` A full Avro schema to use. **Type**: `string` **Default**: `""` ### [](#schema_path)`schema_path` The path of a schema document to apply. Use either this or the `schema` field. URLs must begin with `file://` or `http://`. Note that `file://` URLs must use absolute paths (e.g. `[file:///absolute/path/to/spec.avsc](file:///absolute/path/to/spec.avsc)`); relative paths are not supported. **Type**: `string` **Default**: `""` ```yaml # Examples: schema_path: file:///path/to/spec.avsc # --- schema_path: http://localhost:8081/path/to/spec/versions/1 ``` --- # Page 172: aws_bedrock_chat **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/aws_bedrock_chat.md --- # aws_bedrock_chat > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_bedrock_chat latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/aws_bedrock_chat page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/aws_bedrock_chat.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/aws_bedrock_chat.adoc description: Generates responses to messages in a chat conversation, using the AWS Bedrock API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates responses to messages in a chat conversation, using the [AWS Bedrock API](https://aws.amazon.com/bedrock/). #### Common ```yml processors: label: "" aws_bedrock_chat: model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) max_tokens: "" # No default (optional) temperature: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" aws_bedrock_chat: region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) max_tokens: "" # No default (optional) temperature: "" # No default (optional) stop: [] # No default (optional) top_p: "" # No default (optional) ``` This processor sends prompts to your chosen large language model (LLM) and generates text from the responses, using the AWS Bedrock API. For more information, see the [AWS Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide). ## [](#fields)Fields ### [](#credentials)`credentials` Configure which AWS credentials to use (optional). For more information, see [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` The profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` The role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` The external ID to use when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials you want to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials you want to use. You must enter this value when using short-term credentials. **Type**: `string` ### [](#endpoint)`endpoint` A custom endpoint URL for AWS API requests. Use this to connect to AWS-compatible services or local testing environments instead of the standard AWS endpoints. **Type**: `string` ### [](#max_tokens)`max_tokens` The maximum number of tokens to allow in the generated response. **Type**: `int` ### [](#model)`model` The model ID to use. For a full list, see the [AWS Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/model-ids.html). **Type**: `string` ```yaml # Examples: model: amazon.titan-text-express-v1 # --- model: anthropic.claude-3-5-sonnet-20240620-v1:0 # --- model: cohere.command-text-v14 # --- model: meta.llama3-1-70b-instruct-v1:0 # --- model: mistral.mistral-large-2402-v1:0 ``` ### [](#prompt)`prompt` The prompt you want to generate a response for. By default, the processor submits the entire payload as a string. **Type**: `string` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#stop)`stop[]` A list of stop sequences. A stop sequence is a sequence of characters that causes the model to stop generating the response. **Type**: `array` ### [](#system_prompt)`system_prompt` The system prompt to submit to the AWS Bedrock LLM. **Type**: `string` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#temperature)`temperature` The likelihood of the model selecting higher-probability options while generating a response. A lower value makes the model more likely to choose higher-probability options. A higher value makes the model more likely to choose lower-probability options. **Type**: `float` ### [](#top_p)`top_p` The percentage of most-likely candidates that the model considers for the next token. For example, if you choose a value of `0.8`, the model selects from the top 80% of the probability distribution of tokens that could be next in the sequence. **Type**: `float` --- # Page 173: aws_bedrock_embeddings **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/aws_bedrock_embeddings.md --- # aws_bedrock_embeddings > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_bedrock_embeddings page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/aws_bedrock_embeddings page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/aws_bedrock_embeddings.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/aws_bedrock_embeddings.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-10-16" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Generates vector embeddings from text prompts, using the [AWS Bedrock API](https://aws.amazon.com/bedrock/). #### Common ```yaml # Common config fields, showing default values label: "" aws_bedrock_embeddings: model: amazon.titan-embed-text-v1 # No default (required) text: "" # No default (optional) ``` #### Advanced ```yaml # All config fields, showing default values label: "" aws_bedrock_embeddings: region: "" endpoint: "" credentials: from_ec2_role: false role: "" role_external_id: "" model: amazon.titan-embed-text-v1 # No default (required) text: "" # No default (optional) ``` This processor sends text prompts to your chosen large language model (LLM), which generates vector embeddings for them using the AWS Bedrock API. For more information, see the [AWS Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide). ## [](#fields)Fields ### [](#credentials)`credentials` Manually configure the AWS credentials to use (optional). For more information, see the [Amazon Web Services guide](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of the AWS credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` The profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` The role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to use when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the AWS credentials in use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the AWS credentials in use. This is a required value for short-term credentials. **Type**: `string` ### [](#endpoint)`endpoint` A custom endpoint URL for AWS API requests. Use this to connect to AWS-compatible services or local testing environments instead of the standard AWS endpoints. **Type**: `string` ### [](#input_type)`input_type` Specifies the type of input passed to the model. Required by Cohere embedding models; ignored by Amazon Titan models. **Type**: `string` | Option | Summary | | --- | --- | | classification | Used for embeddings passed through a text classifier. | | clustering | Used for the embeddings run through a clustering algorithm. | | search_document | Used for embeddings stored in a vector database for search use-cases. | | search_query | Used for embeddings of search queries run against a vector DB to find relevant documents. | ### [](#model)`model` The ID of the LLM that you want to use to generate vector embeddings. For a full list, see the [AWS Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/model-ids.html). **Type**: `string` ```yaml # Examples: model: amazon.titan-embed-text-v1 # --- model: amazon.titan-embed-text-v2:0 # --- model: cohere.embed-english-v3 # --- model: cohere.embed-multilingual-v3 # --- model: cohere.embed-v4:0 ``` ### [](#region)`region` The region in which your AWS resources are hosted. **Type**: `string` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#text)`text` The prompt you want to generate a vector embedding for. The processor submits the entire payload as a string. **Type**: `string` --- # Page 174: aws_dynamodb_partiql **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/aws_dynamodb_partiql.md --- # aws_dynamodb_partiql > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_dynamodb_partiql latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/aws_dynamodb_partiql page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/aws_dynamodb_partiql.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/aws_dynamodb_partiql.adoc description: Executes a PartiQL expression against a DynamoDB table for each message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a PartiQL expression against a DynamoDB table for each message. #### Common ```yml processors: label: "" aws_dynamodb_partiql: query: "" # No default (required) args_mapping: "" ``` #### Advanced ```yml processors: label: "" aws_dynamodb_partiql: query: "" # No default (required) unsafe_dynamic_query: false use_batch: true args_mapping: "" region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) ``` Both writes or reads are supported, when the query is a read the contents of the message will be replaced with the result. This processor is more efficient when messages are pre-batched as the whole batch will be executed in a single call. ## [](#examples)Examples ### [](#insert)Insert The following example inserts rows into the table footable with the columns foo, bar and baz populated with values extracted from messages: ```yaml pipeline: processors: - aws_dynamodb_partiql: query: "INSERT INTO footable VALUE {'foo':'?','bar':'?','baz':'?'}" args_mapping: | root = [ { "S": this.foo }, { "S": meta("kafka_topic") }, { "S": this.document.content }, ] ``` ### [](#query-a-gsi-for-a-single-record)Query a GSI for a single record The following example looks up a single record from the table footable using the global secondary index index\_name, matching on the field bar. BatchExecuteStatement can’t query a GSI, so use\_batch is disabled: ```yaml pipeline: processors: - aws_dynamodb_partiql: query: "SELECT * FROM \"footable\".\"index_name\" WHERE bar = ?" use_batch: false args_mapping: | root = [ { "S": this.bar }, ] ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that, for each message, creates a list of arguments to use with the query. **Type**: `string` **Default**: `""` ### [](#credentials)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#query)`query` A PartiQL query to execute for each message. **Type**: `string` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#unsafe_dynamic_query)`unsafe_dynamic_query` Whether to enable dynamic queries that support interpolation functions. **Type**: `bool` **Default**: `false` ### [](#use_batch)`use_batch` Whether to execute all messages in a batch as a single `BatchExecuteStatement` call. Set this to `false` to execute one `ExecuteStatement` call per message instead, which is required for PartiQL `SELECT` queries against a global secondary index (GSI), because `BatchExecuteStatement` does not support querying a GSI. Only the first result row is used when a query returns multiple items. **Type**: `bool` **Default**: `true` --- # Page 175: aws_lambda **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/aws_lambda.md --- # aws_lambda > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: aws_lambda latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/aws_lambda page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/aws_lambda.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/aws_lambda.adoc description: Invokes an AWS lambda for each message. The contents of the message is the payload of the request, and the result of the invocation will become the new contents of the message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Invokes an AWS lambda for each message. The contents of the message is the payload of the request, and the result of the invocation will become the new contents of the message. #### Common ```yml processors: label: "" aws_lambda: parallel: false function: "" # No default (required) ``` #### Advanced ```yml processors: label: "" aws_lambda: parallel: false function: "" # No default (required) rate_limit: "" region: "" # No default (optional) endpoint: "" # No default (optional) tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s credentials: profile: "" # No default (optional) id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) from_ec2_role: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) timeout: 5s retries: 3 ``` The `rate_limit` field can be used to specify a rate limit [resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) to cap the rate of requests across parallel components service wide. In order to map or encode the payload to a specific request body, and map the response back into the original payload instead of replacing it entirely, you can use the [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/). ## [](#error-handling)Error handling When Redpanda Connect is unable to connect to the AWS endpoint or is otherwise unable to invoke the target lambda function it will retry the request according to the configured number of retries. Once these attempts have been exhausted the failed message will continue through the pipeline with it’s contents unchanged, but flagged as having failed, allowing you to use [standard processor error handling patterns](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). However, if the invocation of the function is successful but the function itself throws an error, then the message will have it’s contents updated with a JSON payload describing the reason for the failure, and a metadata field `lambda_function_error` will be added to the message allowing you to detect and handle function errors with a [`branch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/): ```yaml pipeline: processors: - branch: processors: - aws_lambda: function: foo result_map: | root = if meta().exists("lambda_function_error") { throw("Invocation failed due to %v: %v".format(this.errorType, this.errorMessage)) } else { this } output: switch: retry_until_success: false cases: - check: errored() output: reject: ${! error() } - output: resource: somewhere_else ``` ## [](#credentials)Credentials By default Redpanda Connect will use a shared credentials file when connecting to AWS services. It’s also possible to set them explicitly at the component level, allowing you to transfer data across accounts. You can find out more in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). ## [](#examples)Examples ### [](#branched-invoke)Branched Invoke This example uses a [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) to map a new payload for triggering a lambda function with an ID and username from the original message, and the result of the lambda is discarded, meaning the original message is unchanged. ```yaml pipeline: processors: - branch: request_map: '{"id":this.doc.id,"username":this.user.name}' processors: - aws_lambda: function: trigger_user_update ``` ## [](#fields)Fields ### [](#credentials-2)`credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#credentials-from_ec2_role)`credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#credentials-id)`credentials.id` The ID of credentials to use. **Type**: `string` ### [](#credentials-profile)`credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#credentials-role)`credentials.role` A role ARN to assume. **Type**: `string` ### [](#credentials-role_external_id)`credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#credentials-secret)`credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#credentials-token)`credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#endpoint)`endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#function)`function` The function to invoke. **Type**: `string` ### [](#parallel)`parallel` Whether messages of a batch should be dispatched in parallel. **Type**: `bool` **Default**: `false` ### [](#rate_limit)`rate_limit` An optional [`rate_limit`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) to throttle invocations by. **Type**: `string` **Default**: `""` ### [](#region)`region` The AWS region to target. **Type**: `string` ### [](#retries)`retries` The maximum number of retry attempts for each message. **Type**: `int` **Default**: `3` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` The maximum period of time to wait before abandoning an invocation. **Type**: `string` **Default**: `5s` --- # Page 176: azure_cosmosdb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/azure_cosmosdb.md --- # azure_cosmosdb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: azure_cosmosdb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/azure_cosmosdb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/azure_cosmosdb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/azure_cosmosdb.adoc description: Creates or updates messages as JSON documents in Azure CosmosDB. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Creates or updates messages as JSON documents in [Azure CosmosDB](https://learn.microsoft.com/en-us/azure/cosmos-db/introduction). ### Common ```yml processors: label: "" azure_cosmosdb: endpoint: "" # No default (optional) account_key: "" # No default (optional) connection_string: "" # No default (optional) database: "" # No default (required) container: "" # No default (required) partition_keys_map: "" # No default (required) operation: Create item_id: "" # No default (optional) ``` ### Advanced ```yml processors: label: "" azure_cosmosdb: endpoint: "" # No default (optional) account_key: "" # No default (optional) connection_string: "" # No default (optional) database: "" # No default (required) container: "" # No default (required) partition_keys_map: "" # No default (required) operation: Create patch_operations: [] # No default (optional) patch_condition: "" # No default (optional) auto_id: true item_id: "" # No default (optional) enable_content_response_on_write: true ``` When creating documents, each message must have the `id` property (case-sensitive) set (or use `auto_id: true`). It is the unique name that identifies the document, that is, no two documents share the same `id` within a logical partition. The `id` field must not exceed 255 characters. [See details](https://learn.microsoft.com/en-us/rest/api/cosmos-db/documents). The `partition_keys` field must resolve to the same value(s) across the entire message batch. ## [](#credentials)Credentials You can use one of the following authentication mechanisms: - Set the `endpoint` field and the `account_key` field - Set only the `endpoint` field to use [DefaultAzureCredential](https://pkg.go.dev/github.com/Azure/azure-sdk-for-go/sdk/azidentity#DefaultAzureCredential) - Set the `connection_string` field ## [](#metadata)Metadata This component adds the following metadata fields to each message: - `activity_id` - `request_charge` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#batching)Batching CosmosDB limits the maximum batch size to 100 messages and the payload must not exceed 2MB ([details here](https://learn.microsoft.com/en-us/azure/cosmos-db/concepts-limits#per-request-limits)). ## [](#examples)Examples ### [](#patch-documents)Patch documents Query documents from a container and patch them. ```yaml input: azure_cosmosdb: endpoint: http://localhost:8080 account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== database: blobbase container: blobfish partition_keys_map: root = "AbyssalPlain" query: SELECT * FROM blobfish processors: - mapping: | root = "" meta habitat = json("habitat") meta id = this.id - azure_cosmosdb: endpoint: http://localhost:8080 account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== database: testdb container: blobfish partition_keys_map: root = json("habitat") item_id: ${! meta("id") } operation: Patch patch_operations: # Add a new /diet field - operation: Add path: /diet value_map: root = json("diet") # Remove the first location from the /locations array field - operation: Remove path: /locations/0 # Add new location at the end of the /locations array field - operation: Add path: /locations/- value_map: root = "Challenger Deep" # Return the updated document enable_content_response_on_write: true ``` ## [](#fields)Fields ### [](#account_key)`account_key` Account key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: account_key: C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw== ``` ### [](#auto_id)`auto_id` Automatically set the item `id` field to a random UUID v4. If the `id` field is already set, then it will not be overwritten. Setting this to `false` can improve performance, since the messages will not have to be parsed. **Type**: `bool` **Default**: `true` ### [](#connection_string)`connection_string` Connection string. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: connection_string: AccountEndpoint=https://localhost:8081/;AccountKey=C2y6yDjf5/R+ob0N8A7Cgv30VRDJIWEHLM+4QDU5DE2nQ9nDuVTqobD4b8mGGyPMbIZnqyMsEcaGQy67XIw/Jw==; ``` ### [](#container)`container` Container. **Type**: `string` ```yaml # Examples: container: testcontainer ``` ### [](#database)`database` Database. **Type**: `string` ```yaml # Examples: database: testdb ``` ### [](#enable_content_response_on_write)`enable_content_response_on_write` Enable content response on write operations. To save some bandwidth, set this to false if you don’t need to receive the updated message(s) from the server, in which case the processor will not modify the content of the messages which are fed into it. Applies to every operation except Read. **Type**: `bool` **Default**: `true` ### [](#endpoint)`endpoint` CosmosDB endpoint. **Type**: `string` ```yaml # Examples: endpoint: https://localhost:8081 ``` ### [](#item_id)`item_id` ID of item to replace or delete. Only used by the Replace and Delete operations This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: item_id: ${! json("id") } ``` ### [](#operation)`operation` Operation. **Type**: `string` **Default**: `Create` | Option | Summary | | --- | --- | | Create | Create operation. | | Delete | Delete operation. | | Patch | Patch operation. | | Read | Read operation. | | Replace | Replace operation. | | Upsert | Upsert operation. | ### [](#partition_keys_map)`partition_keys_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to a single partition key value or an array of partition key values of type string, integer or boolean. Currently, hierarchical partition keys are not supported so only one value may be provided. **Type**: `string` ```yaml # Examples: partition_keys_map: root = "blobfish" # --- partition_keys_map: root = 41 # --- partition_keys_map: root = true # --- partition_keys_map: root = null # --- partition_keys_map: root = json("blobfish").depth ``` ### [](#patch_condition)`patch_condition` Patch operation condition. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: patch_condition: from c where not is_defined(c.blobfish) ``` ### [](#patch_operations)`patch_operations[]` Patch operations to be performed when `operation: Patch` . **Type**: `array` ### [](#patch_operations-operation)`patch_operations[].operation` Operation. **Type**: `string` **Default**: `Add` | Option | Summary | | --- | --- | | Add | Add patch operation. | | Increment | Increment patch operation. | | Remove | Remove patch operation. | | Replace | Replace patch operation. | | Set | Set patch operation. | ### [](#patch_operations-path)`patch_operations[].path` Path. **Type**: `string` ```yaml # Examples: path: /foo/bar/baz ``` ### [](#patch_operations-value_map)`patch_operations[].value_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to a value of any type that is supported by CosmosDB. **Type**: `string` ```yaml # Examples: value_map: root = "blobfish" # --- value_map: root = 41 # --- value_map: root = true # --- value_map: root = json("blobfish").depth # --- value_map: root = [1, 2, 3] ``` ## [](#cosmosdb-emulator)CosmosDB emulator If you wish to run the CosmosDB emulator that is referenced in the documentation [here](https://learn.microsoft.com/en-us/azure/cosmos-db/linux-emulator), the following Docker command should do the trick: ```bash > docker run --rm -it -p 8081:8081 --name=cosmosdb -e AZURE_COSMOS_EMULATOR_PARTITION_COUNT=10 -e AZURE_COSMOS_EMULATOR_ENABLE_DATA_PERSISTENCE=false mcr.microsoft.com/cosmosdb/linux/azure-cosmos-emulator ``` Note: `AZURE_COSMOS_EMULATOR_PARTITION_COUNT` controls the number of partitions that will be supported by the emulator. The bigger the value, the longer it takes for the container to start up. Additionally, instead of installing the container self-signed certificate which is exposed via `[https://localhost:8081/_explorer/emulator.pem](https://localhost:8081/_explorer/emulator.pem)`, you can run [mitmproxy](https://mitmproxy.org/) like so: ```bash > mitmproxy -k --mode "reverse:https://localhost:8081" ``` Then you can access the CosmosDB UI via `[http://localhost:8080/_explorer/index.html](http://localhost:8080/_explorer/index.html)` and use `[http://localhost:8080](http://localhost:8080)` as the CosmosDB endpoint. --- # Page 177: benchmark **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/benchmark.md --- # benchmark > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: benchmark latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/benchmark page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/benchmark.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/benchmark.adoc description: Logs basic throughput statistics of messages that pass through this processor. page-git-created-date: "2024-12-16" page-git-modified-date: "2026-05-26" --- Logs throughput statistics for processed messages, and provides a summary of those statistics over the lifetime of the processor. ```yml # Configuration fields, showing default values label: "" benchmark: interval: 5s count_bytes: true ``` ## [](#throughput-statistics)Throughput statistics This processor logs the following rolling statistics at a [configurable interval](#interval) to help you to understand the current performance of your pipeline: - The number of messages processed per second. - The number of bytes processed per second (optional). For example: ```bash INFO rolling stats: 1 msg/sec, 407 B/sec ``` When the processor shuts down, it also logs a summary of the number and size of messages processed during its lifetime. For example: ```bash INFO total stats: 1.00186 msg/sec, 425 B/sec ``` ## [](#fields)Fields ### [](#count_bytes)`count_bytes` Whether to measure the number of bytes per second of throughput. If set to `true`, Redpanda Connect must serialize structured data to count the number of bytes processed, which can unnecessarily degrade performance if serialization is not required elsewhere in your pipeline. **Type**: `bool` **Default**: `true` ### [](#interval)`interval` How often to emit rolling statistics. Set to `0`, if you only want to log summary statistics when the processor shuts down. **Type**: `string` **Default**: `5s` --- # Page 178: bloblang **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/bloblang.md --- # bloblang > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: bloblang latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/bloblang page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/bloblang.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/bloblang.adoc description: Executes a Bloblang mapping on messages. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) mapping on messages. ```yml # Config fields, showing default values label: "" bloblang: "" ``` Bloblang is a powerful language that enables a wide range of mapping, transformation and filtering tasks. For more information see [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/). If your mapping is large and you’d prefer for it to live in a separate file then you can execute a mapping directly from a file with the expression `from ""`, where the path must be absolute, or relative from the location that Redpanda Connect is executed from. ## [](#component-rename)Component rename This processor was recently renamed to the [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/) in order to make the purpose of the processor more prominent. It is still valid to use the existing `bloblang` name but eventually it will be deprecated and replaced by the new name in example configs. ## [](#examples)Examples ### [](#mapping)Mapping Given JSON documents containing an array of fans: ```json { "id":"foo", "description":"a show about foo", "fans":[ {"name":"bev","obsession":0.57}, {"name":"grace","obsession":0.21}, {"name":"ali","obsession":0.89}, {"name":"vic","obsession":0.43} ] } ``` We can reduce the fans to only those with an obsession score above 0.5, giving us: ```json { "id":"foo", "description":"a show about foo", "fans":[ {"name":"bev","obsession":0.57}, {"name":"ali","obsession":0.89} ] } ``` With the following config: ```yaml pipeline: processors: - bloblang: | root = this root.fans = this.fans.filter(fan -> fan.obsession > 0.5) ``` ### [](#more-mapping)More Mapping When receiving JSON documents of the form: ```json { "locations": [ {"name": "Seattle", "state": "WA"}, {"name": "New York", "state": "NY"}, {"name": "Bellevue", "state": "WA"}, {"name": "Olympia", "state": "WA"} ] } ``` We could collapse the location names from the state of Washington into a field `Cities`: ```json {"Cities": "Bellevue, Olympia, Seattle"} ``` With the following config: ```yaml pipeline: processors: - bloblang: | root.Cities = this.locations. filter(loc -> loc.state == "WA"). map_each(loc -> loc.name). sort().join(", ") ``` ## [](#error-handling)Error handling Bloblang mappings can fail, in which case the message remains unchanged, errors are logged, and the message is flagged as having failed, allowing you to use [standard processor error handling patterns](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). However, Bloblang itself also provides powerful ways of ensuring your mappings do not fail by specifying desired fallback behavior, which you can read about in [Error handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#error-handling.adoc). --- # Page 179: bounds_check **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/bounds_check.md --- # bounds_check > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: bounds_check latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/bounds_check page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/bounds_check.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/bounds_check.adoc description: Removes messages (and batches) that do not fit within certain size boundaries. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Removes messages (and batches) that do not fit within certain size boundaries. #### Common ```yml processors: label: "" bounds_check: max_part_size: 1073741824 min_part_size: 1 ``` #### Advanced ```yml processors: label: "" bounds_check: max_part_size: 1073741824 min_part_size: 1 max_parts: 100 min_parts: 1 ``` ## [](#fields)Fields ### [](#max_part_size)`max_part_size` The maximum size of a message to allow (in bytes) **Type**: `int` **Default**: `1073741824` ### [](#max_parts)`max_parts` The maximum size of message batches to allow (in message count) **Type**: `int` **Default**: `100` ### [](#min_part_size)`min_part_size` The minimum size of a message to allow (in bytes) **Type**: `int` **Default**: `1` ### [](#min_parts)`min_parts` The minimum size of message batches to allow (in message count) **Type**: `int` **Default**: `1` --- # Page 180: branch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch.md --- # branch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: branch latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/branch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/branch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/branch.adoc description: The branch processor allows you to create a new request message via a Bloblang mapping, execute a list of processors on the request messages, and, finally, map the result back into the source message using another mapping. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- The `branch` processor allows you to create a new request message via a [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/), execute a list of processors on the request messages, and, finally, map the result back into the source message using another mapping. ```yml # Config fields, showing default values label: "" branch: request_map: "" processors: [] # No default (required) result_map: "" ``` This is useful for preserving the original message contents when using processors that would otherwise replace the entire contents. ## [](#metadata)Metadata Metadata fields that are added to messages during branch processing will not be automatically copied into the resulting message. In order to do this you should explicitly declare in your `result_map` either a wholesale copy with `meta = metadata()`, or selective copies with `meta foo = metadata("bar")` and so on. It is also possible to reference the metadata of the origin message in the `result_map` using the [`@` operator](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#metadata). ## [](#error-handling)Error handling If the `request_map` fails the child processors will not be executed. If the child processors themselves result in an (uncaught) error then the `result_map` will not be executed. If the `result_map` fails the message will remain unchanged. Under any of these conditions standard [error handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/) can be used in order to filter, DLQ or recover the failed messages. ## [](#conditional-branching)Conditional branching If the root of your request map is set to `deleted()` then the branch processors are skipped for the given message, this allows you to conditionally branch messages. ## [](#fields)Fields ### [](#processors)`processors[]` A list of processors to apply to mapped requests. When processing message batches the resulting batch must match the size and ordering of the input batch, therefore filtering, grouping should not be performed within these processors. **Type**: `array` ### [](#request_map)`request_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that describes how to create a request payload suitable for the child processors of this branch. If left empty then the branch will begin with an exact copy of the origin message (including metadata). **Type**: `string` **Default**: `""` ```yaml # Examples: request_map: |- root = { "id": this.doc.id, "content": this.doc.body.text } # --- request_map: |- root = if this.type == "foo" { this.foo.request } else { deleted() } ``` ### [](#result_map)`result_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that describes how the resulting messages from branched processing should be mapped back into the original payload. If left empty the origin message will remain unchanged (including metadata). **Type**: `string` **Default**: `""` ```yaml # Examples: result_map: |- meta foo_code = metadata("code") root.foo_result = this # --- result_map: |- meta = metadata() root.bar.body = this.body root.bar.id = this.user.id # --- result_map: root.raw_result = content().string() # --- result_map: |- root.enrichments.foo = if metadata("request_failed") != null { throw(metadata("request_failed")) } else { this } # --- result_map: |- # Retain only the updated metadata fields which were present in the origin message meta = metadata().filter(v -> @.get(v.key) != null) ``` ## [](#examples)Examples ### [](#http-request)HTTP Request This example strips the request message into an empty body, grabs an HTTP payload, and places the result back into the original message at the path `image.pull_count`: ```yaml pipeline: processors: - branch: request_map: 'root = ""' processors: - http: url: https://hub.docker.com/v2/repositories/jeffail/benthos verb: GET headers: Content-Type: application/json result_map: root.image.pull_count = this.pull_count # Example input: {"id":"foo","some":"pre-existing data"} # Example output: {"id":"foo","some":"pre-existing data","image":{"pull_count":1234}} ``` ### [](#non-structured-results)Non Structured Results When the result of your branch processors is unstructured and you wish to simply set a resulting field to the raw output use the content function to obtain the raw bytes of the resulting message and then coerce it into your value type of choice: ```yaml pipeline: processors: - branch: request_map: 'root = this.document.id' processors: - cache: resource: descriptions_cache key: ${! content() } operator: get result_map: root.document.description = content().string() # Example input: {"document":{"id":"foo","content":"hello world"}} # Example output: {"document":{"id":"foo","content":"hello world","description":"this is a cool doc"}} ``` ### [](#lambda-function)Lambda Function This example maps a new payload for triggering a lambda function with an ID and username from the original message, and the result of the lambda is discarded, meaning the original message is unchanged. ```yaml pipeline: processors: - branch: request_map: '{"id":this.doc.id,"username":this.user.name}' processors: - aws_lambda: function: trigger_user_update # Example input: {"doc":{"id":"foo","body":"hello world"},"user":{"name":"fooey"}} # Output matches the input, which is unchanged ``` ### [](#conditional-caching)Conditional Caching This example caches a document by a message ID only when the type of the document is a foo: ```yaml pipeline: processors: - branch: request_map: | meta id = this.id root = if this.type == "foo" { this.document } else { deleted() } processors: - cache: resource: TODO operator: set key: ${! @id } value: ${! content() } ``` --- # Page 181: cache **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cache.md --- # cache > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: cache latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/cache page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/cache.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/cache.adoc description: Performs operations against a cache resource for each message, allowing you to store or retrieve data within message payloads. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Performs operations against a [cache resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) for each message, allowing you to store or retrieve data within message payloads. #### Common ```yml processors: label: "" cache: resource: "" # No default (required) operator: "" # No default (required) key: "" # No default (required) value: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" cache: resource: "" # No default (required) operator: "" # No default (required) key: "" # No default (required) value: "" # No default (optional) ttl: "" # No default (optional) ``` For use cases where you wish to cache the result of processors, consider using the [`cached` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cached/) instead. This processor will interpolate functions within the `key` and `value` fields individually for each message. This allows you to specify dynamic keys and values based on the contents of the message payloads and metadata. You can find a list of functions in [Bloblang queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#examples)Examples ### [](#deduplication)Deduplication Deduplication can be done using the add operator with a key extracted from the message payload, since it fails when a key already exists we can remove the duplicates using a [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/): ```yaml pipeline: processors: - cache: resource: foocache operator: add key: '${! json("message.id") }' value: "storeme" - mapping: root = if errored() { deleted() } cache_resources: - label: foocache redis: url: tcp://TODO:6379 ``` ### [](#deduplication-batch-wide)Deduplication Batch-Wide Sometimes it’s necessary to deduplicate a batch of messages (also known as a window) by a single identifying value. This can be done by introducing a [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/), which executes the cache only once on behalf of the batch, in this case with a value make from a field extracted from the first and last messages of the batch: ```yaml pipeline: processors: # Try and add one message to a cache that identifies the whole batch - branch: request_map: | root = if batch_index() == 0 { json("id").from(0) + json("meta.tail_id").from(-1) } else { deleted() } processors: - cache: resource: foocache operator: add key: ${! content() } value: t # Delete all messages if we failed - mapping: | root = if errored().from(0) { deleted() } ``` ### [](#hydration)Hydration It’s possible to enrich payloads with content previously stored in a cache by using the [`branch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) processor: ```yaml pipeline: processors: - branch: processors: - cache: resource: foocache operator: get key: '${! json("message.document_id") }' result_map: 'root.message.document = this' # NOTE: If the data stored in the cache is not valid JSON then use # something like this instead: # result_map: 'root.message.document = content().string()' cache_resources: - label: foocache memcached: addresses: [ "TODO:11211" ] ``` ## [](#fields)Fields ### [](#key)`key` A key to use with the cache. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#operator)`operator` The [operation](#operators) to perform with the cache. **Type**: `string` **Options**: `set`, `add`, `get`, `delete`, `exists` ### [](#resource)`resource` The [`cache` resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) to target with this processor. **Type**: `string` ### [](#ttl)`ttl` The time to live (TTL) of each individual item as a duration string. After this period an item will be eligible for removal during the next compaction. Not all caches support per-key TTLs, those that do will have a configuration field `default_ttl`, and those that do not will fall back to their generally configured TTL setting. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: ttl: 60s # --- ttl: 5m # --- ttl: 36h ``` ### [](#value)`value` A value to use with the cache (when applicable). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ## [](#operators)Operators ### [](#set)`set` Set a key in the cache to a value. If the key already exists the contents are overridden. ### [](#add)`add` Set a key in the cache to a value. If the key already exists the action fails with a 'key already exists' error, which can be detected with [processor error handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ### [](#get)`get` Retrieve the contents of a cached key and replace the original message payload with the result. If the key does not exist the action fails with an error, which can be detected with [processor error handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ### [](#exists)`exists` Check whether a specific key is in the cache and replace the original message payload with `true` if the key exists, or `false` if it doesn’t. ### [](#delete)`delete` Delete a key and its contents from the cache. If the key does not exist the action is a no-op and will not fail with an error. --- # Page 182: cached **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cached.md --- # cached > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: cached latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/cached page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/cached.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/cached.adoc description: Cache the result of applying one or more processors to messages identified by a key. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Cache the result of applying one or more processors to messages identified by a key. If the key already exists within the cache the contents of the message will be replaced with the cached result instead of applying the processors. This component is therefore useful in situations where an expensive set of processors need only be executed periodically. ```yml # Config fields, showing default values label: "" cached: cache: "" # No default (required) skip_on: errored() # No default (optional) key: my_foo_result # No default (required) ttl: "" # No default (optional) processors: [] # No default (required) ``` The format of the data when stored within the cache is a custom and versioned schema chosen to balance performance and storage space. It is therefore not possible to point this processor to a cache that is pre-populated with data that this processor has not created itself. ## [](#examples)Examples ### [](#cached-enrichment)Cached Enrichment In the following example we want to we enrich messages consumed from Kafka with data specific to the origin topic partition, we do this by placing an `http` processor within a `branch`, where the HTTP URL contains interpolation functions with the topic and partition in the path. However, it would be inefficient to make this HTTP request for every single message as the result is consistent for all data of a given topic partition. We can solve this by placing our enrichment call within a `cached` processor where the key contains the topic and partition, resulting in messages that originate from the same topic/partition combination using the cached result of the prior. ```yaml pipeline: processors: - branch: processors: - cached: key: '${! meta("kafka_topic") }-${! meta("kafka_partition") }' cache: foo_cache processors: - mapping: 'root = ""' - http: url: http://example.com/enrichment/${! meta("kafka_topic") }/${! meta("kafka_partition") } verb: GET result_map: 'root.enrichment = this' cache_resources: - label: foo_cache memory: # Disable compaction so that cached items never expire compaction_interval: "" ``` ### [](#periodic-global-enrichment)Periodic Global Enrichment In the following example we enrich all messages with the same data obtained from a static URL with an `http` processor within a `branch`. However, we expect the data from this URL to change roughly every 10 minutes, so we configure a `cached` processor with a static key (since this request is consistent for all messages) and a TTL of `10m`. ```yaml pipeline: processors: - branch: request_map: 'root = ""' processors: - cached: key: static_foo cache: foo_cache ttl: 10m processors: - http: url: http://example.com/get/foo.json verb: GET result_map: 'root.foo = this' cache_resources: - label: foo_cache memory: {} ``` ## [](#fields)Fields ### [](#cache)`cache` The cache resource to read and write processor results from. **Type**: `string` ### [](#key)`key` A key to be resolved for each message, if the key already exists in the cache then the cached result is used, otherwise the processors are applied and the result is cached under this key. The key could be static and therefore apply generally to all messages or it could be an interpolated expression that is potentially unique for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: key: my_foo_result # --- key: ${! this.document.id } # --- key: ${! meta("kafka_key") } # --- key: ${! meta("kafka_topic") } ``` ### [](#processors)`processors[]` The list of processors whose result will be cached. **Type**: `array` ### [](#skip_on)`skip_on` A condition that can be used to skip caching the results from the processors. **Type**: `string` ```yaml # Examples: skip_on: errored() ``` ### [](#ttl)`ttl` An optional expiry period to set for each cache entry. Some caches only have a general TTL and will therefore ignore this setting. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 183: catch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/catch.md --- # catch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: catch latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/catch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/catch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/catch.adoc description: Applies a list of child processors _only_ when a previous processing step has failed. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Applies a list of child processors _only_ when a previous processing step has failed. ```yml # Config fields, showing default values label: "" catch: [] ``` Behaves similarly to the [`for_each`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/for_each/) processor, where a list of child processors are applied to individual messages of a batch. However, processors are only applied to messages that failed a processing step prior to the catch. For example, with the following config: ```yaml pipeline: processors: - resource: foo - catch: - resource: bar - resource: baz ``` If the processor `foo` fails for a particular message, that message will be fed into the processors `bar` and `baz`. Messages that do not fail for the processor `foo` will skip these processors. When messages leave the catch block their fail flags are cleared. This processor is useful for when it’s possible to recover failed messages, or when special actions (such as logging/metrics) are required before dropping them. More information about error handling can be found in [Error Handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). --- # Page 184: cohere_chat **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cohere_chat.md --- # cohere_chat > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: cohere_chat page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/cohere_chat page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/cohere_chat.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/cohere_chat.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-10-16" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Generates responses to messages in a chat conversation, using the [Cohere API](https://docs.cohere.com/docs/chat-api) and external tools. ### Common ```yml processors: label: "" cohere_chat: base_url: https://api.cohere.com api_key: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) max_tokens: "" # No default (optional) temperature: "" # No default (optional) response_format: text json_schema: "" # No default (optional) max_tool_calls: 10 tools: [] ``` ### Advanced ```yml processors: label: "" cohere_chat: base_url: https://api.cohere.com api_key: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) max_tokens: "" # No default (optional) temperature: "" # No default (optional) response_format: text json_schema: "" # No default (optional) schema_registry: url: "" # No default (required) subject: "" # No default (required) refresh_interval: "" # No default (optional) tls: skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} top_p: "" # No default (optional) frequency_penalty: "" # No default (optional) presence_penalty: "" # No default (optional) seed: "" # No default (optional) stop: [] # No default (optional) max_tool_calls: 10 tools: [] ``` This processor sends the contents of user prompts to the Cohere API, which generates responses using all available context, including supplementary data provided by external tools. By default, the processor submits the entire payload of each message as a string, unless you use the `prompt` field to customize it. To learn more about chat completion, see the [Cohere API documentation](https://docs.cohere.com/docs/chat-api). ## [](#fields)Fields ### [](#api_key)`api_key` The API key for the Cohere API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#base_url)`base_url` The base URL to use for API requests. **Type**: `string` **Default**: `[https://api.cohere.com](https://api.cohere.com)` ### [](#frequency_penalty)`frequency_penalty` A number between `-2.0` and `2.0`. Positive values penalize new tokens based on the frequency of their appearance in the text so far. This decreases the model’s likelihood to repeat the same line verbatim. **Type**: `float` ### [](#json_schema)`json_schema` The JSON schema to use when responding in `json_schema` format. To learn more about the JSON schema features supported, see the [Cohere documentation](https://docs.cohere.com/docs/structured-outputs-json). **Type**: `string` ### [](#max_tokens)`max_tokens` The maximum number of tokens to allow in the chat completion. **Type**: `int` ### [](#max_tool_calls)`max_tool_calls` The maximum number of tool calls the model can perform. **Type**: `int` **Default**: `10` ### [](#model)`model` The name of the Cohere large language model (LLM) you want to use. **Type**: `string` ```yaml # Examples: model: command-r-plus # --- model: command-r # --- model: command # --- model: command-light ``` ### [](#presence_penalty)`presence_penalty` A number between `-2.0` and `2.0`. Positive values penalize new tokens based on the frequency of their appearance in the text so far. This increases the model’s likelihood to talk about new topics. **Type**: `float` ### [](#prompt)`prompt` The user prompt you want to generate a response for. By default, the processor submits the entire payload as a string. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#response_format)`response_format` Choose the model’s output format. If `json_schema` is specified, then you must also configure a `json_schema` or `schema_registry`. **Type**: `string` **Default**: `text` **Options**: `text`, `json`, `json_schema` ### [](#schema_registry)`schema_registry` The schema registry to dynamically load schemas from when responding in `json_schema` format. Schemas themselves must be in JSON format. To learn more about the JSON schema features supported, see the [Cohere documentation](https://docs.cohere.com/docs/structured-outputs-json). **Type**: `object` ### [](#schema_registry-basic_auth)`schema_registry.basic_auth` Configure basic authentication for requests from this component to your schema registry. **Type**: `object` ### [](#schema_registry-basic_auth-enabled)`schema_registry.basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-basic_auth-password)`schema_registry.basic_auth.password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-basic_auth-username)`schema_registry.basic_auth.username` The username of the account credentials to authenticate as. Used together with `password` for basic authentication. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt)`schema_registry.jwt` (beta) Configure JSON Web Token (JWT) authentication for secure data transmission from your schema registry to this component. This feature is in beta and may change in future releases. **Type**: `object` ### [](#schema_registry-jwt-claims)`schema_registry.jwt.claims` Values used to pass the identity of the authenticated entity to the service provider. In this case, between this component and the schema registry. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-enabled)`schema_registry.jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-jwt-headers)`schema_registry.jwt.headers` The key/value pairs that identify the type of token and signing algorithm. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-private_key_file)`schema_registry.jwt.private_key_file` Path to a file containing the PEM-encoded private key using PKCS#1 or PKCS#8 format. The private key must be compatible with the algorithm specified in the `signing_method` field. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt-signing_method)`schema_registry.jwt.signing_method` The cryptographic algorithm used to sign the JWT token. Supported algorithms include RS256, RS384, RS512, and EdDSA. This algorithm must be compatible with the private key specified in the `private_key_file` field. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth)`schema_registry.oauth` Configure OAuth version 1.0 to give this component authorized access to your schema registry. **Type**: `object` ### [](#schema_registry-oauth-access_token)`schema_registry.oauth.access_token` The value this component can use to gain access to the data in the schema registry. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-access_token_secret)`schema_registry.oauth.access_token_secret` The secret that establishes ownership of the `oauth.access_token` in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_key)`schema_registry.oauth.consumer_key` The value used to identify this component or client to your schema registry. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_secret)`schema_registry.oauth.consumer_secret` The secret that establishes ownership of the consumer key in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-enabled)`schema_registry.oauth.enabled` Whether to enable OAuth version 1.0 authentication for requests to the schema registry. **Type**: `bool` **Default**: `false` ### [](#schema_registry-refresh_interval)`schema_registry.refresh_interval` The refresh rate for fetching the latest schema. If not specified the schema does not refresh. **Type**: `string` ### [](#schema_registry-subject)`schema_registry.subject` The subject name to fetch the schema for. **Type**: `string` ### [](#schema_registry-tls)`schema_registry.tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#schema_registry-tls-client_certs)`schema_registry.tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#schema_registry-tls-client_certs-cert)`schema_registry.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-cert_file)`schema_registry.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key)`schema_registry.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key_file)`schema_registry.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-password)`schema_registry.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#schema_registry-tls-enable_renegotiation)`schema_registry.tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-root_cas)`schema_registry.tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#schema_registry-tls-root_cas_file)`schema_registry.tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#schema_registry-tls-skip_cert_verify)`schema_registry.tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#schema_registry-url)`schema_registry.url` The base URL of the schema registry service. **Type**: `string` ### [](#seed)`seed` If specified, Redpanda Connect makes a best effort to sample deterministically. Repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed. **Type**: `int` ### [](#stop)`stop[]` Specify up to four sequences to stop the API from generating further tokens. **Type**: `array` ### [](#system_prompt)`system_prompt` The system prompt to submit along with the user prompt. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#temperature)`temperature` Choose a sampling temperature between `0` and `2`: - Higher values, such as `0.8` make the output more random. - Lower values, such as `0.2` make the output more focused and deterministic. Redpanda recommends adding a value for this field or `top_p`, but not both. **Type**: `float` ### [](#tools)`tools[]` External tools that the model can invoke, such as functions, APIs, or web browsing. You can define a series of processors that describe these tools, enabling the model to use agent-like behavior to decide when and how to invoke them to enhance response generation. **Type**: `array` **Default**: `[]` ### [](#tools-description)`tools[].description` A description of this tool, the LLM uses this to decide if the tool should be used. **Type**: `string` ### [](#tools-name)`tools[].name` The name of this tool. **Type**: `string` ### [](#tools-parameters)`tools[].parameters` The parameters the LLM needs to provide to invoke this tool. **Type**: `object` ### [](#tools-parameters-properties)`tools[].parameters.properties` The properties for the processor’s input data **Type**: `object` ### [](#tools-parameters-properties-description)`tools[].parameters.properties.description` A description of this parameter. **Type**: `string` ### [](#tools-parameters-properties-enum)`tools[].parameters.properties.enum[]` Specifies that this parameter is an enum and only these specific values should be used. **Type**: `array` **Default**: `[]` ### [](#tools-parameters-properties-type)`tools[].parameters.properties.type` The type of this parameter. **Type**: `string` ### [](#tools-parameters-required)`tools[].parameters.required[]` The required parameters for this pipeline. **Type**: `array` **Default**: `[]` ### [](#tools-processors)`tools[].processors[]` The pipeline to execute when the LLM uses this tool. **Type**: `array` ### [](#top_p)`top_p` An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with `top_p` probability mass. For example, a `top_p` of `0.1` means only the tokens comprising the top 10% probability mass are sampled. Redpanda recommends adding a value for this field or `temperature`, but not both. **Type**: `float` ## [](#example)Example In this pipeline configuration, the Command R+ model executes a number of processors, which make a tool call to retrieve weather data for a specific city. ```yaml input: generate: count: 1 mapping: | root = "What is the weather like in Chicago?" pipeline: processors: - cohere_chat: auth_token: my_cohere_api_token model: command-r-plus prompt: "${!content().string()}" tools: - name: GetWeather description: "Retrieve the weather for a specific city" parameters: required: ["city"] properties: city: type: string description: the city to look up the weather for processors: - http: verb: GET url: 'https://wttr.in/${!this.city}?T' headers: User-Agent: curl/8.11.1 # Returns a text string from the weather website output: stdout: {} ``` --- # Page 185: cohere_embeddings **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cohere_embeddings.md --- # cohere_embeddings > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: cohere_embeddings page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/cohere_embeddings page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/cohere_embeddings.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/cohere_embeddings.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-10-16" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Generates vector embeddings to represent input text, using the [Cohere API](https://docs.cohere.com/docs/embeddings). ```yml # Configuration fields, showing default values label: "" cohere_embeddings: base_url: https://api.cohere.com auth_token: "" # No default (required) model: embed-english-v3.0 # No default (required) text_mapping: "" # No default (optional) input_type: search_document dimensions: "" # No default (optional) ``` This processor sends text strings to your chosen large language model (LLM), which generates vector embeddings for them using the Cohere API. By default, the processor submits the entire payload of each message as a string, unless you use the `text_mapping` field to customize it. To learn more about vector embeddings, see the [Cohere API documentation](https://docs.cohere.com/docs/embeddings). ## [](#examples)Examples ### [](#store-embedding-vectors-in-qdrant)Store embedding vectors in Qdrant Compute embeddings for some generated data and store it within xrefs:component:outputs/qdrant.adoc\[Qdrant\] ```yaml input: generate: interval: 1s mapping: | root = {"text": fake("paragraph")} pipeline: processors: - cohere_embeddings: model: embed-english-v3 api_key: "${COHERE_API_KEY}" text_mapping: "root = this.text" output: qdrant: grpc_host: localhost:6334 collection_name: "example_collection" id: "root = uuid_v4()" vector_mapping: "root = this" ``` ## [](#fields)Fields ### [](#api_key)`api_key` The API key for the Cohere API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#base_url)`base_url` The base URL to use for API requests. **Type**: `string` **Default**: `[https://api.cohere.com](https://api.cohere.com)` ### [](#dimensions)`dimensions` The number of dimensions (numerical values) in each vector embedding generated by this processor. This parameter only supports [`embed-v4.0`](https://docs.cohere.com/v2/docs/embeddings) and newer models. **Type**: `int` ### [](#input_type)`input_type` The type of text input passed to the model. **Type**: `string` **Default**: `search_document` | Option | Summary | | --- | --- | | classification | Used for embeddings passed through a text classifier. | | clustering | Used for the embeddings run through a clustering algorithm. | | search_document | Used for embeddings stored in a vector database for search use-cases. | | search_query | Used for embeddings of search queries run against a vector DB to find relevant documents. | ### [](#model)`model` The name of the Cohere LLM you want to use. **Type**: `string` ```yaml # Examples: model: embed-english-v3.0 # --- model: embed-english-light-v3.0 # --- model: embed-multilingual-v3.0 # --- model: embed-multilingual-light-v3.0 ``` ### [](#text_mapping)`text_mapping` The text you want to generate a vector embedding for. By default, the processor submits the entire payload as a string. **Type**: `string` --- # Page 186: cohere_rerank **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cohere_rerank.md --- # cohere_rerank > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: cohere_rerank latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/cohere_rerank page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/cohere_rerank.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/cohere_rerank.adoc page-git-created-date: "2025-05-19" page-git-modified-date: "2026-05-26" --- Sends document strings to the [Cohere API](https://docs.cohere.com/reference/rerank), which returns them [ranked by their relevance to a specified query](https://docs.cohere.com/docs/rerank-2). The output of this processor is an array of strings, ordered by their relevance to the query. ```yml # Configuration fields, showing default values label: "" cohere_rerank: base_url: https://api.cohere.com api_key: "" # No default (required) model: rerank-v3.5 # No default (required) query: "" # No default (required) documents: "" # No default (required) top_n: 0 max_tokens_per_doc: 4096 ``` ## [](#metadata)Metadata - `relevance_scores`: An array of scores for each input document that indicates how relevant it is to the query. The scores are in the same order as the documents in the input. The higher the score, the more relevant the document. ## [](#examples)Examples ### [](#rerank-some-documents-based-on-a-query)Rerank some documents based on a query Rerank some documents based on a query ```yaml input: generate: interval: 1s mapping: | root = { "query": fake("sentence"), "docs": [fake("paragraph"), fake("paragraph"), fake("paragraph")], } pipeline: processors: - cohere_rerank: model: rerank-v3.5 api_key: "${COHERE_API_KEY}" query: "${!this.query}" documents: "root = this.docs" output: stdout: {} ``` ## [](#fields)Fields ### [](#api_key)`api_key` Your API key for the Cohere API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#base_url)`base_url` The base URL to use for API requests. **Type**: `string` **Default**: `[https://api.cohere.com](https://api.cohere.com)` ### [](#documents)`documents` A list of text strings that are compared to the specified query. For optimal performance: - Send fewer than 1000 documents in a single request - Send structured data in YAML format **Type**: `string` ### [](#max_tokens_per_doc)`max_tokens_per_doc` This processor automatically truncates long documents to the specified number of tokens. **Type**: `int` **Default**: `4096` ### [](#model)`model` The name of the Cohere LLM you want to use. **Type**: `string` ```yaml # Examples: model: rerank-v3.5 ``` ### [](#query)`query` The search query you want to execute. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#top_n)`top_n` The number of documents to return when the query is executed. If set to `0`, all documents are returned. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `0` --- # Page 187: compress **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/compress.md --- # compress > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: compress latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/compress page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/compress.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/compress.adoc description: "Compresses messages according to the selected algorithm. Supported compression algorithms are: [flate gzip lz4 pgzip snappy zlib]." page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Compresses messages according to the selected algorithm. Supported compression algorithms are: \[flate gzip lz4 pgzip snappy zlib\] ```yml # Config fields, showing default values label: "" compress: algorithm: "" # No default (required) level: -1 ``` The 'level' field might not apply to all algorithms. ## [](#fields)Fields ### [](#algorithm)`algorithm` The compression algorithm to use. **Type**: `string` **Options**: `flate`, `gzip`, `lz4`, `pgzip`, `snappy`, `zlib` ### [](#level)`level` The level of compression to use. May not be applicable to all algorithms. **Type**: `int` **Default**: `-1` --- # Page 188: decompress **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/decompress.md --- # decompress > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: decompress latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/decompress page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/decompress.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/decompress.adoc description: "Decompresses messages according to the selected algorithm. Supported decompression algorithms are: [bzip2 flate gzip lz4 pgzip snappy zlib]." page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Decompresses messages according to the selected algorithm. Supported decompression algorithms are: \[bzip2 flate gzip lz4 pgzip snappy zlib\] ```yml # Config fields, showing default values label: "" decompress: algorithm: "" # No default (required) ``` ## [](#fields)Fields ### [](#algorithm)`algorithm` The decompression algorithm to use. **Type**: `string` **Options**: `bzip2`, `flate`, `gzip`, `lz4`, `pgzip`, `snappy`, `zlib` --- # Page 189: dedupe **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/dedupe.md --- # dedupe > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: dedupe latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/dedupe page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/dedupe.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/dedupe.adoc description: Deduplicates messages by storing a key value in a cache using the add operator. If the key already exists within the cache it is dropped. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Deduplicates messages by storing a key value in a cache using the `add` operator. If the key already exists within the cache it is dropped. ```yml # Config fields, showing default values label: "" dedupe: cache: "" # No default (required) key: ${! meta("kafka_key") } # No default (required) drop_on_err: true ``` Caches must be configured as resources, for more information check out the [cache documentation](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/). When using this processor with an output target that might fail you should always wrap the output within an indefinite [`retry`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/retry/) block. This ensures that during outages your messages aren’t reprocessed after failures, which would result in messages being dropped. ## [](#batch-deduplication)Batch deduplication This processor enacts on individual messages only, in order to perform a deduplication on behalf of a batch (or window) of messages instead use the [`cache` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cache/#examples). ## [](#delivery-guarantees)Delivery guarantees Performing deduplication on a stream using a distributed cache voids any at-least-once guarantees that it previously had. This is because the cache will preserve message signatures even if the message fails to leave the Redpanda Connect pipeline, which would cause message loss in the event of an outage at the output sink followed by a restart of the Redpanda Connect instance (or a server crash, etc). This problem can be mitigated by using an in-memory cache and distributing messages to horizontally scaled Redpanda Connect pipelines partitioned by the deduplication key. However, in situations where at-least-once delivery guarantees are important it is worth avoiding deduplication in favour of implement idempotent behavior at the edge of your stream pipelines. ## [](#fields)Fields ### [](#cache)`cache` The [`cache` resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) to target with this processor. **Type**: `string` ### [](#drop_on_err)`drop_on_err` Whether messages should be dropped when the cache returns a general error such as a network issue. **Type**: `bool` **Default**: `true` ### [](#key)`key` An interpolated string yielding the key to deduplicate by for each message. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: key: ${! meta("kafka_key") } # --- key: ${! content().hash("xxhash64") } ``` ## [](#examples)Examples ### [](#deduplicate-based-on-kafka-key)Deduplicate based on Kafka key The following configuration demonstrates a pipeline that deduplicates messages based on the Kafka key. ```yaml pipeline: processors: - dedupe: cache: keycache key: ${! meta("kafka_key") } cache_resources: - label: keycache memory: default_ttl: 60s ``` --- # Page 190: for_each **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/for_each.md --- # for_each > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: for_each latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/for_each page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/for_each.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/for_each.adoc description: A processor that applies a list of child processors to messages of a batch as though they were each a batch of one message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A processor that applies a list of child processors to messages of a batch as though they were each a batch of one message. ```yml # Config fields, showing default values label: "" for_each: [] ``` This is useful for forcing batch wide processors such as [`dedupe`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/dedupe/) or interpolations such as the `value` field of the `metadata` processor to execute on individual message parts of a batch instead. Please note that most processors already process per message of a batch, and this processor is not needed in those cases. --- # Page 191: gcp_bigquery_select **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/gcp_bigquery_select.md --- # gcp_bigquery_select > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_bigquery_select latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/gcp_bigquery_select page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/gcp_bigquery_select.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/gcp_bigquery_select.adoc description: Executes a SELECT query against BigQuery and replaces messages with the rows returned. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a `SELECT` query against BigQuery and replaces messages with the rows returned. ```yml # Config fields, showing default values label: "" gcp_bigquery_select: project: "" # No default (required) credentials_json: "" # No default (optional) table: bigquery-public-data.samples.shakespeare # No default (required) columns: [] # No default (required) where: type = ? and created_at > ? # No default (optional) job_labels: {} args_mapping: root = [ "article", now().ts_format("2006-01-02") ] # No default (optional) prefix: "" # No default (optional) suffix: "" # No default (optional) ``` ## [](#examples)Examples ### [](#word-count)Word count Given a stream of English terms, enrich the messages with the word count from Shakespeare’s public works: ```yaml pipeline: processors: - branch: processors: - gcp_bigquery_select: project: test-project table: bigquery-public-data.samples.shakespeare columns: - word - sum(word_count) as total_count where: word = ? suffix: | GROUP BY word ORDER BY total_count DESC LIMIT 10 args_mapping: root = [ this.term ] result_map: | root.count = this.get("0.total_count") ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of placeholder arguments in the field `where`. **Type**: `string` ```yaml # Examples: args_mapping: root = [ "article", now().ts_format("2006-01-02") ] ``` ### [](#columns)`columns[]` A list of columns to query. **Type**: `array` ### [](#credentials_json)`credentials_json` Base64-encoded Google Service Account credentials in JSON format (optional). Use this field to authenticate with Google Cloud services. For more information about creating service account credentials, see [Google’s service account documentation](https://developers.google.com/workspace/guides/create-credentials#create_credentials_for_a_service_account). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#job_labels)`job_labels` A list of labels to add to the query job. **Type**: `object` **Default**: `{}` ### [](#prefix)`prefix` An optional prefix to prepend to the select query (before SELECT). **Type**: `string` ### [](#project)`project` GCP project where the query job will execute. **Type**: `string` ### [](#suffix)`suffix` An optional suffix to append to the select query. **Type**: `string` ### [](#table)`table` Fully-qualified BigQuery table name to query. **Type**: `string` ```yaml # Examples: table: bigquery-public-data.samples.shakespeare ``` ### [](#where)`where` An optional where clause to add. Placeholder arguments are populated with the `args_mapping` field. Placeholders should always be question marks (`?`). **Type**: `string` ```yaml # Examples: where: type = ? and created_at > ? # --- where: user_id = ? ``` --- # Page 192: gcp_vertex_ai_chat **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/gcp_vertex_ai_chat.md --- # gcp_vertex_ai_chat > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_vertex_ai_chat latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/gcp_vertex_ai_chat page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/gcp_vertex_ai_chat.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/gcp_vertex_ai_chat.adoc description: Generates responses to messages in a chat conversation, using the Vertex AI API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates responses to messages in a chat conversation, using the [Vertex API AI](https://cloud.google.com/vertex-ai/docs/start/introduction-unified-platform). #### Common ```yml processors: label: "" gcp_vertex_ai_chat: project: "" # No default (required) credentials_json: "" # No default (optional) location: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) history: "" # No default (optional) attachment: "" # No default (optional) temperature: "" # No default (optional) max_tokens: "" # No default (optional) response_format: text tools: [] ``` #### Advanced ```yml processors: label: "" gcp_vertex_ai_chat: project: "" # No default (required) credentials_json: "" # No default (optional) location: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) history: "" # No default (optional) attachment: "" # No default (optional) temperature: "" # No default (optional) max_tokens: "" # No default (optional) response_format: text top_p: "" # No default (optional) top_k: "" # No default (optional) stop: [] # No default (optional) presence_penalty: "" # No default (optional) frequency_penalty: "" # No default (optional) max_tool_calls: 10 tools: [] ``` This processor sends prompts to your chosen large language model (LLM) and generates text from the responses, using the Vertex AI API. For more information, see the [Vertex AI documentation](https://cloud.google.com/vertex-ai/docs). ## [](#fields)Fields ### [](#attachment)`attachment` Additional data like an image to send with the prompt to the model. The result of the mapping must be a byte array, and the content type is automatically detected. **Type**: `string` ```yaml # Examples: attachment: root = this.image.decode("base64") # decode base64 encoded image ``` ### [](#credentials_json)`credentials_json` An optional field to set a Google Service Account Credentials JSON. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#frequency_penalty)`frequency_penalty` Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim. **Type**: `float` ### [](#history)`history` Historical messages to include in the chat request. The result of the bloblang query should be an array of objects of the form of \[{"role": "", "content":""}\], where role is "user" or "model". **Type**: `string` ### [](#location)`location` Specify the location of a fine tuned model. For base models, you can omit this field. **Type**: `string` ```yaml # Examples: location: us-central1 ``` ### [](#max_tokens)`max_tokens` The maximum number of output tokens to generate per message. **Type**: `int` ### [](#max_tool_calls)`max_tool_calls` The maximum number of sequential tool calls. **Type**: `int` **Default**: `10` ### [](#model)`model` The name of the LLM to use. For a full list of models, see the [Vertex AI Model Garden](https://console.cloud.google.com/vertex-ai/model-garden). **Type**: `string` ```yaml # Examples: model: gemini-1.5-pro-001 # --- model: gemini-1.5-flash-001 ``` ### [](#presence_penalty)`presence_penalty` Positive values penalize new tokens if they appear in the text already, increasing the model’s likelihood to include new topics. **Type**: `float` ### [](#project)`project` The GCP project ID to use. **Type**: `string` ### [](#prompt)`prompt` The prompt you want to generate a response for. By default, the processor submits the entire payload as a string. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#response_format)`response_format` The format of the generated response. You must also prompt the model to output the appropriate response type. **Type**: `string` **Default**: `text` **Options**: `text`, `json` ### [](#stop)`stop[]` Sets the stop sequences to use. When this pattern is encountered the LLM stops generating text and returns the final response. **Type**: `array` ### [](#system_prompt)`system_prompt` The system prompt to submit to the Vertex AI LLM. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#temperature)`temperature` Controls the randomness of predictions. **Type**: `float` ### [](#tools)`tools[]` The tools to allow the LLM to invoke. This allows building subpipelines that the LLM can choose to invoke to execute agentic-like actions. **Type**: `array` **Default**: `[]` ### [](#tools-description)`tools[].description` A description of this tool, the LLM uses this to decide if the tool should be used. **Type**: `string` ### [](#tools-name)`tools[].name` The name of this tool. **Type**: `string` ### [](#tools-parameters)`tools[].parameters` The parameters the LLM needs to provide to invoke this tool. **Type**: `object` ### [](#tools-parameters-properties)`tools[].parameters.properties` The properties for the processor’s input data **Type**: `object` ### [](#tools-parameters-properties-description)`tools[].parameters.properties.description` A description of this parameter. **Type**: `string` ### [](#tools-parameters-properties-enum)`tools[].parameters.properties.enum[]` Specifies that this parameter is an enum and only these specific values should be used. **Type**: `array` **Default**: `[]` ### [](#tools-parameters-properties-type)`tools[].parameters.properties.type` The type of this parameter. **Type**: `string` ### [](#tools-parameters-required)`tools[].parameters.required[]` The required parameters for this pipeline. **Type**: `array` **Default**: `[]` ### [](#tools-processors)`tools[].processors[]` The pipeline to execute when the LLM uses this tool. **Type**: `array` ### [](#top_k)`top_k` Enables top-k sampling (optional). **Type**: `float` ### [](#top_p)`top_p` Enables nucleus sampling (optional). **Type**: `float` --- # Page 193: gcp_vertex_ai_embeddings **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/gcp_vertex_ai_embeddings.md --- # gcp_vertex_ai_embeddings > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_vertex_ai_embeddings page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/gcp_vertex_ai_embeddings page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/gcp_vertex_ai_embeddings.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/gcp_vertex_ai_embeddings.adoc # Beta release status page-beta: "true" page-git-created-date: "2024-10-16" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- Generates vector embeddings to represent a text string, using the [Vertex AI API](https://cloud.google.com/vertex-ai/generative-ai/docs/embeddings). ```yml # Configuration fields, showing default values label: "" gcp_vertex_ai_embeddings: project: "" # No default (required) credentials_json: "" # No default (optional) location: us-central1 model: text-embedding-004 # No default (required) task_type: RETRIEVAL_DOCUMENT text: "" # No default (optional) output_dimensions: 0 # No default (optional) ``` This processor sends text strings to the Vertex AI API, which generates vector embeddings for them. By default, the processor submits the entire payload of each message as a string, unless you use the `text` field to customize it. For more information, see the [Vertex AI documentation](https://cloud.google.com/vertex-ai/generative-ai/docs/embeddings). ## [](#fields)Fields ### [](#credentials_json)`credentials_json` Set your Google Service Account Credentials as JSON. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#location)`location` The location of the Vertex AI large language model (LLM) that you want to use. **Type**: `string` **Default**: `us-central1` ### [](#model)`model` The name of the LLM to use. For a full list of models, see the [Vertex AI Model Garden](https://console.cloud.google.com/vertex-ai/model-garden). **Type**: `string` ```yaml # Examples: model: text-embedding-004 # --- model: text-multilingual-embedding-002 ``` ### [](#output_dimensions)`output_dimensions` The maximum length of a generated vector embedding. If this value is set, generated embeddings are truncated to this size. **Type**: `int` ### [](#project)`project` The ID of your Google Cloud project. **Type**: `string` ### [](#task_type)`task_type` Use the following options to optimize embeddings that the model generates for specific use cases. **Type**: `string` **Default**: `RETRIEVAL_DOCUMENT` | Option | Summary | | --- | --- | | CLASSIFICATION | optimize for being able classify texts according to preset labels | | CLUSTERING | optimize for clustering texts based on their similarities | | FACT_VERIFICATION | optimize for queries that are proving or disproving a fact such as "apples grow underground" | | QUESTION_ANSWERING | optimize for search proper questions such as "Why is the sky blue?" | | RETRIEVAL_DOCUMENT | optimize for documents that will be searched (also known as a corpus) | | RETRIEVAL_QUERY | optimize for queries such as "What is the best fish recipe?" or "best restaurant in Chicago" | | SEMANTIC_SIMILARITY | optimize for text similarity | ### [](#text)`text` The text you want to generate vector embeddings for. By default, the processor submits the entire payload as a string. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 194: google_drive_download **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/google_drive_download.md --- # google_drive_download > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: google_drive_download latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/google_drive_download page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/google_drive_download.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/google_drive_download.adoc page-git-created-date: "2025-05-19" page-git-modified-date: "2026-05-26" --- Downloads files from Google Drive that contain matching file IDs. Try out the [example pipeline on this page](#example), which downloads all files from your Google Drive. ### Common ```yml processors: label: "" google_drive_download: credentials_json: "" # No default (optional) file_id: "" # No default (required) mime_type: "" # No default (required) shared_drives: false ``` ### Advanced ```yml processors: label: "" google_drive_download: credentials_json: "" # No default (optional) file_id: "" # No default (required) mime_type: "" # No default (required) export_mime_types: application/vnd.google-apps.document: "text/markdown" application/vnd.google-apps.drawing: "image/png" application/vnd.google-apps.presentation: "application/pdf" application/vnd.google-apps.script: "application/vnd.google-apps.script+json" application/vnd.google-apps.spreadsheet: "text/csv" shared_drives: false ``` ## [](#authentication)Authentication By default, this processor uses [Google Application Default Credentials (ADC)](https://cloud.google.com/docs/authentication/application-default-credentials) to authenticate with Google APIs. To set up local ADC authentication, use the following `gcloud` commands: - Authenticate using Application Default Credentials and grant read-only access to your Google Drive. ```bash gcloud auth application-default login --scopes='openid,https://www.googleapis.com/auth/userinfo.email,https://www.googleapis.com/auth/cloud-platform,https://www.googleapis.com/auth/drive.readonly' ``` - Assign a quota project to the Application Default Credentials when using a user account. ```bash gcloud auth application-default set-quota-project ``` Replace the `` placeholder with your Google Cloud project ID To use a service account instead, create a JSON key for the account and add it to the [`credentials_json`](#credentials_json) field. To access Google Drive files using a service account, either: - Explicitly share files with the service account’s email account - Use [domain-wide delegation](https://support.google.com/a/answer/162106) to share all files within a Google Workspace ## [](#fields)Fields ### [](#credentials_json)`credentials_json` The JSON key for your service account (optional). If left empty, Application Default Credentials are used. For more details, see [Authentication](#authentication). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#export_mime_types)`export_mime_types` Maps Google Drive MIME types to [supported file export formats](https://developers.google.com/workspace/drive/api/guides/ref-export-formats). The MIME type is the key, and the export format is the value. **Type**: `object` **Default**: ```yaml application/vnd.google-apps.document: "text/markdown" application/vnd.google-apps.drawing: "image/png" application/vnd.google-apps.presentation: "application/pdf" application/vnd.google-apps.script: "application/vnd.google-apps.script+json" application/vnd.google-apps.spreadsheet: "text/csv" ``` ```yaml # Examples: export_mime_types: application/vnd.google-apps.document: application/pdf application/vnd.google-apps.drawing: application/pdf application/vnd.google-apps.presentation: application/pdf application/vnd.google-apps.spreadsheet: application/pdf # --- export_mime_types: application/vnd.google-apps.document: application/vnd.openxmlformats-officedocument.wordprocessingml.document application/vnd.google-apps.drawing: image/svg+xml application/vnd.google-apps.presentation: application/vnd.openxmlformats-officedocument.presentationml.presentation application/vnd.google-apps.spreadsheet: application/vnd.openxmlformats-officedocument.spreadsheetml.sheet ``` ### [](#file_id)`file_id` The ID of the file to download from Google Drive. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#mime_type)`mime_type` The [MIME type](https://developers.google.com/workspace/drive/api/guides/mime-types) of the file for download. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#shared_drives)`shared_drives` Whether or not to include shared drives. **Type**: `bool` **Default**: `false` ## [](#example)Example This example downloads all files from a Google Drive. ```yaml input: stdin: {} pipeline: processors: - google_drive_search: query: "${!content().string()}" - mutation: 'meta path = this.name' - google_drive_download: file_id: "${!this.id}" mime_type: "${!this.mimeType}" output: file: path: "${!@path}" codec: all-bytes ``` --- # Page 195: google_drive_list_labels **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/google_drive_list_labels.md --- # google_drive_list_labels > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: google_drive_list_labels latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/google_drive_list_labels page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/google_drive_list_labels.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/google_drive_list_labels.adoc description: Lists labels for a file in Google Drive. page-git-created-date: "2025-05-19" page-git-modified-date: "2026-05-26" --- Lists [labels](https://developers.google.com/workspace/drive/api/guides/about-labels) for files on a Google Drive. ```yml # Configuration fields, showing default values label: "" google_drive_list_labels: credentials_json: "" # No default (optional) ``` ## [](#authentication)Authentication By default, this processor uses [Google Application Default Credentials (ADC)](https://cloud.google.com/docs/authentication/application-default-credentials) to authenticate with Google APIs. To set up local ADC authentication, use the following `gcloud` commands: - Authenticate using Application Default Credentials and grant read-only access to your Google Drive. ```bash gcloud auth application-default login --scopes='openid,https://www.googleapis.com/auth/userinfo.email,https://www.googleapis.com/auth/cloud-platform,https://www.googleapis.com/auth/drive.readonly' ``` - Assign a quota project to the Application Default Credentials when using a user account. ```bash gcloud auth application-default set-quota-project ``` Replace the `` placeholder with your Google Cloud project ID To use a service account instead, create a JSON key for the account and add it to the [`credentials_json`](#credentials_json) field. To access Google Drive files using a service account, either: - Explicitly share files with the service account’s email account - Use [domain-wide delegation](https://support.google.com/a/answer/162106) to share all files within a Google Workspace ## [](#fields)Fields ### [](#credentials_json)`credentials_json` The JSON key for your service account (optional). If left empty, Application Default Credentials are used. For more details, see [Authentication](#authentication). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` --- # Page 196: google_drive_search **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/google_drive_search.md --- # google_drive_search > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: google_drive_search latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/google_drive_search page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/google_drive_search.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/google_drive_search.adoc page-git-created-date: "2025-05-19" page-git-modified-date: "2026-05-26" --- Searches Google Drive for files that match a specified query and emits the results as a batch of messages. Each message contains the [metadata of a Google Drive file](https://developers.google.com/workspace/drive/api/reference/rest/v3/files#File). Try out the [example pipeline on this page](#example), which searches for and downloads all Google Drive files that match the specified query. ```yml # Configuration fields, showing default values label: "" google_drive_search: credentials_json: "" # No default (optional) query: "" # No default (required) projection: - id - name - mimeType - size - labelInfo include_label_ids: "" # No default (optional) max_results: 64 ``` ## [](#authentication)Authentication By default, this processor uses [Google Application Default Credentials (ADC)](https://cloud.google.com/docs/authentication/application-default-credentials) to authenticate with Google APIs. To set up local ADC authentication, use the following `gcloud` commands: - Authenticate using Application Default Credentials and grant read-only access to your Google Drive. ```bash gcloud auth application-default login --scopes='openid,https://www.googleapis.com/auth/userinfo.email,https://www.googleapis.com/auth/cloud-platform,https://www.googleapis.com/auth/drive.readonly' ``` - Assign a quota project to the Application Default Credentials when using a user account. ```bash gcloud auth application-default set-quota-project ``` Replace the `` placeholder with your Google Cloud project ID To use a service account instead, create a JSON key for the account and add it to the [`credentials_json`](#credentials_json) field. To access Google Drive files using a service account, either: - Explicitly share files with the service account’s email account - Use [domain-wide delegation](https://support.google.com/a/answer/162106) to share all files within a Google Workspace ## [](#fields)Fields ### [](#credentials_json)`credentials_json` The JSON key for your service account (optional). If left empty, Application Default Credentials are used. For more details, see [Authentication](#authentication). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#include_label_ids)`include_label_ids` A comma delimited list of label IDs to include in the Google Drive search result. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#max_results)`max_results` The maximum number of search results to return. **Type**: `int` **Default**: `64` ### [](#projection)`projection[]` Partial fields to include in the Google Drive search result. **Type**: `array` **Default**: ```yaml - "id" - "name" - "mimeType" - "size" - "labelInfo" ``` ### [](#query)`query` Specify a search query to locate matching files in Google Drive. This field supports: - The same query syntax as the Google Drive UI - [Bloblang interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) for dynamic query generation **Type**: `string` ### [](#shared_drives)`shared_drives` Whether or not to include shared drives in the result. **Type**: `bool` **Default**: `false` ## [](#example)Example This example searches Google Drive for files matching a query and downloads each file to a specified location. It uses the `google_drive_search` processor to perform the search and the [`google_drive_download` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/google_drive_download/) to retrieve the files. ```yaml input: stdin: {} pipeline: processors: - google_drive_search: query: "${!content().string()}" - mutation: 'meta path = this.name' - google_drive_download: file_id: "${!this.id}" mime_type: "${!this.mimeType}" output: file: path: "${!@path}" codec: all-bytes ``` --- # Page 197: group_by_value **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/group_by_value.md --- # group_by_value > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: group_by_value latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/group_by_value page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/group_by_value.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/group_by_value.adoc description: Splits a batch of messages into N batches, where each resulting batch contains a group of messages determined by a function interpolated string evaluated per message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Splits a batch of messages into N batches, where each resulting batch contains a group of messages determined by a [function interpolated string](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) evaluated per message. ```yml # Config fields, showing default values label: "" group_by_value: value: ${! meta("kafka_key") } # No default (required) ``` This allows you to group messages using arbitrary fields within their content or metadata, process them individually, and send them to unique locations as per their group. The functionality of this processor depends on being applied across messages that are batched. You can find out more about batching [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#value)`value` The interpolated string to group based on. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: value: ${! meta("kafka_key") } # --- value: ${! json("foo.bar") }-${! meta("baz") } ``` ## [](#examples)Examples If we were consuming Kafka messages and needed to group them by their key, archive the groups, and send them to S3 with the key as part of the path we could achieve that with the following: ```yaml pipeline: processors: - group_by_value: value: ${! meta("kafka_key") } - archive: format: tar - compress: algorithm: gzip output: aws_s3: bucket: TODO path: docs/${! meta("kafka_key") }/${! count("files") }-${! timestamp_unix_nano() }.tar.gz ``` --- # Page 198: group_by **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/group_by.md --- # group_by > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: group_by latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/group_by page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/group_by.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/group_by.adoc description: Splits a batch of messages into N batches, where each resulting batch contains a group of messages determined by a Bloblang query. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Splits a [batch of messages](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/) into N batches, where each resulting batch contains a group of messages determined by a [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/). ```yml # Config fields, showing default values label: "" group_by: [] # No default (required) ``` Once the groups are established a list of processors are applied to their respective grouped batch, which can be used to label the batch as per their grouping. Messages that do not pass the check of any specified group are placed in their own group. The functionality of this processor depends on being applied across messages that are batched. You can find out more about batching [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#check)`check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message belongs to a given group. **Type**: `string` ```yaml # Examples: check: this.type == "foo" # --- check: this.contents.urls.contains("https://benthos.dev/") # --- check: true ``` ### [](#processors)`processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to execute on the newly formed group. **Type**: `array` **Default**: `[]` ## [](#examples)Examples ### [](#grouped-processing)Grouped Processing Imagine we have a batch of messages that we wish to split into a group of foos and everything else, which should be sent to different output destinations based on those groupings. We also need to send the foos as a tar gzip archive. For this purpose we can use the `group_by` processor with a [`switch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/switch/) output: ```yaml pipeline: processors: - group_by: - check: content().contains("this is a foo") processors: - archive: format: tar - compress: algorithm: gzip - mapping: 'meta grouping = "foo"' output: switch: cases: - check: meta("grouping") == "foo" output: gcp_pubsub: project: foo_prod topic: only_the_foos - output: gcp_pubsub: project: somewhere_else topic: no_foos_here ``` --- # Page 199: http **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/http.md --- # http > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: http latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/http page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/http.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/http.adoc page-git-created-date: "2025-03-04" page-git-modified-date: "2026-05-26" --- Performs a HTTP request using a message batch as the request body, and replaces the original message parts with the body of the response. #### Common ```yml processors: label: "" http: url: "" # No default (required) verb: POST headers: {} rate_limit: "" # No default (optional) timeout: 5s parallel: false ``` #### Advanced ```yml processors: label: "" http: url: "" # No default (required) verb: POST headers: {} metadata: include_prefixes: [] include_patterns: [] dump_request_log_level: "" oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] extract_headers: include_prefixes: [] include_patterns: [] rate_limit: "" # No default (optional) timeout: 5s retry_period: 1s max_retry_backoff: 300s retries: 3 follow_redirects: true backoff_on: - 429 drop_on: [] successful_on: [] proxy_url: "" # No default (optional) disable_http2: false batch_as_multipart: false parallel: false ``` ## [](#rate-limit-requests)Rate limit requests You can use the `rate_limit` field to specify a [rate limit resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/), which restricts the number of requests processed service-wide, regardless of how many components you run in parallel. ## [](#dynamic-url-and-header-settings)Dynamic URL and header settings You can set the [`url`](#url) and [`headers`](#headers) values dynamically using [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#map-payloads-with-the-branch-processor)Map payloads with the branch processor You can use the [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) to transform or encode the payload into a specific request body format, and map the response back into the original payload instead of replacing it entirely. This example uses a [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) to strip the request message into an empty body (`request_map: 'root = ""'`), grab an HTTP payload, and place the result back into the original message at the path `repo.status`: ```yaml pipeline: processors: - branch: request_map: 'root = ""' processors: - http: url: https://hub.docker.com/v2/repositories/jeffail/benthos verb: GET headers: Content-Type: application/json result_map: 'root.repo.status = this' ``` ## [](#response-codes)Response codes HTTP response codes in the 200-299 range indicate a successful response. You can use the [`successful_on`](#successful_on) field to add more success status codes. HTTP status codes in the 300-399 range are redirects. The [`follow_redirects` field](#follow_redirects) determines how these responses are handled. If a request returns a response code that matches an entry in: - The [`backoff_on` field](#backoff_on), the request is retried after increasing intervals. - The [`drop_on` field](#drop_on), the request is immediately treated as a failure. ## [](#add-metadata-to-errors)Add metadata to errors If a request returns an error response code, this processor sets a `http_status_code` metadata field in the resulting message. > 💡 **TIP** > > You can use the [`extract_headers`](#extract_headers) field to define rules for copying headers into messages generated from the response. ## [](#error-handling)Error handling When all retry attempts for a message are exhausted, this processor cancels the attempt. By default, the failed message continues through the pipeline unchanged unless you configure other error-handling. For example, you might want to drop failed messages or route them to a dead letter queue. For more information, see [Error Handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ## [](#fields)Fields ### [](#backoff_on)`backoff_on[]` A list of status codes that indicate a request failure, and trigger retries with an increasing backoff period between attempts. **Type**: `array` **Default**: ```yaml - 429 ``` ### [](#basic_auth)`basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#batch_as_multipart)`batch_as_multipart` When set to `true`, sends all message in a batch as a single request using [RFC1341](https://www.w3.org/Protocols/rfc1341/7_2_Multipart.html). When set to `false`, sends messages in a batch as individual requests. **Type**: `bool` **Default**: `false` ### [](#disable_http2)`disable_http2` Whether to disable HTTP/2. By default, HTTP/2 is enabled. **Type**: `bool` **Default**: `false` ### [](#drop_on)`drop_on[]` A list of status codes that indicate a request failure, where the input should not attempt retries. This helps avoid unnecessary retries for requests that are unlikely to succeed. > 📝 **NOTE** > > In these cases, the _request_ is dropped, but the _message_ that triggered the request is retained. **Type**: `array` **Default**: `[]` ### [](#dump_request_log_level)`dump_request_log_level` EXPERIMENTAL: Set the logging level for the request and response payloads of each HTTP request. **Type**: `string` **Default**: `""` **Options**: `TRACE`, `DEBUG`, `INFO`, `WARN`, `ERROR`, `FATAL`, \`\` ### [](#extract_headers)`extract_headers` Specify which response headers to add to the resulting messages as metadata. Header keys are automatically converted to lowercase before matching, so make sure that your patterns target the lowercase versions of the expected header keys. **Type**: `object` ### [](#extract_headers-include_patterns)`extract_headers.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#extract_headers-include_prefixes)`extract_headers.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#follow_redirects)`follow_redirects` Whether to follow redirects, including all responses with HTTP status codes in the 300-399 range. If set to `false`, the response message includes only the body, status, and headers from the redirect response, and this processor does not make a request to the URL specified in the `Location` header. **Type**: `bool` **Default**: `true` ### [](#headers)`headers` A map of headers to add to the request. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: Content-Type: application/octet-stream traceparent: ${! tracing_span().traceparent } ``` ### [](#jwt)`jwt` (beta) Configure JSON Web Token (JWT) authentication. This feature is in beta and may change in future releases. JWT tokens provide secure, stateless authentication between services. **Type**: `object` ### [](#jwt-claims)`jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` Additional key-value pairs to include in the JWT header (optional). These headers provide extra metadata for JWT processing. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` Path to a file containing the PEM-encoded private key using PKCS#1 or PKCS#8 format. The private key must be compatible with the algorithm specified in the `signing_method` field. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` The cryptographic algorithm used to sign the JWT token. Supported algorithms include RS256, RS384, RS512, and EdDSA. This algorithm must be compatible with the private key specified in the `private_key_file` field. **Type**: `string` **Default**: `""` ### [](#max_retry_backoff)`max_retry_backoff` The maximum period to wait between failed requests. **Type**: `string` **Default**: `300s` ### [](#metadata)`metadata` Specify matching rules that determine which metadata keys should be added to the HTTP request as headers. **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#oauth)`oauth` Configure OAuth version 1.0 authentication for secure API access. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` The value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` The secret that establishes ownership of the `oauth.access_token` in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2)`oauth2` Allows you to specify open authentication using OAuth version 2 and the client credentials token flow. **Type**: `object` ### [](#oauth2-client_key)`oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#oauth2-client_secret)`oauth2.client_secret` The secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth2-enabled)`oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#oauth2-endpoint_params)`oauth2.endpoint_params` A list of endpoint parameters specified as arrays of strings (optional). **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: bar: - woof foo: - meow - quack ``` ### [](#oauth2-scopes)`oauth2.scopes[]` A list of requested permissions (optional). **Type**: `array` **Default**: `[]` ### [](#oauth2-token_url)`oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#parallel)`parallel` When processing batched messages, this field determines whether messages in the batch are sent in parallel. If set to `false`, messages are sent serially. **Type**: `bool` **Default**: `false` ### [](#proxy_url)`proxy_url` A HTTP proxy URL (optional). **Type**: `string` ### [](#rate_limit)`rate_limit` A [rate limit](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) to throttle requests by (optional). **Type**: `string` ### [](#retries)`retries` The maximum number of retry attempts to make. **Type**: `int` **Default**: `3` ### [](#retry_period)`retry_period` The initial period to wait between failed requests before retrying. **Type**: `string` **Default**: `1s` ### [](#successful_on)`successful_on[]` A list of HTTP status codes that should be considered as successful, even if they are not 2XX codes. This is useful for handling cases where non-2XX codes indicate that the request was processed successfully, such as `303 See Other` or `409 Conflict`. By default, all 2XX codes are considered successful unless they are specified in `backoff_on` or `drop_on` fields. **Type**: `array` **Default**: `[]` ### [](#timeout)`timeout` A static timeout to apply to requests. **Type**: `string` **Default**: `5s` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL to connect to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#verb)`verb` A verb to connect with. **Type**: `string` **Default**: `POST` ```yaml # Examples: verb: POST # --- verb: GET # --- verb: DELETE ``` --- # Page 200: insert_part **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/insert_part.md --- # insert_part > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: insert_part latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/insert_part page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/insert_part.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/insert_part.adoc description: Insert a new message into a batch at an index. If the specified index is greater than the length of the existing batch it will be appended to the end. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Insert a new message into a batch at an index. If the specified index is greater than the length of the existing batch it will be appended to the end. ```yml # Config fields, showing default values label: "" insert_part: index: -1 content: "" ``` The index can be negative, and if so the message will be inserted from the end counting backwards starting from -1. E.g. if index = -1 then the new message will become the last of the batch, if index = -2 then the new message will be inserted before the last message, and so on. If the negative index is greater than the length of the existing batch it will be inserted at the beginning. The new message will have metadata copied from the first pre-existing message of the batch. This processor will interpolate functions within the 'content' field, you can find a list of functions [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#content)`content` The content of the message being inserted. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ### [](#index)`index` The index within the batch to insert the message at. **Type**: `int` **Default**: `-1` --- # Page 201: jira **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/jira.md --- # jira > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: jira latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/jira page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/jira.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/jira.adoc description: Queries Jira resources and returns structured data. page-git-created-date: "2025-11-03" page-git-modified-date: "2026-08-11" --- > ⚠️ **WARNING: Deprecated in 4.100.0** > > Deprecated in 4.100.0 > > This component is deprecated and will be removed in the next major version release. To stream Jira issues, comments, or changelog entries, use the [`jira` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/jira/) instead. Queries Jira resources and returns structured data. #### Common ```yaml processors: label: "" jira: username: "" # No default (required) api_token: "" # No default (required) max_results_per_page: 50 base_url: "" # No default (required) timeout: 5s ``` #### Advanced ```yaml processors: label: "" jira: username: "" # No default (required) api_token: "" # No default (required) max_results_per_page: 50 base_url: "" # No default (required) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] proxy_url: "" disable_http2: false tps_limit: 0 tps_burst: 1 backoff: initial_interval: 1s max_interval: 30s max_retries: 3 tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s http: max_idle_conns: 100 max_idle_conns_per_host: 0 max_conns_per_host: 64 idle_conn_timeout: 1m30s tls_handshake_timeout: 10s expect_continue_timeout: 1s response_header_timeout: 0s disable_keep_alives: false disable_compression: false max_response_header_bytes: 1048576 max_response_body_bytes: 10485760 write_buffer_size: 4096 read_buffer_size: 4096 h2: strict_max_concurrent_requests: false max_decoder_header_table_size: 4096 max_encoder_header_table_size: 4096 max_read_frame_size: 16384 max_receive_buffer_per_connection: 1048576 max_receive_buffer_per_stream: 1048576 send_ping_timeout: 0s ping_timeout: 15s write_byte_timeout: 0s access_log_level: "" access_log_body_limit: 0 ``` Executes Jira API queries based on input messages and returns structured results. The processor handles pagination, retries, and field expansion automatically. Supports querying the following Jira resources: - Issues (JQL queries) - Issue transitions - Users - Roles - Project versions - Project categories - Project types - Projects The processor authenticates using basic authentication with username and API token. Input messages should contain valid Jira queries in JSON format. ## [](#fields)Fields ### [](#access_log_body_limit)`access_log_body_limit` Maximum bytes of request/response body to include in logs. 0 to skip body logging. **Type**: `int` **Default**: `0` ### [](#access_log_level)`access_log_level` Log level for HTTP request/response logging. Empty disables logging. **Type**: `string` **Default**: `""` **Options**: `` `, `TRACE ``, `DEBUG`, `INFO`, `WARN`, `ERROR` ### [](#api_token)`api_token` The Jira API token for the specified account. You can generate an API token from your [Atlassian account settings](https://id.atlassian.com/manage-profile/security/api-tokens). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#backoff)`backoff` Adaptive backoff configuration for 429 (Too Many Requests) responses. Always active. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` Initial interval between retries on 429 responses. **Type**: `string` **Default**: `1s` ### [](#backoff-max_interval)`backoff.max_interval` Maximum interval between retries on 429 responses. **Type**: `string` **Default**: `30s` ### [](#backoff-max_retries)`backoff.max_retries` Maximum number of retries on 429 responses. **Type**: `int` **Default**: `3` ### [](#base_url)`base_url` The base URL of the Jira instance (for example, `[https://your-domain.atlassian.net](https://your-domain.atlassian.net)`). **Type**: `string` ### [](#disable_http2)`disable_http2` Disable HTTP/2 and force HTTP/1.1. **Type**: `bool` **Default**: `false` ### [](#http)`http` HTTP transport settings controlling connection pooling, timeouts, and HTTP/2. **Type**: `object` ### [](#http-disable_compression)`http.disable_compression` Disable automatic decompression of gzip responses. **Type**: `bool` **Default**: `false` ### [](#http-disable_keep_alives)`http.disable_keep_alives` Disable HTTP keep-alive connections; each request uses a new connection. **Type**: `bool` **Default**: `false` ### [](#http-expect_continue_timeout)`http.expect_continue_timeout` Maximum time to wait for a server’s 100-continue response before sending the body. 0 means the body is sent immediately. **Type**: `string` **Default**: `1s` ### [](#http-h2)`http.h2` HTTP/2-specific transport settings. Only applied when HTTP/2 is enabled. **Type**: `object` ### [](#http-h2-max_decoder_header_table_size)`http.h2.max_decoder_header_table_size` Upper limit in bytes for the HPACK header table used to decode headers from the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-h2-max_encoder_header_table_size)`http.h2.max_encoder_header_table_size` Upper limit in bytes for the HPACK header table used to encode headers sent to the peer. Must be less than 4 MiB. **Type**: `int` **Default**: `4096` ### [](#http-h2-max_read_frame_size)`http.h2.max_read_frame_size` Largest HTTP/2 frame this endpoint will read. Valid range: 16 KiB to 16 MiB. **Type**: `int` **Default**: `16384` ### [](#http-h2-max_receive_buffer_per_connection)`http.h2.max_receive_buffer_per_connection` Maximum flow-control window size in bytes for data received on a connection. Must be at least 64 KiB and less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-h2-max_receive_buffer_per_stream)`http.h2.max_receive_buffer_per_stream` Maximum flow-control window size in bytes for data received on a single stream. Must be less than 4 MiB. **Type**: `int` **Default**: `1048576` ### [](#http-h2-ping_timeout)`http.h2.ping_timeout` Timeout waiting for a PING response before closing the connection. **Type**: `string` **Default**: `15s` ### [](#http-h2-send_ping_timeout)`http.h2.send_ping_timeout` Idle timeout after which a PING frame is sent to verify connection health. 0 disables health checks. **Type**: `string` **Default**: `0s` ### [](#http-h2-strict_max_concurrent_requests)`http.h2.strict_max_concurrent_requests` When true, new requests block when a connection’s concurrency limit is reached instead of opening a new connection. **Type**: `bool` **Default**: `false` ### [](#http-h2-write_byte_timeout)`http.h2.write_byte_timeout` Timeout for writing data to a connection. The timer resets whenever bytes are written. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-idle_conn_timeout)`http.idle_conn_timeout` How long an idle connection remains in the pool before being closed. 0 disables the timeout. **Type**: `string` **Default**: `1m30s` ### [](#http-max_conns_per_host)`http.max_conns_per_host` Maximum total connections (active + idle) per host. 0 means unlimited. **Type**: `int` **Default**: `64` ### [](#http-max_idle_conns)`http.max_idle_conns` Maximum total number of idle (keep-alive) connections across all hosts. 0 means unlimited. **Type**: `int` **Default**: `100` ### [](#http-max_idle_conns_per_host)`http.max_idle_conns_per_host` Maximum idle connections to keep per host. 0 (the default) uses GOMAXPROCS+1. **Type**: `int` **Default**: `0` ### [](#http-max_response_body_bytes)`http.max_response_body_bytes` Maximum bytes of response body the client will read. The response body is wrapped with a limit reader; reads beyond this cap return EOF. 0 disables the limit. **Type**: `int` **Default**: `10485760` ### [](#http-max_response_header_bytes)`http.max_response_header_bytes` Maximum bytes of response headers to allow. **Type**: `int` **Default**: `1048576` ### [](#http-read_buffer_size)`http.read_buffer_size` Size in bytes of the per-connection read buffer. **Type**: `int` **Default**: `4096` ### [](#http-response_header_timeout)`http.response_header_timeout` Maximum time to wait for response headers after writing the full request. 0 disables the timeout. **Type**: `string` **Default**: `0s` ### [](#http-tls_handshake_timeout)`http.tls_handshake_timeout` Maximum time to wait for a TLS handshake to complete. 0 disables the timeout. **Type**: `string` **Default**: `10s` ### [](#http-write_buffer_size)`http.write_buffer_size` Size in bytes of the per-connection write buffer. **Type**: `int` **Default**: `4096` ### [](#max_results_per_page)`max_results_per_page` The maximum number of results to return per page when calling the Jira API. [Pagination](https://docs.atlassian.com/software/jira/docs/api/REST/9.17.0/#pagination) in the Jira API is zero-based, so the first page starts at `0`. **Type**: `int` **Default**: `50` ### [](#proxy_url)`proxy_url` HTTP proxy URL. Empty string disables proxying. **Type**: `string` **Default**: `""` ### [](#tcp)`tcp` TCP socket configuration. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` HTTP request timeout. **Type**: `string` **Default**: `5s` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tps_burst)`tps_burst` Maximum burst size for rate limiting. **Type**: `int` **Default**: `1` ### [](#tps_limit)`tps_limit` Rate limit in requests per second. 0 disables rate limiting. **Type**: `float` **Default**: `0` ### [](#username)`username` The username or email address of the Jira account. **Type**: `string` ## [](#examples)Examples ### [](#minimal-configuration)Minimal configuration Basic Jira processor setup with required fields only ```yaml pipeline: processors: - jira: base_url: "https://your-domain.atlassian.net" username: "${JIRA_USERNAME}" api_token: "${JIRA_API_TOKEN}" ``` ### [](#full-configuration-with-tuning)Full configuration with tuning Complete configuration with pagination and timeout settings ```yaml pipeline: processors: - jira: base_url: "https://your-domain.atlassian.net" username: "${JIRA_USERNAME}" api_token: "${JIRA_API_TOKEN}" max_results_per_page: 200 timeout: "30s" ``` --- # Page 202: jmespath **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/jmespath.md --- # jmespath > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: jmespath latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/jmespath page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/jmespath.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/jmespath.adoc description: Executes a http://jmespath.org/[JMESPath query] on JSON documents and replaces the message with the resulting document. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a [JMESPath query](http://jmespath.org/) on JSON documents and replaces the message with the resulting document. ```yml # Config fields, showing default values label: "" jmespath: query: "" # No default (required) ``` > 💡 **TIP: Try out Bloblang** > > Try out Bloblang > > For better performance and improved capabilities try native Redpanda Connect mapping with the [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/). ## [](#fields)Fields ### [](#query)`query` The JMESPath query to apply to messages. **Type**: `string` nclude::connect:components:partial$examples/processors/jmespath.adoc\[\] --- # Page 203: jq **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/jq.md --- # jq > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: jq latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/jq page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/jq.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/jq.adoc description: Transforms and filters messages using jq queries. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Transforms and filters messages using jq queries. #### Common ```yml processors: label: "" jq: query: "" # No default (required) ``` #### Advanced ```yml processors: label: "" jq: query: "" # No default (required) raw: false output_raw: false ``` > 💡 **TIP: Try out Bloblang** > > Try out Bloblang > > For better performance and improved capabilities try out native Redpanda Connect mapping with the [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/). The provided query is executed on each message, targeting either the contents as a structured JSON value or as a raw string using the field `raw`, and the message is replaced with the query result. Message metadata is also accessible within the query from the variable `$metadata`. This processor uses the [gojq library](https://github.com/itchyny/gojq), and therefore does not require jq to be installed as a dependency. However, this also means there are some [differences in how these queries are executed](https://github.com/itchyny/gojq#difference-to-jq) versus the jq cli. If the query does not emit any value then the message is filtered, if the query returns multiple values then the resulting message will be an array containing all values. The full query syntax is described in [jq’s documentation](https://stedolan.github.io/jq/manual/). ## [](#error-handling)Error handling Queries can fail, in which case the message remains unchanged, errors are logged, and the message is flagged as having failed, allowing you to use [standard processor error handling patterns](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ## [](#fields)Fields ### [](#output_raw)`output_raw` Whether to output raw text (unquoted) instead of JSON strings when the emitted values are string types. **Type**: `bool` **Default**: `false` ### [](#query)`query` The jq query to filter and transform messages with. **Type**: `string` ### [](#raw)`raw` Whether to process the input as a raw string instead of as JSON. **Type**: `bool` **Default**: `false` ## [](#examples)Examples ### [](#mapping)Mapping When receiving JSON documents of the form: ```json { "locations": [ {"name": "Seattle", "state": "WA"}, {"name": "New York", "state": "NY"}, {"name": "Bellevue", "state": "WA"}, {"name": "Olympia", "state": "WA"} ] } ``` We could collapse the location names from the state of Washington into a field `Cities`: ```json {"Cities": "Bellevue, Olympia, Seattle"} ``` With the following config: ```yaml pipeline: processors: - jq: query: '{Cities: .locations | map(select(.state == "WA").name) | sort | join(", ") }' ``` --- # Page 204: json_schema **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/json_schema.md --- # json_schema > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: json_schema latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/json_schema page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/json_schema.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/json_schema.adoc description: Checks messages against a provided JSONSchema definition but does not change the payload under any circumstances. If a message does not match the schema it can be caught using error handling methods. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Checks messages against a provided JSONSchema definition but does not change the payload under any circumstances. If a message does not match the schema it can be caught using [error handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ```yml # Config fields, showing default values label: "" json_schema: schema: "" # No default (optional) schema_path: "" # No default (optional) ``` Please refer to the [JSON Schema website](https://json-schema.org/) for information and tutorials regarding the syntax of the schema. ## [](#fields)Fields ### [](#schema)`schema` A schema to apply. Use either this or the `schema_path` field. **Type**: `string` ### [](#schema_path)`schema_path` The path of a schema document to apply. Use either this or the `schema` field. **Type**: `string` ## [](#examples)Examples With the following JSONSchema document: ```json { "$id": "https://example.com/person.schema.json", "$schema": "http://json-schema.org/draft-07/schema#", "title": "Person", "type": "object", "properties": { "firstName": { "type": "string", "description": "The person's first name." }, "lastName": { "type": "string", "description": "The person's last name." }, "age": { "description": "Age in years which must be equal to or greater than zero.", "type": "integer", "minimum": 0 } } } ``` And the following Redpanda Connect configuration: ```yaml pipeline: processors: - json_schema: schema_path: "file://path_to_schema.json" - catch: - log: level: ERROR message: "Schema validation failed due to: ${!error()}" - mapping: 'root = deleted()' # Drop messages that fail ``` If a payload being processed looked like: ```json {"firstName":"John","lastName":"Doe","age":-21} ``` Then a log message would appear explaining the fault and the payload would be dropped. --- # Page 205: log **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/log.md --- # log > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: log latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/log page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/log.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/log.adoc description: Prints a log event for each message. Messages always remain unchanged. The log message can be set using function interpolations described in Bloblang queries which allows you to log the contents and metadata of messages. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Prints a log event for each message. Messages always remain unchanged. The log message can be set using function interpolations described in [Bloblang queries](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) which allows you to log the contents and metadata of messages. ```yml # Config fields, showing default values label: "" log: level: INFO fields_mapping: |- # No default (optional) root.reason = "cus I wana" root.id = this.id root.age = this.user.age.number() root.kafka_topic = meta("kafka_topic") message: "" ``` The `level` field determines the log level of the printed events and can be any of the following values: TRACE, DEBUG, INFO, WARN, ERROR. ## [](#structured-fields)Structured fields It’s also possible add custom fields to logs when the format is set to a structured form such as `json` or `logfmt` with the config field [`fields_mapping`](#fields_mapping): ```yaml pipeline: processors: - log: level: DEBUG message: hello world fields_mapping: | root.reason = "cus I wana" root.id = this.id root.age = this.user.age root.kafka_topic = meta("kafka_topic") ``` ## [](#fields)Fields ### [](#fields_mapping)`fields_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that can be used to specify extra fields to add to the log. If log fields are also added with the `fields` field then those values will override matching keys from this mapping. **Type**: `string` ```yaml # Examples: fields_mapping: |- root.reason = "cus I wana" root.id = this.id root.age = this.user.age.number() root.kafka_topic = meta("kafka_topic") ``` ### [](#level)`level` The log level to use. **Type**: `string` **Default**: `INFO` **Options**: `ERROR`, `WARN`, `INFO`, `DEBUG`, `TRACE` ### [](#message)`message` The message to print. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` --- # Page 206: mapping **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping.md --- # mapping > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mapping latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/mapping page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/mapping.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/mapping.adoc description: Executes a Bloblang mapping on messages, creating a new document that replaces (or filters) the original message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) mapping on messages, creating a new document that replaces (or filters) the original message. ```yml # Config fields, showing default values label: "" mapping: "" # No default (required) ``` Bloblang is a powerful language that enables a wide range of mapping, transformation and filtering tasks. For more information, see [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/). If your mapping is large and you’d prefer for it to live in a separate file then you can execute a mapping directly from a file with the expression `from ""`, where the path must be absolute, or relative from the location that Redpanda Connect is executed from. Note: This processor is equivalent to the [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/bloblang/#component-rename) one. The latter will be deprecated in a future release. ## [](#input-document-immutability)Input document immutability Mapping operates by creating an entirely new object during assignments, this has the advantage of treating the original referenced document as immutable and therefore queryable at any stage of your mapping. For example, with the following mapping: ```bloblang root.id = this.id root.invitees = this.invitees.filter(i -> i.mood >= 0.5) root.rejected = this.invitees.filter(i -> i.mood < 0.5) # In: {"id":"party-2024","invitees":[{"name":"Alice","mood":0.8},{"name":"Bob","mood":0.3},{"name":"Carol","mood":0.9}]} ``` Notice that we mutate the value of `invitees` in the resulting document by filtering out objects with a lower mood. However, even after doing so we’re still able to reference the unchanged original contents of this value from the input document in order to populate a second field. Within this mapping we also have the flexibility to reference the mutable mapped document by using the keyword `root` (i.e. `root.invitees`) on the right-hand side instead. Mapping documents is advantageous in situations where the result is a document with a dramatically different shape to the input document, since we are effectively rebuilding the document in its entirety and might as well keep a reference to the unchanged input document throughout. However, in situations where we are only performing minor alterations to the input document, the rest of which is unchanged, it might be more efficient to use the [`mutation` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mutation/) instead. ## [](#error-handling)Error handling Bloblang mappings can fail, in which case the message remains unchanged, errors are logged, and the message is flagged as having failed, allowing you to use [standard processor error handling patterns](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). However, Bloblang itself also provides powerful ways of ensuring your mappings do not fail by specifying desired [fallback behavior](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#error-handling). ## [](#examples)Examples ### [](#mapping)Mapping Given JSON documents containing an array of fans: ```json { "id":"foo", "description":"a show about foo", "fans":[ {"name":"bev","obsession":0.57}, {"name":"grace","obsession":0.21}, {"name":"ali","obsession":0.89}, {"name":"vic","obsession":0.43} ] } ``` We can reduce the documents down to just the ID and only those fans with an obsession score above 0.5, giving us: ```json { "id":"foo", "fans":[ {"name":"bev","obsession":0.57}, {"name":"ali","obsession":0.89} ] } ``` With the following config: ```yaml pipeline: processors: - mapping: | root.id = this.id root.fans = this.fans.filter(fan -> fan.obsession > 0.5) ``` ### [](#more-mapping)More Mapping When receiving JSON documents of the form: ```json { "locations": [ {"name": "Seattle", "state": "WA"}, {"name": "New York", "state": "NY"}, {"name": "Bellevue", "state": "WA"}, {"name": "Olympia", "state": "WA"} ] } ``` We could collapse the location names from the state of Washington into a field `Cities`: ```json {"Cities": "Bellevue, Olympia, Seattle"} ``` With the following config: ```yaml pipeline: processors: - mapping: | root.Cities = this.locations. filter(loc -> loc.state == "WA"). map_each(loc -> loc.name). sort().join(", ") ``` --- # Page 207: metric **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/metric.md --- # metric > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: metric latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/metric page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/metric.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/metric.adoc description: Emit custom metrics by extracting values from messages. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Emit custom metrics by extracting values from messages. ```yml # Config fields, showing default values label: "" metric: type: "" # No default (required) name: "" # No default (required) labels: {} # No default (optional) value: "" ``` This processor works by evaluating an [interpolated field `value`](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) for each message and updating a emitted metric according to the [type](#types). Custom metrics such as these are emitted along with Redpanda Connect internal metrics, where you can customize where metrics are sent, which metric names are emitted and rename them as/when appropriate. For more information see the [metrics docs](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/). ## [](#fields)Fields ### [](#labels)`labels` A map of label names and values that can be used to enrich metrics. Labels are not supported by some metric destinations, in which case the metrics series are combined. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` ```yaml # Examples: labels: topic: ${! meta("kafka_topic") } type: ${! json("doc.type") } ``` ### [](#name)`name` The name of the metric to create, this must be unique across all Redpanda Connect components otherwise it will overwrite those other metrics. **Type**: `string` ### [](#type)`type` The metric [type](#types) to create. **Type**: `string` **Options**: `counter`, `counter_by`, `gauge`, `timing` ### [](#value)`value` For some metric types specifies a value to set, increment. Certain metrics exporters such as Prometheus support floating point values, but those that do not will cast a floating point value into an integer. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` **Default**: `""` ## [](#examples)Examples ### [](#counter)Counter In this example we emit a counter metric called `Foos`, which increments for every message processed, and we label the metric with some metadata about where the message came from and a field from the document that states what type it is. We also configure our metrics to emit to CloudWatch, and explicitly only allow our custom metric and some internal Redpanda Connect metrics to emit. ```yaml pipeline: processors: - metric: name: Foos type: counter labels: topic: ${! meta("kafka_topic") } partition: ${! meta("kafka_partition") } type: ${! json("document.type").or("unknown") } metrics: mapping: | root = if ![ "Foos", "input_received", "output_sent" ].contains(this) { deleted() } aws_cloudwatch: namespace: ProdConsumer ``` ### [](#gauge)Gauge In this example we emit a gauge metric called `FooSize`, which is given a value extracted from JSON messages at the path `foo.size`. We then also configure our Prometheus metric exporter to only emit this custom metric and nothing else. We also label the metric with some metadata. ```yaml pipeline: processors: - metric: name: FooSize type: gauge labels: topic: ${! meta("kafka_topic") } value: ${! json("foo.size") } metrics: mapping: 'if this != "FooSize" { deleted() }' prometheus: {} ``` ## [](#types)Types ### [](#counter-2)`counter` Increments a counter by exactly 1, the contents of `value` are ignored by this type. ### [](#counter_by)`counter_by` If the contents of `value` can be parsed as a positive integer value then the counter is incremented by this value. For example, the following configuration will increment the value of the `count.custom.field` metric by the contents of `field.some.value`: ```yaml pipeline: processors: - metric: type: counter_by name: CountCustomField value: ${!json("field.some.value")} ``` ### [](#gauge-2)`gauge` If the contents of `value` can be parsed as a positive integer value then the gauge is set to this value. For example, the following configuration will set the value of the `gauge.custom.field` metric to the contents of `field.some.value`: ```yaml pipeline: processors: - metric: type: gauge name: GaugeCustomField value: ${!json("field.some.value")} ``` ### [](#timing)`timing` Equivalent to `gauge` where instead the metric is a timing. It is recommended that timing values are recorded in nanoseconds in order to be consistent with standard Redpanda Connect timing metrics, as in some cases these values are automatically converted into other units such as when exporting timings as histograms with Prometheus metrics. --- # Page 208: mongodb **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mongodb.md --- # mongodb > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mongodb latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/mongodb page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/mongodb.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/mongodb.adoc description: Performs operations against MongoDB for each message, allowing you to store or retrieve data within message payloads. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- Performs operations against MongoDB for each message, allowing you to store or retrieve data within message payloads. #### Common ```yml processors: label: "" mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" collection: "" # No default (required) operation: insert-one write_concern: w: majority j: false w_timeout: "" document_map: "" filter_map: "" hint_map: "" upsert: false ``` #### Advanced ```yml processors: label: "" mongodb: url: "" # No default (required) database: "" # No default (required) username: "" password: "" app_name: benthos aws: enabled: false region: "" # No default (optional) session_duration: 1h id: "" # No default (optional) secret: "" # No default (optional) token: "" # No default (optional) role: "" # No default (optional) role_external_id: "" # No default (optional) roles: [] # No default (optional) collection: "" # No default (required) operation: insert-one write_concern: w: majority j: false w_timeout: "" document_map: "" filter_map: "" hint_map: "" upsert: false json_marshal_mode: canonical ``` ## [](#fields)Fields ### [](#app_name)`app_name` The client application name. **Type**: `string` **Default**: `benthos` ### [](#aws)`aws` AWS IAM authentication using the `MONGODB-AWS` mechanism, for example against MongoDB Atlas. When enabled, IAM credentials are used instead of a static username and password. Role-derived session credentials are resolved when the component connects and are re-resolved whenever it reconnects. The `mongodb` processor and cache establish their client once at creation and cannot refresh expiring session credentials, so `role`, `roles` and session tokens are rejected for those components; use the ambient credential chain or long-lived access keys with them. For long-running pipelines, prefer the ambient credential chain (leave keys and roles unset), which the driver refreshes automatically. **Type**: `object` ### [](#aws-enabled)`aws.enabled` Enable AWS IAM authentication using the driver-native `MONGODB-AWS` mechanism. The MongoDB Atlas database user must be created with the AWS IAM authentication type, and connections require TLS. When no static credentials or roles are configured, the ambient AWS credential chain (environment variables, EC2 instance profile, EKS pod role) is used and expiring credentials are refreshed automatically. **Type**: `bool` **Default**: `false` ### [](#aws-id)`aws.id` The ID of credentials to use. **Type**: `string` ### [](#aws-region)`aws.region` The AWS region used when assuming roles (for STS calls). Only used when `role` or `roles` are configured; the ambient and static-key paths ignore it. If no region is specified then the environment default is used. **Type**: `string` ### [](#aws-role)`aws.role` Optional AWS IAM role ARN to assume for authentication. Cannot be combined with `roles`; use the `roles` array instead when chaining multiple roles. **Type**: `string` ### [](#aws-role_external_id)`aws.role_external_id` Optional external ID for the role assumption. Only used with the `role` field, which cannot be combined with `roles`. **Type**: `string` ### [](#aws-roles)`aws.roles[]` Optional array of AWS IAM roles to assume for authentication. Roles can be assumed in sequence, enabling chaining for purposes such as cross-account access. Each role can optionally specify an external ID. Cannot be combined with `role`. **Type**: `array` ### [](#aws-roles-role)`aws.roles[].role` AWS IAM role ARN to assume. **Type**: `string` **Default**: `""` ### [](#aws-roles-role_external_id)`aws.roles[].role_external_id` Optional external ID for the role assumption. **Type**: `string` **Default**: `""` ### [](#aws-secret)`aws.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#aws-session_duration)`aws.session_duration` The duration of the STS session requested when assuming roles. AWS requires at least 15 minutes and caps sessions created through role chaining at one hour. Only used when `role` or `roles` are configured. For long-running pipelines, prefer the ambient credential chain over a fixed session duration, since the driver refreshes ambient credentials automatically as they near expiry. **Type**: `string` **Default**: `1h` ### [](#aws-token)`aws.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#collection)`collection` The name of the target collection. **Type**: `string` ### [](#database)`database` The name of the target MongoDB database. **Type**: `string` ### [](#document_map)`document_map` A Bloblang map that represents a document to store in MongoDB, expressed as [extended JSON in canonical form](https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/). The `document_map` parameter is required for the following database operations: `insert-one`, `replace-one`, `update-one`, and `aggregate`. **Type**: `string` **Default**: `""` ```yaml # Examples: document_map: |- root.a = this.foo root.b = this.bar ``` ### [](#filter_map)`filter_map` A Bloblang map that represents a filter for a MongoDB command, expressed as [extended JSON in canonical form](https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/). The `filter_map` parameter is required for all database operations except `insert-one`. This output uses `filter_map` to find documents for the specified operation. For example, for a `delete-one` operation, the filter map should include the fields required to locate the document for deletion. **Type**: `string` **Default**: `""` ```yaml # Examples: filter_map: |- root.a = this.foo root.b = this.bar ``` ### [](#hint_map)`hint_map` A Bloblang map that represents a hint or index for a MongoDB command to use, expressed as [extended JSON in canonical form](https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/). This map is optional, and is used with all operations except `insert-one`. Define a `hint_map` to improve performance when finding documents in the MongoDB database. **Type**: `string` **Default**: `""` ```yaml # Examples: hint_map: |- root.a = this.foo root.b = this.bar ``` ### [](#json_marshal_mode)`json_marshal_mode` Controls the format of the output message (optional). **Type**: `string` **Default**: `canonical` | Option | Summary | | --- | --- | | canonical | A string format that emphasizes type preservation at the expense of readability and interoperability. That is, conversion from canonical to BSON will generally preserve type information except in certain specific cases. | | relaxed | A string format that emphasizes readability and interoperability at the expense of type preservation. That is, conversion from relaxed format to BSON can lose type information. | ### [](#operation)`operation` The MongoDB database operation to perform. **Type**: `string` **Default**: `insert-one` **Options**: `insert-one`, `delete-one`, `delete-many`, `replace-one`, `update-one`, `find-one`, `aggregate` ### [](#password)`password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#upsert)`upsert` The `upsert` parameter is optional, and only applies for `update-one` and `replace-one` operations. If the filter specified in `filter_map` matches an existing document, this operation updates or replaces the document, otherwise a new document is created. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target MongoDB server. **Type**: `string` ```yaml # Examples: url: mongodb://localhost:27017 ``` ### [](#username)`username` The username required to connect to the database. **Type**: `string` **Default**: `""` ### [](#write_concern)`write_concern` The [write concern settings](https://www.mongodb.com/docs/manual/reference/write-concern/) for the MongoDB connection. **Type**: `object` ### [](#write_concern-j)`write_concern.j` The `j` requests acknowledgement from MongoDB, which is created when write operations are written to the journal. **Type**: `bool` **Default**: `false` ### [](#write_concern-w)`write_concern.w` The `w` requests acknowledgement, which write operations propagate to the specified number of MongoDB instances. **Type**: `string` **Default**: `majority` ### [](#write_concern-w_timeout)`write_concern.w_timeout` The write concern timeout. **Type**: `string` **Default**: `""` --- # Page 209: mutation **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mutation.md --- # mutation > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: mutation latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/mutation page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/mutation.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/mutation.adoc description: Executes a Bloblang mapping and directly transforms the contents of messages, mutating (or deleting) them. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) mapping and directly transforms the contents of messages, mutating (or deleting) them. ```yml # Config fields, showing default values label: "" mutation: "" # No default (required) ``` Bloblang is a powerful language that enables a wide range of mapping, transformation and filtering tasks. For more information, see [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/). If your mapping is large and you’d prefer for it to live in a separate file then you can execute a mapping directly from a file with the expression `from ""`, where the path must be absolute, or relative from the location that Redpanda Connect is executed from. ## [](#input-document-mutability)Input document mutability A mutation is a mapping that transforms input documents directly, this has the advantage of reducing the need to copy the data fed into the mapping. However, this also means that the referenced document is mutable and therefore changes throughout the mapping. For example, with the following Bloblang: ```bloblang root.rejected = this.invitees.filter(i -> i.mood < 0.5) root.invitees = this.invitees.filter(i -> i.mood >= 0.5) # In: {"invitees":[{"name":"Alice","mood":0.8},{"name":"Bob","mood":0.3},{"name":"Carol","mood":0.9}]} ``` Notice that we create a field `rejected` by copying the array field `invitees` and filtering out objects with a high mood. We then overwrite the field `invitees` by filtering out objects with a low mood, resulting in two array fields that are each a subset of the original. If we were to reverse the ordering of these assignments like so: ```bloblang root.invitees = this.invitees.filter(i -> i.mood >= 0.5) root.rejected = this.invitees.filter(i -> i.mood < 0.5) # In: {"invitees":[{"name":"Alice","mood":0.8},{"name":"Bob","mood":0.3},{"name":"Carol","mood":0.9}]} ``` Then the new field `rejected` would be empty as we have already mutated `invitees` to exclude the objects that it would be populated by. We can solve this problem either by carefully ordering our assignments or by capturing the original array using a variable (`let invitees = this.invitees`). Mutations are advantageous over a standard mapping in situations where the result is a document with mostly the same shape as the input document, since we can avoid unnecessarily copying data from the referenced input document. However, in situations where we are creating an entirely new document shape it can be more convenient to use the traditional [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/) instead. ## [](#error-handling)Error handling Bloblang mappings can fail, in which case the error is logged and the message is flagged as having failed, allowing you to use [standard processor error handling patterns](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). However, Bloblang itself also provides powerful ways of ensuring your mappings do not fail by specifying desired [fallback behavior](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#error-handling). ## [](#examples)Examples ### [](#mapping)Mapping Given JSON documents containing an array of fans: ```json { "id":"foo", "description":"a show about foo", "fans":[ {"name":"bev","obsession":0.57}, {"name":"grace","obsession":0.21}, {"name":"ali","obsession":0.89}, {"name":"vic","obsession":0.43} ] } ``` We can reduce the documents down to just the ID and only those fans with an obsession score above 0.5, giving us: ```json { "id":"foo", "fans":[ {"name":"bev","obsession":0.57}, {"name":"ali","obsession":0.89} ] } ``` With the following config: ```yaml pipeline: processors: - mutation: | root.description = deleted() root.fans = this.fans.filter(fan -> fan.obsession > 0.5) ``` ### [](#more-mapping)More Mapping When receiving JSON documents of the form: ```json { "locations": [ {"name": "Seattle", "state": "WA"}, {"name": "New York", "state": "NY"}, {"name": "Bellevue", "state": "WA"}, {"name": "Olympia", "state": "WA"} ] } ``` We could collapse the location names from the state of Washington into a field `Cities`: ```json {"Cities": "Bellevue, Olympia, Seattle"} ``` With the following config: ```yaml pipeline: processors: - mutation: | root.Cities = this.locations. filter(loc -> loc.state == "WA"). map_each(loc -> loc.name). sort().join(", ") ``` --- # Page 210: nats_kv **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/nats_kv.md --- # nats_kv > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats_kv latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/nats_kv page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/nats_kv.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/nats_kv.adoc description: Perform operations on a NATS key-value bucket. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Perform operations on a NATS key-value bucket. #### Common ```yml processors: label: "" nats_kv: urls: [] # No default (required) bucket: "" # No default (required) operation: "" # No default (required) key: "" # No default (required) ``` #### Advanced ```yml processors: label: "" nats_kv: urls: [] # No default (required) max_reconnects: "" # No default (optional) bucket: "" # No default (required) operation: "" # No default (required) key: "" # No default (required) revision: "" # No default (optional) timeout: 5s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) ``` ## [](#kv-operations)KV operations The NATS KV processor supports many KV operations using the [`operation`](#operation) field. Along with `get`, `put`, and `delete`, this processor supports atomic operations like `update` and `create`, as well as utility operations like `purge`, `history`, and `keys`. ## [](#metadata)Metadata This processor adds the following metadata fields to each message, depending on the chosen `operation`: ### [](#get-get_revision)get, get_revision - `nats_kv_key` - `nats_kv_bucket` - `nats_kv_revision` - `nats_kv_delta` - `nats_kv_operation` - `nats_kv_created` ### [](#create-update-delete-purge)create, update, delete, purge - `nats_kv_key` - `nats_kv_bucket` - `nats_kv_revision` - `nats_kv_operation` ### [](#keys)keys - `nats_kv_bucket` ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the corresponding user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#bucket)`bucket` The name of the KV bucket. **Type**: `string` ```yaml # Examples: bucket: my_kv_bucket ``` ### [](#key)`key` The key for each message. Supports [wildcards](https://docs.nats.io/nats-concepts/subjects#wildcards) for the `history` and `keys` operations. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: key: foo # --- key: foo.bar.baz # --- key: foo.* # --- key: foo.> # --- key: foo.${! json("meta.type") } ``` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#operation)`operation` The operation to perform on the KV bucket. **Type**: `string` | Option | Summary | | --- | --- | | create | Adds the key/value pair if it does not exist. Returns an error if it already exists. | | delete | Deletes the key/value pair, but keeps historical values. | | get | Returns the latest value for key. | | get_revision | Returns the value of key for the specified revision. | | history | Returns historical values of key as an array of objects containing the following fields: key, value, bucket, revision, delta, operation, created. | | keys | Returns the keys in the bucket which match the keys_filter as an array of strings. | | purge | Deletes the key/value pair and all historical values. | | put | Places a new value for the key into the store. | | update | Updates the value for key only if the revision matches the latest revision. | ### [](#revision)`revision` The revision of the key to operate on. Used for `get_revision` and `update` operations. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: revision: 42 # --- revision: ${! @nats_kv_revision } ``` ### [](#timeout)`timeout` The maximum period to wait on an operation before aborting and returning an error. **Type**: `string` **Default**: `5s` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 211: nats_request_reply **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/nats_request_reply.md --- # nats_request_reply > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: nats_request_reply latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/nats_request_reply page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/nats_request_reply.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/nats_request_reply.adoc description: Sends a message to a NATS subject and expects a reply, from a NATS subscriber acting as a responder, back. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sends a message to a NATS subject and expects a reply back from a NATS subscriber acting as a responder. #### Common ```yml processors: label: "" nats_request_reply: urls: [] # No default (required) subject: "" # No default (required) headers: {} metadata: include_prefixes: [] include_patterns: [] timeout: 3s ``` #### Advanced ```yml processors: label: "" nats_request_reply: urls: [] # No default (required) max_reconnects: "" # No default (optional) subject: "" # No default (required) inbox_prefix: "" # No default (optional) headers: {} metadata: include_prefixes: [] include_patterns: [] timeout: 3s tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] tls_handshake_first: false auth: nkey_file: "" # No default (optional) nkey: "" # No default (optional) user_credentials_file: "" # No default (optional) user_jwt: "" # No default (optional) user_nkey_seed: "" # No default (optional) user: "" # No default (optional) password: "" # No default (optional) token: "" # No default (optional) ``` ## [](#metadata)Metadata This input adds the following metadata fields to each message: - `nats_subject` - `nats_sequence_stream` - `nats_sequence_consumer` - `nats_num_delivered` - `nats_num_pending` - `nats_domain` - `nats_timestamp_unix_nano` You can access these metadata fields using [function interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ## [](#fields)Fields ### [](#auth)`auth` Optional configuration of NATS authentication parameters. **Type**: `object` ### [](#auth-nkey)`auth.nkey` Your NKey seed or private key for NATS authentication. NKeys provide secure, cryptographic authentication without passwords. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ```yaml # Examples: nkey: UDXU4RCSJNZOIQHZNWXHXORDPRTGNJAHAHFRGZNEEJCPQTT2M7NLCNF4 ``` ### [](#auth-nkey_file)`auth.nkey_file` An optional file containing a NKey seed. **Type**: `string` ```yaml # Examples: nkey_file: ./seed.nk ``` ### [](#auth-password)`auth.password` An optional plain text password (given along with the corresponding user name). > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-token)`auth.token` An optional plain text token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user)`auth.user` An optional plain text user name (given along with the corresponding user password). **Type**: `string` ### [](#auth-user_credentials_file)`auth.user_credentials_file` An optional file containing user credentials which consist of a user JWT and corresponding NKey seed. **Type**: `string` ```yaml # Examples: user_credentials_file: ./user.creds ``` ### [](#auth-user_jwt)`auth.user_jwt` An optional plaintext user JWT to use along with the corresponding user NKey seed. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#auth-user_nkey_seed)`auth.user_nkey_seed` An optional plaintext user NKey seed to use along with the corresponding user JWT. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#headers)`headers` Explicit message headers to add to messages. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `object` **Default**: `{}` ```yaml # Examples: headers: Content-Type: application/json Timestamp: ${!meta("Timestamp")} ``` ### [](#inbox_prefix)`inbox_prefix` Set an explicit inbox prefix for the response subject **Type**: `string` ```yaml # Examples: inbox_prefix: _INBOX_joe ``` ### [](#max_reconnects)`max_reconnects` The maximum number of times to attempt to reconnect to the server. If negative, it will never stop trying to reconnect. **Type**: `int` ### [](#metadata-2)`metadata` Determine which (if any) metadata values should be added to messages as headers. **Type**: `object` ### [](#metadata-include_patterns)`metadata.include_patterns[]` Provide a list of explicit metadata key regular expression (re2) patterns to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_patterns: - .* # --- include_patterns: - _timestamp_unix$ ``` ### [](#metadata-include_prefixes)`metadata.include_prefixes[]` Provide a list of explicit metadata key prefixes to match against. **Type**: `array` **Default**: `[]` ```yaml # Examples: include_prefixes: - foo_ - bar_ # --- include_prefixes: - kafka_ # --- include_prefixes: - content- ``` ### [](#subject)`subject` A subject to write to. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: subject: foo.bar.baz # --- subject: ${! meta("kafka_topic") } # --- subject: foo.${! json("meta.type") } ``` ### [](#timeout)`timeout` A duration string is a possibly signed sequence of decimal numbers, each with optional fraction and a unit suffix, such as 300ms, -1.5h or 2h45m. Valid time units are ns, us (or µs), ms, s, m, h. **Type**: `string` **Default**: `3s` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls_handshake_first)`tls_handshake_first` Whether to perform the initial TLS handshake before sending the NATS INFO protocol message. This is required when connecting to some NATS servers that expect TLS to be established immediately after connection, before any protocol negotiation. **Type**: `bool` **Default**: `false` ### [](#urls)`urls[]` A list of URLs to connect to. If a list item contains commas, it will be expanded into multiple URLs. **Type**: `array` ```yaml # Examples: urls: - "nats://127.0.0.1:4222" # --- urls: - "nats://username:password@127.0.0.1:4222" ``` --- # Page 212: noop **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/noop.md --- # noop > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: noop latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/noop page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/noop.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/noop.adoc description: Noop is a processor that does nothing, the message passes through unchanged. Why? Sometimes doing nothing is the braver option. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Noop is a processor that does nothing, the message passes through unchanged. Why? Sometimes doing nothing is the braver option. ```yml # Config fields, showing default values label: "" noop: {} ``` --- # Page 213: ollama_chat **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/ollama_chat.md --- # ollama_chat > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: ollama_chat latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/ollama_chat page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/ollama_chat.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/ollama_chat.adoc description: Generates responses to messages in a chat conversation, using the Ollama API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > 📝 **NOTE** > > Ollama connectors are currently only available on BYOC GCP clusters. > ⚠️ **CAUTION** > > When Redpanda Connect runs a data pipeline with a Ollama processor in it, Redpanda Cloud deploys a GPU-powered instance for the exclusive use of that pipeline. As pricing is based on resource consumption, this can have cost implications. Generates responses to messages in a chat conversation using the Ollama API and external tools. #### Common ```yml processors: label: "" ollama_chat: model: "" # No default (required) prompt: "" # No default (optional) image: "" # No default (optional) response_format: text max_tokens: "" # No default (optional) temperature: "" # No default (optional) save_prompt_metadata: false history: "" # No default (optional) tools: [] runner: context_size: "" # No default (optional) batch_size: "" # No default (optional) gpu_layers: "" # No default (optional) threads: "" # No default (optional) use_mmap: "" # No default (optional) server_address: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" ollama_chat: model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) image: "" # No default (optional) response_format: text max_tokens: "" # No default (optional) temperature: "" # No default (optional) num_keep: "" # No default (optional) seed: "" # No default (optional) top_k: "" # No default (optional) top_p: "" # No default (optional) repeat_penalty: "" # No default (optional) presence_penalty: "" # No default (optional) frequency_penalty: "" # No default (optional) stop: [] # No default (optional) save_prompt_metadata: false history: "" # No default (optional) max_tool_calls: 3 tools: [] runner: context_size: "" # No default (optional) batch_size: "" # No default (optional) gpu_layers: "" # No default (optional) threads: "" # No default (optional) use_mmap: "" # No default (optional) server_address: "" # No default (optional) cache_directory: "" # No default (optional) download_url: "" # No default (optional) ``` This processor sends prompts to your chosen Ollama large language model (LLM) and generates text from the responses using the Ollama API and external tools. By default, the processor starts and runs a locally-installed Ollama server. Alternatively, to use an already running Ollama server, add your server details to the `server_address` field. You can [download and install Ollama from the Ollama website](https://ollama.com/download). For more information, see the [Ollama documentation](https://github.com/ollama/ollama/tree/main/docs) and [examples](#examples). ## [](#fields)Fields ### [](#cache_directory)`cache_directory` If `server_address` is not set - the directory to download the Ollama binary and use as a model cache. **Type**: `string` ```yaml # Examples: cache_directory: /opt/cache/connect/ollama ``` ### [](#download_url)`download_url` If `server_address` is not set - the URL to download the Ollama binary from. Defaults to the official Ollama GitHub release for this platform. **Type**: `string` ### [](#frequency_penalty)`frequency_penalty` Positive values penalize new tokens based on the frequency of their appearance in the text so far. This decreases the model’s likelihood to repeat the same line verbatim. **Type**: `float` ### [](#history)`history` Include historical messages in a chat request. You must use a Bloblang query to create an array of objects in the form of `[{"role": "", "content":""}]` where: - `role` is the sender of the original messages, either `system`, `user`, `assistant`, or `tool`. - `content` is the text of the original messages. **Type**: `string` ### [](#image)`image` An optional image to submit along with the [`prompt`](#prompt) value. The result is a byte array. **Type**: `string` ```yaml # Examples: image: root = this.image.decode("base64") # decode base64 encoded image ``` ### [](#max_tokens)`max_tokens` The maximum number of tokens to predict and output. Limiting the amount of output means that requests are processed faster and have a fixed limit on the cost. **Type**: `int` ### [](#max_tool_calls)`max_tool_calls` The maximum number of sequential calls you can make to external tools to retrieve additional information to answer a prompt. **Type**: `int` **Default**: `3` ### [](#model)`model` The name of the Ollama LLM to use. For a full list of models, see the [Ollama website](https://ollama.com/models). **Type**: `string` ```yaml # Examples: model: llama3.1 # --- model: gemma2 # --- model: qwen2 # --- model: phi3 ``` ### [](#num_keep)`num_keep` Specify the number of tokens from the initial prompt to retain when the model resets its internal context. By default, this value is set to `4`. Use `-1` to retain all tokens from the initial prompt. **Type**: `int` ### [](#presence_penalty)`presence_penalty` Positive values penalize new tokens if they have appeared in the text so far. This increases the model’s likelihood to talk about new topics. **Type**: `float` ### [](#prompt)`prompt` The prompt you want to generate a response for. By default, the processor submits the entire payload as a string. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#repeat_penalty)`repeat_penalty` Sets how strongly to penalize repetitions. A higher value, for example 1.5, will penalize repetitions more strongly. A lower value, for example 0.9, will be more lenient. **Type**: `float` ### [](#response_format)`response_format` The format of the response the Ollama model generates. If specifying JSON output, then the `prompt` should specify that the output should be in JSON as well. **Type**: `string` **Default**: `text` **Options**: `text`, `json` ### [](#runner)`runner` Options for the model runner that are used when the model is first loaded into memory. **Type**: `object` ### [](#runner-batch_size)`runner.batch_size` The maximum number of requests to process in parallel. **Type**: `int` ### [](#runner-context_size)`runner.context_size` Sets the size of the context window used to generate the next token. Using a larger context window uses more memory and takes longer to process. **Type**: `int` ### [](#runner-gpu_layers)`runner.gpu_layers` This option allows offloading some layers to the GPU for computation. This generally results in increased performance. By default, the runtime decides the number of layers dynamically. **Type**: `int` ### [](#runner-threads)`runner.threads` Set the number of threads to use during generation. For optimal performance, it is recommended to set this value to the number of physical CPU cores your system has. By default, the runtime decides the optimal number of threads. **Type**: `int` ### [](#runner-use_mmap)`runner.use_mmap` Map the model into memory. This is only support on unix systems and allows loading only the necessary parts of the model as needed. **Type**: `bool` ### [](#save_prompt_metadata)`save_prompt_metadata` Set to `true` to save the prompt value to a metadata field (`@prompt`) on the corresponding output message. If you use the `system_prompt` field, its value is also saved to an `@system_prompt` metadata field on each output message. **Type**: `bool` **Default**: `false` ### [](#seed)`seed` Sets the random number seed to use for generation. Setting this to a specific number will make the model generate the same text for the same prompt. **Type**: `int` ```yaml # Examples: seed: 42 ``` ### [](#server_address)`server_address` The address of the Ollama server to use. Leave the field blank and the processor starts and runs a local Ollama server or specify the address of your own local or remote server. **Type**: `string` ```yaml # Examples: server_address: http://127.0.0.1:11434 ``` ### [](#stop)`stop[]` Sets the stop sequences to use. When this pattern is encountered, the LLM stops generating text and returns the final response. **Type**: `array` ### [](#system_prompt)`system_prompt` The system prompt to submit to the Ollama LLM. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#temperature)`temperature` The temperature of the model. Increasing the temperature makes the model answer more creatively. **Type**: `int` ### [](#tools)`tools[]` The external tools the LLM can invoke, such as functions, APIs, or web browsing. You can build a series of processors that include definitions of these tools, and the specified LLM can choose when to invoke them to help answer a prompt. For more information, see [examples](#examples). **Type**: `array` **Default**: `[]` ### [](#tools-description)`tools[].description` A description of this tool, the LLM uses this to decide if the tool should be used. **Type**: `string` ### [](#tools-name)`tools[].name` The name of this tool. **Type**: `string` ### [](#tools-parameters)`tools[].parameters` The parameters the LLM needs to provide to invoke this tool. **Type**: `object` ### [](#tools-parameters-properties)`tools[].parameters.properties` The properties for the processor’s input data **Type**: `object` ### [](#tools-parameters-properties-description)`tools[].parameters.properties.description` A description of this parameter. **Type**: `string` ### [](#tools-parameters-properties-enum)`tools[].parameters.properties.enum[]` Specifies that this parameter is an enum and only these specific values should be used. **Type**: `array` **Default**: `[]` ### [](#tools-parameters-properties-type)`tools[].parameters.properties.type` The type of this parameter. **Type**: `string` ### [](#tools-parameters-required)`tools[].parameters.required[]` The required parameters for this pipeline. **Type**: `array` **Default**: `[]` ### [](#tools-processors)`tools[].processors[]` The pipeline to execute when the LLM uses this tool. **Type**: `array` ### [](#top_k)`top_k` Reduces the probability of generating nonsense. A higher value, for example `100`, will give more diverse answers. A lower value, for example `10`, will be more conservative. **Type**: `int` ### [](#top_p)`top_p` Works together with `top-k`. A higher value, for example 0.95, will lead to more diverse text. A lower value, for example 0.5, will generate more focused and conservative text. **Type**: `float` ## [](#examples)Examples ### [](#use-llava-to-analyze-an-image)Use Llava to analyze an image This example fetches image URLs from stdin and has a multimodal LLM describe the image. ```yaml input: stdin: scanner: lines: {} pipeline: processors: - http: verb: GET url: "${!content().string()}" - ollama_chat: model: llava prompt: "Describe the following image" image: "root = content()" output: stdout: codec: lines ``` ### [](#use-subpipelines-as-tool-calls)Use subpipelines as tool calls This example allows llama3.2 to execute a subpipeline as a tool call to get more data. ```yaml input: generate: count: 1 mapping: | root = "What is the weather like in Chicago?" pipeline: processors: - ollama_chat: model: llama3.2 prompt: "${!content().string()}" tools: - name: GetWeather description: "Retrieve the weather for a specific city" parameters: required: ["city"] properties: city: type: string description: the city to lookup the weather for processors: - http: verb: GET url: 'https://wttr.in/${!this.city}?T' headers: # Spoof curl user-ageent to get a plaintext text User-Agent: curl/8.11.1 output: stdout: {} ``` --- # Page 214: ollama_embeddings **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/ollama_embeddings.md --- # ollama_embeddings > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: ollama_embeddings latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/ollama_embeddings page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/ollama_embeddings.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/ollama_embeddings.adoc description: Generates vector embeddings from text, using the Ollama API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > 📝 **NOTE** > > Ollama connectors are currently only available on BYOC GCP clusters. > ⚠️ **CAUTION** > > When Redpanda Connect runs a data pipeline with a Ollama processor in it, Redpanda Cloud deploys a GPU-powered instance for the exclusive use of that pipeline. As pricing is based on resource consumption, this can have cost implications. Generates vector embeddings from text, using the Ollama API. #### Common ```yml processors: label: "" ollama_embeddings: model: "" # No default (required) text: "" # No default (optional) runner: context_size: "" # No default (optional) batch_size: "" # No default (optional) gpu_layers: "" # No default (optional) threads: "" # No default (optional) use_mmap: "" # No default (optional) server_address: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" ollama_embeddings: model: "" # No default (required) text: "" # No default (optional) runner: context_size: "" # No default (optional) batch_size: "" # No default (optional) gpu_layers: "" # No default (optional) threads: "" # No default (optional) use_mmap: "" # No default (optional) server_address: "" # No default (optional) cache_directory: "" # No default (optional) download_url: "" # No default (optional) ``` This processor sends text to your chosen Ollama large language model (LLM) and creates vector embeddings, using the Ollama API. Vector embeddings are long arrays of numbers that represent values or objects, in this case text. By default, the processor starts and runs a locally installed Ollama server. Alternatively, to use an already running Ollama server, add your server details to the `server_address` field. You can [download and install Ollama from the Ollama website](https://ollama.com/download). For more information, see the [Ollama documentation](https://github.com/ollama/ollama/tree/main/docs). ## [](#fields)Fields ### [](#cache_directory)`cache_directory` If `server_address` is not set - the directory to download the ollama binary and use as a model cache. **Type**: `string` ```yaml # Examples: cache_directory: /opt/cache/connect/ollama ``` ### [](#download_url)`download_url` If `server_address` is not set - the URL to download the ollama binary from. Defaults to the official Ollama GitHub release for this platform. **Type**: `string` ### [](#model)`model` The name of the Ollama LLM to use. For a full list of models, see the [Ollama website](https://ollama.com/models). **Type**: `string` ```yaml # Examples: model: nomic-embed-text # --- model: mxbai-embed-large # --- model: snowflake-artic-embed # --- model: all-minilm ``` ### [](#runner)`runner` Options for the model runner that are used when the model is first loaded into memory. **Type**: `object` ### [](#runner-batch_size)`runner.batch_size` The maximum number of requests to process in parallel. **Type**: `int` ### [](#runner-context_size)`runner.context_size` Sets the size of the context window used to generate the next token. Using a larger context window uses more memory and takes longer to processor. **Type**: `int` ### [](#runner-gpu_layers)`runner.gpu_layers` This option allows offloading some layers to the GPU for computation. This generally results in increased performance. By default, the runtime decides the number of layers dynamically. **Type**: `int` ### [](#runner-threads)`runner.threads` Set the number of threads to use during generation. For optimal performance, it is recommended to set this value to the number of physical CPU cores your system has. By default, the runtime decides the optimal number of threads. **Type**: `int` ### [](#runner-use_mmap)`runner.use_mmap` Map the model into memory. This is only support on unix systems and allows loading only the necessary parts of the model as needed. **Type**: `bool` ### [](#server_address)`server_address` The address of the Ollama server to use. Leave the field blank and the processor starts and runs a local Ollama server or specify the address of your own local or remote server. **Type**: `string` ```yaml # Examples: server_address: http://127.0.0.1:11434 ``` ### [](#text)`text` The text you want to create vector embeddings for. By default, the processor submits the entire payload as a string. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 215: ollama_moderation **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/ollama_moderation.md --- # ollama_moderation > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: ollama_moderation page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/ollama_moderation page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/ollama_moderation.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/ollama_moderation.adoc # Beta release status page-beta: "true" page-git-created-date: "2025-01-28" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- > 📝 **NOTE** > > Ollama connectors are currently only available on BYOC GCP clusters. > ⚠️ **CAUTION** > > When Redpanda Connect runs a data pipeline with a Ollama processor in it, Redpanda Cloud deploys a GPU-powered instance for the exclusive use of that pipeline. As pricing is based on resource consumption, this can have cost implications. Generates responses to messages in a chat conversation using the Ollama API, and checks the responses to make sure they do not violate [safety or security standards](https://mlcommons.org/2024/04/mlc-aisafety-v0-5-poc/). #### Common ```yml processors: label: "" ollama_moderation: model: "" # No default (required) prompt: "" # No default (required) response: "" # No default (required) runner: context_size: "" # No default (optional) batch_size: "" # No default (optional) gpu_layers: "" # No default (optional) threads: "" # No default (optional) use_mmap: "" # No default (optional) server_address: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" ollama_moderation: model: "" # No default (required) prompt: "" # No default (required) response: "" # No default (required) runner: context_size: "" # No default (optional) batch_size: "" # No default (optional) gpu_layers: "" # No default (optional) threads: "" # No default (optional) use_mmap: "" # No default (optional) server_address: "" # No default (optional) cache_directory: "" # No default (optional) download_url: "" # No default (optional) ``` This processor checks the safety of responses from your chosen large language model (LLM) using either [Llama Guard 3](https://ollama.com/library/llama-guard3) or [ShieldGemma](https://ollama.com/library/shieldgemma). By default, the processor starts and runs a locally-installed Ollama server. Alternatively, to use an already running Ollama server, add your server details to the `server_address` field. You can [download and install Ollama from the Ollama website](https://ollama.com/download). For more information, see the [Ollama documentation](https://github.com/ollama/ollama/tree/main/docs) and [Examples](#examples). To check the safety of your prompts, see the [`ollama_chat` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/ollama_chat/#examples) documentation. ## [](#fields)Fields ### [](#cache_directory)`cache_directory` If the `server_address` is not set, download the Ollama binary to this directory and use it as a model cache. **Type**: `string` ```yaml # Examples: cache_directory: /opt/cache/connect/ollama ``` ### [](#download_url)`download_url` If `server_address` is not set, download the Ollama binary from this URL. The default value is the official Ollama GitHub release for this platform. **Type**: `string` ### [](#model)`model` The name of the Ollama LLM to use. **Type**: `string` | Option | Summary | | --- | --- | | llama-guard3 | When using llama-guard3, two pieces of metadata is added: @safe with the value of yes or no and the second being @category for the safety category violation. For more information see the Llama Guard 3 Model Card. | | shieldgemma | When using shieldgemma, the model output is a single piece of metadata of @safe with a value of yes or no if the response is not in violation of its defined safety policies. | ```yaml # Examples: model: llama-guard3 # --- model: shieldgemma ``` ### [](#prompt)`prompt` The prompt you used to generate a response from an LLM. If you’re using the `ollama_chat` processor, you can set the `save_prompt_metadata` field to save the contents of your prompts. You can then run them through `ollama_moderation` processor to check the model responses for safety. For more details, see [Examples](#examples). You can also check the safety of your prompts. For more information, see the [`ollama_chat` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/ollama_chat/#examples) documentation. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#response)`response` The LLM’s response that you want to check for safety. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#runner)`runner` Options for the model runner that are used when the model is first loaded into memory. **Type**: `object` ### [](#runner-batch_size)`runner.batch_size` The maximum number of requests to process in parallel. **Type**: `int` ### [](#runner-context_size)`runner.context_size` Sets the size of the context window used to generate the next token. Using a larger context window uses more memory and takes longer to process. **Type**: `int` ### [](#runner-gpu_layers)`runner.gpu_layers` Sets the number of layers to offload to the GPU for computation. This generally results in increased performance. By default, the runtime decides the number of layers dynamically. **Type**: `int` ### [](#runner-threads)`runner.threads` Sets the number of threads to use during response generation. For optimal performance, set this value to the number of physical CPU cores your system has. By default, the runtime decides the optimal number of threads. **Type**: `int` ### [](#runner-use_mmap)`runner.use_mmap` Map the model into memory. Set to `true` to load only the necessary parts of the model into memory. This setting is only supported on Unix systems. **Type**: `bool` ### [](#server_address)`server_address` The address of the Ollama server to use. Leave this field blank and the processor starts and runs a local Ollama server, or specify the address of your own local or remote server. **Type**: `string` ```yaml # Examples: server_address: http://127.0.0.1:11434 ``` ## [](#examples)Examples ### [](#use-llama-guard-3-classify-a-llm-response)Use Llama Guard 3 classify a LLM response This example uses Llama Guard 3 to check if another model responded with a safe or unsafe content. ```yaml input: stdin: scanner: lines: {} pipeline: processors: - ollama_chat: model: llava prompt: "${!content().string()}" save_prompt_metadata: true - ollama_moderation: model: llama-guard3 prompt: "${!@prompt}" response: "${!content().string()}" - mapping: | root.response = content().string() root.is_safe = @safe output: stdout: codec: lines ``` --- # Page 216: openai_chat_completion **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/openai_chat_completion.md --- # openai_chat_completion > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: openai_chat_completion latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/openai_chat_completion page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/openai_chat_completion.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/openai_chat_completion.adoc description: Generates responses to messages in a chat conversation, using the OpenAI API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates responses to messages in a chat conversation, using the OpenAI API and external tools. #### Common ```yml processors: label: "" openai_chat_completion: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) history: "" # No default (optional) image: "" # No default (optional) max_tokens: "" # No default (optional) temperature: "" # No default (optional) user: "" # No default (optional) response_format: text json_schema: name: "" # No default (required) description: "" # No default (optional) schema: "" # No default (required) tools: [] # No default (required) ``` #### Advanced ```yml processors: label: "" openai_chat_completion: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) system_prompt: "" # No default (optional) history: "" # No default (optional) image: "" # No default (optional) max_tokens: "" # No default (optional) temperature: "" # No default (optional) user: "" # No default (optional) response_format: text json_schema: name: "" # No default (required) description: "" # No default (optional) schema: "" # No default (required) schema_registry: url: "" # No default (required) name_prefix: schema_registry_id_ subject: "" # No default (required) refresh_interval: "" # No default (optional) tls: skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} top_p: "" # No default (optional) frequency_penalty: "" # No default (optional) presence_penalty: "" # No default (optional) seed: "" # No default (optional) stop: [] # No default (optional) tools: [] # No default (required) ``` This processor sends user prompts to the OpenAI API, and the specified large language model (LLM) generates responses using all available context, including supplementary data provided by [external tools](#tools). By default, the processor submits the entire payload of each message as a string, unless you use the `prompt` configuration field to customize it. To learn more about chat completion, see the [OpenAI API documentation](https://platform.openai.com/docs/guides/chat-completions), and [Examples](#Examples). ## [](#fields)Fields ### [](#api_key)`api_key` The API secret key for OpenAI API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#frequency_penalty)`frequency_penalty` Specify a number between `-2.0` and `2.0`. Positive values penalize new tokens based on the frequency of their appearance in the text so far. This decreases the model’s likelihood to repeat the same line verbatim. **Type**: `float` ### [](#history)`history` Include messages from a prior conversation. You must use a Bloblang query to create an array of objects in the form of `[{"role": "user", "content": ""}, {"role":"assistant", "content":""}]` where: - `role` is the sender of the original messages, either `system`, `user`, or `assistant`. - `content` is the text of the original messages. For more information, see [Examples](#Examples). **Type**: `string` ### [](#image)`image` An optional image to submit along with the prompt. The result of the Bloblang mapping must be a byte array. **Type**: `string` ```yaml # Examples: image: root = this.image.decode("base64") # decode base64 encoded image ``` ### [](#json_schema)`json_schema` The JSON schema used by the model when generating responses in `json_schema` format. To learn more about supported JSON schema features, see the [OpenAI documentation](https://platform.openai.com/docs/guides/structured-outputs/supported-schemas). **Type**: `object` ### [](#json_schema-description)`json_schema.description` An optional description, which helps the model understand the schema’s purpose. **Type**: `string` ### [](#json_schema-name)`json_schema.name` The name of the JSON schema to use. **Type**: `string` ### [](#json_schema-schema)`json_schema.schema` The JSON schema for the model to use when generating the output. **Type**: `string` ### [](#max_tokens)`max_tokens` The maximum number of tokens to generate for chat completion. **Type**: `int` ### [](#model)`model` The name of the OpenAI model to use. **Type**: `string` ```yaml # Examples: model: gpt-4o # --- model: gpt-4o-mini # --- model: gpt-4 # --- model: gpt4-turbo ``` ### [](#presence_penalty)`presence_penalty` Specify a number between `-2.0` and `2.0`. Positive values penalize new tokens if they have appeared in the text so far. This increases the model’s likelihood to talk about new topics. **Type**: `float` ### [](#prompt)`prompt` The user prompt for which a response is generated. By default, the processor sends the entire payload as a string unless customized using this field. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#response_format)`response_format` Specify the configured [model’s](#model) output format. If you choose the `json_schema` option, you must also configure a `json_schema` or `schema_registry`. **Type**: `string` **Default**: `text` **Options**: `text`, `json`, `json_schema` ### [](#schema_registry)`schema_registry` The schema registry to dynamically load schemas for model responses in `json_schema` format. Schemas must be in JSON format. To learn more about supported JSON schema features, see the [OpenAI documentation](https://platform.openai.com/docs/guides/structured-outputs/supported-schemas). **Type**: `object` ### [](#schema_registry-basic_auth)`schema_registry.basic_auth` Configure basic authentication for requests from this component to your schema registry. **Type**: `object` ### [](#schema_registry-basic_auth-enabled)`schema_registry.basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-basic_auth-password)`schema_registry.basic_auth.password` The password to use for authentication. Used together with `username` for basic authentication or with encrypted private keys for secure access. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-basic_auth-username)`schema_registry.basic_auth.username` The username of the account credentials to authenticate as. Used together with `password` for basic authentication. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt)`schema_registry.jwt` (beta) Allows you to specify JWT authentication. **Type**: `object` ### [](#schema_registry-jwt-claims)`schema_registry.jwt.claims` Values used to pass the identity of the authenticated entity to the service provider. In this case, between this component and the schema registry. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-enabled)`schema_registry.jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-jwt-headers)`schema_registry.jwt.headers` The key/value pairs that identify the type of token and signing algorithm (optional). **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-private_key_file)`schema_registry.jwt.private_key_file` Path to a file containing the PEM-encoded private key using PKCS#1 or PKCS#8 format. The private key must be compatible with the algorithm specified in the `signing_method` field. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt-signing_method)`schema_registry.jwt.signing_method` The cryptographic algorithm used to sign the JWT token. Supported algorithms include RS256, RS384, RS512, and EdDSA. This algorithm must be compatible with the private key specified in the `private_key_file` field. **Type**: `string` **Default**: `""` ### [](#schema_registry-name_prefix)`schema_registry.name_prefix` A prefix to add to the schema registry name. To form the complete schema registry name, the schema ID is appended as a suffix. **Type**: `string` **Default**: `schema_registry_id_` ### [](#schema_registry-oauth)`schema_registry.oauth` Configure OAuth version 1.0 to give this component authorized access to your schema registry. **Type**: `object` ### [](#schema_registry-oauth-access_token)`schema_registry.oauth.access_token` The value this component can use to gain access to the data in the schema registry. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-access_token_secret)`schema_registry.oauth.access_token_secret` The secret that establishes ownership of the `oauth.access_token` in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_key)`schema_registry.oauth.consumer_key` The value used to identify this component or client to your schema registry. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_secret)`schema_registry.oauth.consumer_secret` The secret that establishes ownership of the consumer key in OAuth 1.0 authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-enabled)`schema_registry.oauth.enabled` Whether to enable OAuth version 1.0 authentication for requests to the schema registry. **Type**: `bool` **Default**: `false` ### [](#schema_registry-refresh_interval)`schema_registry.refresh_interval` How frequently to poll the schema registry for updates. If not specified, the schema does not refresh automatically. **Type**: `string` ### [](#schema_registry-subject)`schema_registry.subject` The subject name used to fetch the schema from the schema registry. **Type**: `string` ### [](#schema_registry-tls)`schema_registry.tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#schema_registry-tls-client_certs)`schema_registry.tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#schema_registry-tls-client_certs-cert)`schema_registry.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-cert_file)`schema_registry.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key)`schema_registry.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key_file)`schema_registry.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-password)`schema_registry.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#schema_registry-tls-enable_renegotiation)`schema_registry.tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-root_cas)`schema_registry.tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#schema_registry-tls-root_cas_file)`schema_registry.tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#schema_registry-tls-skip_cert_verify)`schema_registry.tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#schema_registry-url)`schema_registry.url` The base URL of the schema registry service. **Type**: `string` ### [](#seed)`seed` When set to a specific number, Redpanda Connect attempts to generate consistent responses for requests that use the same prompt, seed, and parameters. **Type**: `int` ### [](#server_address)`server_address` The OpenAI API endpoint to which the processor sends requests. Update the default value to use a different OpenAI-compatible service. **Type**: `string` **Default**: `[https://api.openai.com/v1](https://api.openai.com/v1)` ### [](#stop)`stop[]` Specify up to four stop sequences to use. When the model encounters a stop pattern, it stops generating text and returns the final response. **Type**: `array` ### [](#system_prompt)`system_prompt` The system prompt to submit along with the user prompt. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#temperature)`temperature` Choose a sampling temperature between `0` and `2`: - Higher values, such as `0.8` make the output more random. - Lower values, such as `0.2` make the output more focused and deterministic. Redpanda recommends adding a value for this field or [`top_p`](#top_p), but not both. **Type**: `float` ### [](#tools)`tools[]` External tools the model can invoke, such as functions, APIs, or web browsing. You can build a series of processors that include definitions of these tools, and the specified model can choose when to invoke them to help answer a prompt. For more information, see [Examples](#Examples). > 📝 **NOTE** > > If you don’t want to use external tools, enter an empty array `tools:[]`. **Type**: `array` ### [](#tools-description)`tools[].description` A description of this tool, the LLM uses this to decide if the tool should be used. **Type**: `string` ### [](#tools-name)`tools[].name` The name of this tool. **Type**: `string` ### [](#tools-parameters)`tools[].parameters` The parameters the LLM needs to provide to invoke this tool. **Type**: `object` **Default**: `[]` ### [](#tools-parameters-properties)`tools[].parameters.properties` The properties for the processor’s input data **Type**: `object` ### [](#tools-parameters-properties-description)`tools[].parameters.properties.description` A description of this parameter. **Type**: `string` ### [](#tools-parameters-properties-enum)`tools[].parameters.properties.enum[]` Specifies that this parameter is an enum and only these specific values should be used. **Type**: `array` **Default**: `[]` ### [](#tools-parameters-properties-type)`tools[].parameters.properties.type` The type of this parameter. **Type**: `string` ### [](#tools-parameters-required)`tools[].parameters.required[]` The required parameters for this pipeline. **Type**: `array` **Default**: `[]` ### [](#tools-processors)`tools[].processors[]` The pipeline to execute when the LLM uses this tool. **Type**: `array` ### [](#top_p)`top_p` An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with `top_p` probability mass. For example, a `top_p` of `0.1` means only the tokens comprising the top 10% probability mass are sampled. Redpanda recommends adding a value for this field or `temperature`, but not both. **Type**: `float` ### [](#user)`user` A unique identifier that represents the end-user generating the prompt. This value can help OpenAI monitor and detect [platform abuse](https://openai.com/policies/usage-policies/). This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` nclude::connect:components:partial$examples/processors/openai\_chat\_completion.adoc\[\] --- # Page 217: openai_embeddings **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/openai_embeddings.md --- # openai_embeddings > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: openai_embeddings latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/openai_embeddings page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/openai_embeddings.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/openai_embeddings.adoc description: Generates vector embeddings to represent input text, using the OpenAI API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates vector embeddings to represent input text, using the OpenAI API. ```yml # Config fields, showing default values label: "" openai_embeddings: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: text-embedding-3-large # No default (required) text_mapping: "" # No default (optional) ``` This processor sends text strings to the OpenAI API, which generates vector embeddings. By default, the processor submits the entire payload of each message as a string, unless you use the `text_mapping` configuration field to customize it. To learn more about vector embeddings, see the [OpenAI API documentation](https://platform.openai.com/docs/guides/embeddings). ## [](#fields)Fields ### [](#api_key)`api_key` The API key for OpenAI API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#dimensions)`dimensions` The number of dimensions the resulting output embeddings should have. Only supported in `text-embedding-3` and later models. **Type**: `int` ### [](#model)`model` The name of the OpenAI model to use. **Type**: `string` ```yaml # Examples: model: text-embedding-3-large # --- model: text-embedding-3-small # --- model: text-embedding-ada-002 ``` ### [](#server_address)`server_address` The Open API endpoint that the processor sends requests to. Update the default value to use another OpenAI compatible service. **Type**: `string` **Default**: `[https://api.openai.com/v1](https://api.openai.com/v1)` ### [](#text_mapping)`text_mapping` The text you want to generate a vector embedding for. By default, the processor submits the entire payload as a string. **Type**: `string` --- # Page 218: openai_image_generation **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/openai_image_generation.md --- # openai_image_generation > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: openai_image_generation latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/openai_image_generation page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/openai_image_generation.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/openai_image_generation.adoc description: Generates an image from a text description and other attributes, using OpenAI API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates an image from a text description and other attributes, using OpenAI API. #### Common ```yml processors: label: "" openai_image_generation: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" openai_image_generation: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) prompt: "" # No default (optional) quality: "" # No default (optional) size: "" # No default (optional) style: "" # No default (optional) ``` This processor sends an image description and other attributes, such as image size and quality to the OpenAI API, which generates an image. By default, the processor submits the entire payload of each message as a string, unless you use the `prompt` configuration field to customize it. To learn more about image generation, see the [OpenAI API documentation](https://platform.openai.com/docs/guides/images). ## [](#fields)Fields ### [](#api_key)`api_key` The API key for OpenAI API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#model)`model` The name of the OpenAI model to use. **Type**: `string` ```yaml # Examples: model: dall-e-3 # --- model: dall-e-2 ``` ### [](#prompt)`prompt` A text description of the image you want to generate. The `prompt` field accepts a maximum of 1000 characters for `dall-e-2` and 4000 characters for `dall-e-3`. **Type**: `string` ### [](#quality)`quality` The quality of the image to generate. Use `hd` to create images with finer details and greater consistency across the image. This parameter is only supported for `dall-e-3` models. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: quality: standard # --- quality: hd ``` ### [](#server_address)`server_address` The Open API endpoint that the processor sends requests to. Update the default value to use another OpenAI compatible service. **Type**: `string` **Default**: `[https://api.openai.com/v1](https://api.openai.com/v1)` ### [](#size)`size` The size of the generated image. Choose from `256x256`, `512x512`, or `1024x1024` for `dall-e-2`. Choose from `1024x1024`, `1792x1024`, or `1024x1792` for `dall-e-3` models. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: size: 1024x1024 # --- size: 512x512 # --- size: 1792x1024 # --- size: 1024x1792 ``` ### [](#style)`style` The style of the generated image. Choose from `vivid` or `natural`. Vivid causes the model to lean towards generating hyperreal and dramatic images. Natural causes the model to produce more natural, less hyperreal looking images. This parameter is only supported for `dall-e-3`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: style: vivid # --- style: natural ``` --- # Page 219: openai_speech **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/openai_speech.md --- # openai_speech > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: openai_speech latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/openai_speech page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/openai_speech.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/openai_speech.adoc description: Generates audio from a text description and other attributes, using OpenAI API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates audio from a text description and other attributes, using OpenAI API. #### Common ```yml processors: label: "" openai_speech: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) input: "" # No default (optional) voice: "" # No default (required) ``` #### Advanced ```yml processors: label: "" openai_speech: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) input: "" # No default (optional) voice: "" # No default (required) response_format: "" # No default (optional) ``` This processor sends a text description and other attributes, such as a voice type and format to the OpenAI API, which generates audio. By default, the processor submits the entire payload of each message as a string, unless you use the `input` configuration field to customize it. To learn more about turning text into spoken audio, see the [OpenAI API documentation](https://platform.openai.com/docs/guides/text-to-speech). ## [](#fields)Fields ### [](#api_key)`api_key` The API key for OpenAI API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#input)`input` A text description of the audio you want to generate. The `input` field accepts a maximum of 4096 characters. **Type**: `string` ### [](#model)`model` The name of the OpenAI model to use. **Type**: `string` ```yaml # Examples: model: tts-1 # --- model: tts-1-hd ``` ### [](#response_format)`response_format` The format to generate audio in. Default is `mp3`. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: response_format: mp3 # --- response_format: opus # --- response_format: aac # --- response_format: flac # --- response_format: wav # --- response_format: pcm ``` ### [](#server_address)`server_address` The Open API endpoint that the processor sends requests to. Update the default value to use another OpenAI compatible service. **Type**: `string` **Default**: `[https://api.openai.com/v1](https://api.openai.com/v1)` ### [](#voice)`voice` The type of voice to use when generating the audio. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: voice: alloy # --- voice: echo # --- voice: fable # --- voice: onyx # --- voice: nova # --- voice: shimmer ``` --- # Page 220: openai_transcription **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/openai_transcription.md --- # openai_transcription > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: openai_transcription latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/openai_transcription page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/openai_transcription.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/openai_transcription.adoc description: Generates a transcription of spoken audio in the input language, using the OpenAI API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Generates a transcription of spoken audio in the input language, using the OpenAI API. #### Common ```yml processors: label: "" openai_transcription: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) file: "" # No default (required) ``` #### Advanced ```yml processors: label: "" openai_transcription: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) file: "" # No default (required) language: "" # No default (optional) prompt: "" # No default (optional) ``` This processor sends an audio file object along with the input language to OpenAI API to generate a transcription. By default, the processor submits the entire payload of each message as a string, unless you use the `file` configuration field to customize it. To learn more about audio transcription, see the: [OpenAI API documentation](https://platform.openai.com/docs/guides/speech-to-text). ## [](#fields)Fields ### [](#api_key)`api_key` The API key for OpenAI API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#file)`file` The audio file object (not file name) to transcribe, in one of the following formats: `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, or `webm`. **Type**: `string` ### [](#language)`language` The language of the input audio. Supplying the input language in ISO-639-1 format improves accuracy and latency. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: language: en # --- language: fr # --- language: de # --- language: zh ``` ### [](#model)`model` The name of the OpenAI model to use. **Type**: `string` ```yaml # Examples: model: whisper-1 ``` ### [](#prompt)`prompt` Optional text to guide the model’s style or continue a previous audio segment. The prompt should match the audio language. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#server_address)`server_address` The Open API endpoint that the processor sends requests to. Update the default value to use another OpenAI compatible service. **Type**: `string` **Default**: `[https://api.openai.com/v1](https://api.openai.com/v1)` --- # Page 221: openai_translation **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/openai_translation.md --- # openai_translation > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: openai_translation latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/openai_translation page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/openai_translation.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/openai_translation.adoc description: Translates spoken audio into English, using the OpenAI API. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Translates spoken audio into English, using the OpenAI API. #### Common ```yml processors: label: "" openai_translation: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) file: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" openai_translation: server_address: https://api.openai.com/v1 api_key: "" # No default (required) model: "" # No default (required) file: "" # No default (optional) prompt: "" # No default (optional) ``` This processor sends an audio file object to OpenAI API to generate a translation. By default, the processor submits the entire payload of each message as a string, unless you use the `file` configuration field to customize it. To learn more about translation, see the [OpenAI API documentation](https://platform.openai.com/docs/guides/speech-to-text). ## [](#fields)Fields ### [](#api_key)`api_key` The API key for OpenAI API. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#file)`file` The audio file object (not file name) to translate, in one of the following formats: `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, or `webm`. **Type**: `string` ### [](#model)`model` The name of the OpenAI model to use. **Type**: `string` ```yaml # Examples: model: whisper-1 ``` ### [](#prompt)`prompt` Optional text to guide the model’s style or continue a previous audio segment. The prompt should match the audio language. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#server_address)`server_address` The Open API endpoint that the processor sends requests to. Update the default value to use another OpenAI compatible service. **Type**: `string` **Default**: `[https://api.openai.com/v1](https://api.openai.com/v1)` --- # Page 222: parallel **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/parallel.md --- # parallel > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: parallel latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/parallel page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/parallel.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/parallel.adoc description: A processor that applies a list of child processors to messages of a batch as though they were each a batch of one message (similar to the for_each processor), but where each message is processed in parallel. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A processor that applies a list of child processors to messages of a batch as though they were each a batch of one message (similar to the [`for_each`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/for_each/) processor), but where each message is processed in parallel. ```yml # Config fields, showing default values label: "" parallel: cap: 0 processors: [] # No default (required) ``` The field `cap`, if greater than zero, caps the maximum number of parallel processing threads. The functionality of this processor depends on being applied across messages that are batched. You can find out more about batching in [Message Batching](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#cap)`cap` The maximum number of messages to have processing at a given time. **Type**: `int` **Default**: `0` ### [](#processors)`processors[]` A list of child processors to apply. **Type**: `array` --- # Page 223: parquet_decode **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/parquet_decode.md --- # parquet_decode > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: parquet_decode latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/parquet_decode page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/parquet_decode.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/parquet_decode.adoc description: Decodes Parquet files into a batch of structured messages. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Decodes [Parquet files](https://parquet.apache.org/docs/) into a batch of structured messages. ```yml # Configuration fields, showing default values label: "" parquet_decode: handle_logical_types: v1 ``` ## [](#fields)Fields ### [](#handle_logical_types)`handle_logical_types` Set to `v2` to enable enhanced decoding of logical types, or keep the default value (`v1`) to ignore logical type metadata when decoding values. In Parquet format, logical types are represented using standard physical types along with metadata that provides additional context. For example, UUIDs are stored as a `FIXED_LEN_BYTE_ARRAY` physical type, but the schema metadata identifies them as UUIDs. By enabling `v2`, this processor uses the metadata descriptions of logical types to produce more meaningful values during decoding. > 📝 **NOTE** > > For backward compatibility, this field enables logical-type handling for the specified Parquet format version, and all earlier versions. When creating new pipelines, Redpanda recommends that you use the newest documented version. **Type**: `string` **Default**: `v1` | Option | Summary | | --- | --- | | v1 | No special handling of logical types | | v2 | TIMESTAMP - decodes as an RFC3339 string describing the time. If the isAdjustedToUTC flag is set to true in the parquet file, the time zone will be set to UTC. If it is set to false the time zone will be set to local time.UUID - decodes as a string, i.e. 00112233-4455-6677-8899-aabbccddeeff. | ```yaml # Examples: handle_logical_types: v2 ``` ## [](#examples)Examples ### [](#reading-parquet-files-from-aws-s3)Reading Parquet Files from AWS S3 In this example we consume files from AWS S3 as they’re written by listening onto an SQS queue for upload events. We make sure to use the `to_the_end` scanner which means files are read into memory in full, which then allows us to use a `parquet_decode` processor to expand each file into a batch of messages. Finally, we write the data out to local files as newline delimited JSON. ```yaml input: aws_s3: bucket: TODO prefix: foos/ scanner: to_the_end: {} sqs: url: TODO processors: - parquet_decode: {} output: file: codec: lines path: './foos/${! meta("s3_key") }.jsonl' ``` --- # Page 224: parquet_encode **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/parquet_encode.md --- # parquet_encode > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: parquet_encode latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/parquet_encode page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/parquet_encode.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/parquet_encode.adoc description: Encodes Parquet files from a batch of structured messages. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Encodes [Parquet files](https://parquet.apache.org/docs/) from a batch of structured messages. #### Common ```yml processors: label: "" parquet_encode: schema: [] # No default (optional) schema_metadata: "" default_compression: uncompressed ``` #### Advanced ```yml processors: label: "" parquet_encode: schema: [] # No default (optional) schema_metadata: "" default_compression: uncompressed default_encoding: DELTA_LENGTH_BYTE_ARRAY default_timestamp_unit: NANOSECOND ``` ## [](#fields)Fields ### [](#default_compression)`default_compression` The default compression type to use for fields. **Type**: `string` **Default**: `uncompressed` **Options**: `uncompressed`, `snappy`, `gzip`, `brotli`, `zstd`, `lz4raw` ### [](#default_encoding)`default_encoding` The default encoding type to use for fields. A custom default encoding is only necessary when consuming data with libraries that do not support `DELTA_LENGTH_BYTE_ARRAY`. **Type**: `string` **Default**: `DELTA_LENGTH_BYTE_ARRAY` **Options**: `DELTA_LENGTH_BYTE_ARRAY`, `PLAIN` ### [](#default_timestamp_unit)`default_timestamp_unit` The precision used when encoding TIMESTAMP logical types. The default `NANOSECOND` matches historical behaviour, but `TIMESTAMP(NANOS)` is not readable by Apache Spark (Databricks), AWS Athena or DuckDB; set this to `MICROSECOND` (or `MILLISECOND`) when writing Parquet files intended for consumption by those engines. **Type**: `string` **Default**: `NANOSECOND` **Options**: `NANOSECOND`, `MICROSECOND`, `MILLISECOND` ### [](#schema)`schema[]` Parquet schema. **Type**: `array` ### [](#schema-fields)`schema[].fields[]` A list of child fields. **Type**: `array` ```yaml # Examples: fields: - name: foo type: INT64 - name: bar type: BYTE_ARRAY ``` ### [](#schema-name)`schema[].name` The name of the column. **Type**: `string` ### [](#schema-optional)`schema[].optional` Whether the field is optional. **Type**: `bool` **Default**: `false` ### [](#schema-repeated)`schema[].repeated` Whether the field is repeated. **Type**: `bool` **Default**: `false` ### [](#schema-type)`schema[].type` The type of the column, only applicable for leaf columns with no child fields. Some logical types can be specified here such as UTF8. **Type**: `string` **Options**: `BOOLEAN`, `INT32`, `INT64`, `FLOAT`, `DOUBLE`, `BYTE_ARRAY`, `UTF8`, `TIMESTAMP`, `BSON`, `ENUM`, `JSON`, `UUID` ### [](#schema_metadata)`schema_metadata` Optionally specify a metadata field containing a schema definition to use for encoding instead of a statically defined schema. For batches of messages, the first message’s schema will be applied to all subsequent messages of the batch. **Type**: `string` **Default**: `""` ## [](#examples)Examples ### [](#writing-parquet-files-to-aws-s3)Writing Parquet Files to AWS S3 In this example we use the batching mechanism of an `aws_s3` output to collect a batch of messages in memory, which then converts it to a parquet file and uploads it. ```yaml output: aws_s3: bucket: TODO path: 'stuff/${! timestamp_unix() }-${! uuid_v4() }.parquet' batching: count: 1000 period: 10s processors: - parquet_encode: schema: - name: id type: INT64 - name: weight type: DOUBLE - name: content type: BYTE_ARRAY default_compression: zstd ``` --- # Page 225: parse_log **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/parse_log.md --- # parse_log > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: parse_log latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/parse_log page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/parse_log.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/parse_log.adoc description: Parses common log formats into structured data. This is easier and often much faster than grok. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Parses common log [Formats](#formats) into [structured data](#codecs). #### Common ```yml processors: label: "" parse_log: format: "" # No default (required) ``` #### Advanced ```yml processors: label: "" parse_log: format: "" # No default (required) best_effort: true allow_rfc3339: true default_year: current default_timezone: UTC ``` ## [](#fields)Fields ### [](#allow_rfc3339)`allow_rfc3339` Also accept timestamps in rfc3339 format while parsing. Applicable to format `syslog_rfc3164`. **Type**: `bool` **Default**: `true` ### [](#best_effort)`best_effort` Still returns partially parsed messages even if an error occurs. **Type**: `bool` **Default**: `true` ### [](#default_timezone)`default_timezone` Sets the strategy to decide the timezone for rfc3164 timestamps. Applicable to format `syslog_rfc3164`. This value should follow the [time.LoadLocation](https://golang.org/pkg/time/#LoadLocation) format. **Type**: `string` **Default**: `UTC` ### [](#default_year)`default_year` Sets the strategy used to set the year for rfc3164 timestamps. Applicable to format `syslog_rfc3164`. When set to `current` the current year will be set, when set to an integer that value will be used. Leave this field empty to not set a default year at all. **Type**: `string` **Default**: `current` ### [](#format)`format` A common log [format](#formats) to parse. **Type**: `string` **Options**: `syslog_rfc5424`, `syslog_rfc3164` ## [](#codecs)Codecs Currently the only supported structured data codec is `json`. ## [](#formats)Formats ### [](#syslog_rfc5424)`syslog_rfc5424` Attempts to parse a log following the [Syslog RFC5424](https://tools.ietf.org/html/rfc5424) spec. The resulting structured document may contain any of the following fields: - `message` (string) - `timestamp` (string, RFC3339) - `facility` (int) - `severity` (int) - `priority` (int) - `version` (int) - `hostname` (string) - `procid` (string) - `appname` (string) - `msgid` (string) - `structureddata` (object) ### [](#syslog_rfc3164)`syslog_rfc3164` Attempts to parse a log following the [Syslog rfc3164](https://tools.ietf.org/html/rfc3164) spec. The resulting structured document may contain any of the following fields: - `message` (string) - `timestamp` (string, RFC3339) - `facility` (int) - `severity` (int) - `priority` (int) - `hostname` (string) - `procid` (string) - `appname` (string) - `msgid` (string) --- # Page 226: processors **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/processors.md --- # processors > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: processors latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/processors page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/processors.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/processors.adoc description: A processor grouping several sub-processors. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A processor grouping several sub-processors. ```yml # Config fields, showing default values label: "" processors: [] ``` This processor is useful in situations where you want to collect several processors under a single resource identifier, whether it is for making your configuration easier to read and navigate, or for improving the testability of your configuration. The behavior of child processors will match exactly the behavior they would have under any other processors block. ## [](#examples)Examples ### [](#grouped-processing)Grouped Processing Imagine we have a collection of processors who cover a specific functionality. We could use this processor to group them together and make it easier to read and mock during testing by giving the whole block a label: ```yaml pipeline: processors: - label: my_super_feature processors: - log: message: "Let's do something cool" - archive: format: json_array - mapping: root.items = this ``` --- # Page 227: protobuf **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/protobuf.md --- # protobuf > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: protobuf latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/protobuf page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/protobuf.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/protobuf.adoc description: Performs conversions to or from a protobuf message. This processor uses reflection, meaning conversions can be made directly from the target .proto files. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Handles conversions between JSON documents and protobuf messages using reflection, which allows you to make conversions from or to the target `.proto` files. For more information about JSON mapping of protobuf messages, see [ProtoJSON Format](https://protobuf.dev/programming-guides/json/) and [Examples](#examples). ```yml # Configuration fields, showing default values label: "" protobuf: operator: "" # No default (required) message: "" # No default (required) discard_unknown: false use_proto_names: false import_paths: [] use_enum_numbers: false ``` ## [](#performance-considerations)Performance considerations Processing protobuf messages using reflection is less performant than using generated native code. For scenarios where performance is critical, consider using [Redpanda Connect plugins](https://github.com/benthosdev/benthos-plugin-example). ## [](#operators)Operators ### [](#to_json)`to_json` Converts protobuf messages into a generic JSON structure, which makes it easier to manipulate the contents of the JSON document within Redpanda Connect. ### [](#from_json)`from_json` Attempts to create a target protobuf message from a generic JSON structure. ## [](#fields)Fields ### [](#bsr)`bsr[]` Buf Schema Registry configuration. Either this field or `import_paths` must be populated. Note that this field is an array, and multiple BSR configurations can be provided. **Type**: `array` **Default**: `[]` ### [](#bsr-api_key)`bsr[].api_key` Buf Schema Registry server API key, can be left blank for a public registry. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#bsr-module)`bsr[].module` Module to fetch from a Buf Schema Registry e.g. 'buf.build/exampleco/mymodule'. **Type**: `string` ### [](#bsr-url)`bsr[].url` Buf Schema Registry URL, leave blank to extract from module. **Type**: `string` **Default**: `""` ### [](#bsr-version)`bsr[].version` Version to retrieve from the Buf Schema Registry, leave blank for latest. **Type**: `string` **Default**: `""` ### [](#discard_unknown)`discard_unknown` When set to `true`, the `from_json` operator discards fields that are unknown to the schema. **Type**: `bool` **Default**: `false` ### [](#import_paths)`import_paths[]` A list of directories that contain `.proto` files, including all definitions required for parsing the target message. If left empty, the current directory is used. This processor imports all `.proto` files listed within specified or default directories. **Type**: `array` **Default**: `[]` ### [](#message)`message` The fully-qualified name of the protobuf message to convert from or to JSON. **Type**: `string` ### [](#operator)`operator` The [operator](#operators) to execute. **Type**: `string` **Options**: `to_json`, `from_json`, `decode` ### [](#use_enum_numbers)`use_enum_numbers` When set to `true`, the `to_json` operator deserializes enumeration fields as their numerical values instead of their string names. For example, an enum field with a value of `ENUM_VALUE_ONE` is represented as `1` in the JSON output. **Type**: `bool` **Default**: `false` ### [](#use_proto_names)`use_proto_names` When set to `true`, the `to_json` operator deserializes fields exactly as named in schema file. **Type**: `bool` **Default**: `false` ## [](#examples)Examples ### [](#json-to-protobuf-using-schema-from-disk)JSON to Protobuf using Schema from Disk If we have the following protobuf definition within a directory called `testing/schema`: ```protobuf syntax = "proto3"; package testing; import "google/protobuf/timestamp.proto"; message Person { string first_name = 1; string last_name = 2; string full_name = 3; int32 age = 4; int32 id = 5; // Unique ID number for this person. string email = 6; google.protobuf.Timestamp last_updated = 7; } ``` And a stream of JSON documents of the form: ```json { "firstName": "caleb", "lastName": "quaye", "email": "caleb@myspace.com" } ``` We can convert the documents into protobuf messages with the following config: ```yaml pipeline: processors: - protobuf: operator: from_json message: testing.Person import_paths: [ testing/schema ] ``` ### [](#protobuf-to-json-using-schema-from-disk)Protobuf to JSON using Schema from Disk If we have the following protobuf definition within a directory called `testing/schema`: ```protobuf syntax = "proto3"; package testing; import "google/protobuf/timestamp.proto"; message Person { string first_name = 1; string last_name = 2; string full_name = 3; int32 age = 4; int32 id = 5; // Unique ID number for this person. string email = 6; google.protobuf.Timestamp last_updated = 7; } ``` And a stream of protobuf messages of the type `Person`, we could convert them into JSON documents of the format: ```json { "firstName": "caleb", "lastName": "quaye", "email": "caleb@myspace.com" } ``` With the following config: ```yaml pipeline: processors: - protobuf: operator: to_json message: testing.Person import_paths: [ testing/schema ] ``` ### [](#json-to-protobuf-using-buf-schema-registry)JSON to Protobuf using Buf Schema Registry If we have the following protobuf definition within a BSR module hosted at `buf.build/exampleco/mymodule`: ```protobuf syntax = "proto3"; package testing; import "google/protobuf/timestamp.proto"; message Person { string first_name = 1; string last_name = 2; string full_name = 3; int32 age = 4; int32 id = 5; // Unique ID number for this person. string email = 6; google.protobuf.Timestamp last_updated = 7; } ``` And a stream of JSON documents of the form: ```json { "firstName": "caleb", "lastName": "quaye", "email": "caleb@myspace.com" } ``` We can convert the documents into protobuf messages with the following config: ```yaml pipeline: processors: - protobuf: operator: from_json message: testing.Person bsr: - module: buf.build/exampleco/mymodule api_key: xxx ``` ### [](#protobuf-to-json-using-buf-schema-registry)Protobuf to JSON using Buf Schema Registry If we have the following protobuf definition within a BSR module hosted at `buf.build/exampleco/mymodule`: ```protobuf syntax = "proto3"; package testing; import "google/protobuf/timestamp.proto"; message Person { string first_name = 1; string last_name = 2; string full_name = 3; int32 age = 4; int32 id = 5; // Unique ID number for this person. string email = 6; google.protobuf.Timestamp last_updated = 7; } ``` And a stream of protobuf messages of the type `Person`, we could convert them into JSON documents of the format: ```json { "firstName": "caleb", "lastName": "quaye", "email": "caleb@myspace.com" } ``` With the following config: ```yaml pipeline: processors: - protobuf: operator: to_json message: testing.Person bsr: - module: buf.build/exampleco/mymodule api_key: xxxx ``` --- # Page 228: qdrant **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/qdrant.md --- # qdrant > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: qdrant latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/qdrant page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/qdrant.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/qdrant.adoc page-git-created-date: "2025-05-19" page-git-modified-date: "2026-05-26" --- Query items within a [Qdrant collection](https://qdrant.tech/documentation/concepts/collections/) and filter the returned results. #### Common ```yml processors: label: "" qdrant: grpc_host: "" # No default (required) api_token: "" collection_name: "" # No default (required) vector_mapping: "" # No default (required) filter: "" # No default (optional) payload_fields: [] payload_filter: include limit: 10 ``` #### Advanced ```yml processors: label: "" qdrant: grpc_host: "" # No default (required) api_token: "" tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] collection_name: "" # No default (required) vector_mapping: "" # No default (required) filter: "" # No default (optional) payload_fields: [] payload_filter: include limit: 10 ``` ## [](#fields)Fields ### [](#api_token)`api_token` The Qdrant API token to use for authentication, which defaults to an empty string. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#collection_name)`collection_name` The name of the Qdrant collection you want to query. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#filter)`filter` Specify additional filtering to perform on returned results. Mappings must return [a valid filter](https://qdrant.tech/documentation/concepts/filtering/) using the proto3-encoded form. **Type**: `string` ```yaml # Examples: filter: |- root.must = [ {"has_id":{"has_id":[{"num": 8}, { "uuid":"1234-5678-90ab-cdef" }]}}, {"field":{"key": "city", "match": {"text": "London"}}}, ] # --- filter: |- root.must = [ {"field":{"key": "city", "match": {"text": "London"}}}, ] root.must_not = [ {"field":{"color": "city", "match": {"text": "red"}}}, ] ``` ### [](#grpc_host)`grpc_host` The gRPC host of the Qdrant server. **Type**: `string` ```yaml # Examples: grpc_host: localhost:6334 # --- grpc_host: xyz-example.eu-central.aws.cloud.qdrant.io:6334 ``` ### [](#limit)`limit` The maximum number of points to return from the collection. **Type**: `int` **Default**: `10` ### [](#payload_fields)`payload_fields[]` The fields to include or exclude in returned results. Use this field in combination with `payload_filter`. **Type**: `array` **Default**: `[]` ### [](#payload_filter)`payload_filter` Whether to include or exclude the fields specified in `payload_fields` from the returned results. **Type**: `string` **Default**: `include` | Option | Summary | | --- | --- | | exclude | Exclude the payload fields specified in payload_fields. | | include | Include the payload fields specified in payload_fields. | ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#vector_mapping)`vector_mapping` A mapping to extract search vectors from the returned document. **Type**: `string` ```yaml # Examples: vector_mapping: root = [1.2, 0.5, 0.76] # --- vector_mapping: root = this.vector # --- vector_mapping: root = [[0.352,0.532,0.532,0.234],[0.352,0.532,0.532,0.234]] # --- vector_mapping: root = {"some_sparse": {"indices":[23,325,532],"values":[0.352,0.532,0.532]}} # --- vector_mapping: root = {"some_multi": [[0.352,0.532,0.532,0.234],[0.352,0.532,0.532,0.234]]} # --- vector_mapping: root = {"some_dense": [0.352,0.532,0.532,0.234]} ``` --- # Page 229: rate_limit **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/rate_limit.md --- # rate_limit > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: rate_limit latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/rate_limit page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/rate_limit.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/rate_limit.adoc description: Throttles the throughput of a pipeline according to a specified rate_limit resource. Rate limits are shared across components and therefore apply globally to all processing pipelines. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Throttles the throughput of a pipeline according to a specified [`rate_limit`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/) resource. Rate limits are shared across components and therefore apply globally to all processing pipelines. ```yml # Config fields, showing default values label: "" rate_limit: resource: "" # No default (required) ``` ## [](#fields)Fields ### [](#resource)`resource` The target [`rate_limit` resource](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about/). **Type**: `string` --- # Page 230: redis_script **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/redis_script.md --- # redis_script > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis_script latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/redis_script page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/redis_script.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/redis_script.adoc description: Performs actions against Redis using LUA scripts. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Performs actions against Redis using [LUA scripts](https://redis.io/docs/latest/develop/programmability/eval-intro/). #### Common ```yml processors: label: "" redis_script: url: "" # No default (required) script: "" # No default (required) args_mapping: "" # No default (required) keys_mapping: "" # No default (required) ``` #### Advanced ```yml processors: label: "" redis_script: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] script: "" # No default (required) args_mapping: "" # No default (required) keys_mapping: "" # No default (required) retries: 3 retry_period: 500ms ``` Actions are performed for each message and the message contents are replaced with the result. In order to merge the result into the original message compose this processor within a [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/). ## [](#examples)Examples ### [](#running-a-script)Running a script The following example will use a script execution to get next element from a sorted set and set its score with timestamp unix nano value. ```yaml pipeline: processors: - redis_script: url: TODO script: | local value = redis.call("ZRANGE", KEYS[1], '0', '0') if next(elements) == nil then return '' end redis.call("ZADD", "XX", KEYS[1], ARGV[1], value) return value keys_mapping: 'root = [ meta("key") ]' args_mapping: 'root = [ timestamp_unix_nano() ]' ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of arguments required for the specified Redis script. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.key ] # --- args_mapping: root = [ meta("kafka_key"), "hardcoded_value" ] ``` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#keys_mapping)`keys_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of keys matching in size to the number of arguments required for the specified Redis script. **Type**: `string` ```yaml # Examples: keys_mapping: root = [ this.key ] # --- keys_mapping: root = [ meta("kafka_key"), this.count ] ``` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#retries)`retries` The maximum number of retries before abandoning a request. **Type**: `int` **Default**: `3` ### [](#retry_period)`retry_period` The time to wait before consecutive retry attempts. **Type**: `string` **Default**: `500ms` ### [](#script)`script` A script to use for the target operator. It has precedence over the 'command' field. **Type**: `string` ```yaml # Examples: script: return redis.call('set', KEYS[1], ARGV[1]) ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 231: redis **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/redis.md --- # redis > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/redis page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/redis.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/redis.adoc description: Performs actions against Redis that aren't possible using a cache processor. Actions are performed for each message and the message contents are replaced with the result. In order to merge the result into the original message compose this processor within a branch processor. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Performs actions against Redis that aren’t possible using a [`cache`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cache/) processor. Actions are performed for each message and the message contents are replaced with the result. In order to merge the result into the original message compose this processor within a [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/). #### Common ```yml processors: label: "" redis: url: "" # No default (required) command: "" # No default (optional) args_mapping: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" redis: url: "" # No default (required) kind: simple master: "" client_name: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] command: "" # No default (optional) args_mapping: "" # No default (optional) retries: 3 retry_period: 500ms ``` ## [](#examples)Examples ### [](#querying-cardinality)Querying Cardinality If given payloads containing a metadata field `set_key` it’s possible to query and store the cardinality of the set for each message using a [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) in order to augment rather than replace the message contents: ```yaml pipeline: processors: - branch: processors: - redis: url: TODO command: scard args_mapping: 'root = [ meta("set_key") ]' result_map: 'root.cardinality = this' ``` ### [](#running-total)Running Total If we have JSON data containing number of friends visited during covid 19: ```json {"name":"ash","month":"feb","year":2019,"friends_visited":10} {"name":"ash","month":"apr","year":2019,"friends_visited":-2} {"name":"bob","month":"feb","year":2019,"friends_visited":3} {"name":"bob","month":"apr","year":2019,"friends_visited":1} ``` We can add a field that contains the running total number of friends visited: ```json {"name":"ash","month":"feb","year":2019,"friends_visited":10,"total":10} {"name":"ash","month":"apr","year":2019,"friends_visited":-2,"total":8} {"name":"bob","month":"feb","year":2019,"friends_visited":3,"total":3} {"name":"bob","month":"apr","year":2019,"friends_visited":1,"total":4} ``` Using the `incrby` command: ```yaml pipeline: processors: - branch: processors: - redis: url: TODO command: incrby args_mapping: 'root = [ this.name, this.friends_visited ]' result_map: 'root.total = this' ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of arguments required for the specified Redis command. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.key ] # --- args_mapping: root = [ meta("kafka_key"), this.count ] ``` ### [](#client_name)`client_name` Set the client name for the Redis connection. **Type**: `string` **Default**: `redpanda-connect` ### [](#command)`command` The command to execute. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: command: scard # --- command: incrby # --- command: ${! meta("command") } ``` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `simple` **Options**: `simple`, `cluster`, `failover` ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yaml # Examples: master: mymaster ``` ### [](#retries)`retries` The maximum number of retries before abandoning a request. **Type**: `int` **Default**: `3` ### [](#retry_period)`retry_period` The time to wait before consecutive retry attempts. **Type**: `string` **Default**: `500ms` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yaml # Examples: url: redis://:6379 # --- url: redis://localhost:6379 # --- url: redis://foousername:foopassword@redisplace:6379 # --- url: redis://:foopassword@redisplace:6379 # --- url: redis://localhost:6379/1 # --- url: redis://localhost:6379/1,redis://localhost:6380/1 ``` --- # Page 232: resource **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/resource.md --- # resource > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: resource latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/resource page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/resource.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/resource.adoc description: Resource is a processor type that runs a processor resource identified by its label. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Resource is a processor type that runs a processor resource identified by its label. ```yml # Config fields, showing default values resource: "" ``` This processor allows you to reference the same configured processor resource in multiple places, and can also tidy up large nested configs. For example, the config: ```yaml pipeline: processors: - mapping: | root.message = this root.meta.link_count = this.links.length() root.user.age = this.user.age.number() ``` Is equivalent to: ```yaml pipeline: processors: - resource: foo_proc processor_resources: - label: foo_proc mapping: | root.message = this root.meta.link_count = this.links.length() root.user.age = this.user.age.number() ``` --- # Page 233: retry **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/retry.md --- # retry > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: retry latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/retry page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/retry.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/retry.adoc description: Attempts to execute a series of child processors until success. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Attempts to execute a series of child processors until success. ```yml # Config fields, showing default values label: "" retry: backoff: initial_interval: 500ms max_interval: 10s max_elapsed_time: 1m processors: [] # No default (required) parallel: false max_retries: 0 ``` Executes child processors and if a resulting message is errored then, after a specified backoff period, the same original message will be attempted again through those same processors. If the child processors result in more than one message then the retry mechanism will kick in if _any_ of the resulting messages are errored. It is important to note that any mutations performed on the message during these child processors will be discarded for the next retry, and therefore it is safe to assume that each execution of the child processors will always be performed on the data as it was when it first reached the retry processor. By default the retry backoff has a specified [`max_elapsed_time`](#backoffmax_elapsed_time), if this time period is reached during retries and an error still occurs these errored messages will proceed through to the next processor after the retry (or your outputs). Normal [error handling patterns](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/) can be used on these messages. In order to avoid permanent loops any error associated with messages as they first enter a retry processor will be cleared. ## [](#metadata)Metadata This processor adds the following metadata fields to each message: - `retry_count` - The number of retry attempts. - `backoff_duration` - The total time (in nanoseconds) elapsed while performing retries. > ⚠️ **CAUTION: Batching** > > Batching > > If you wish to wrap a batch-aware series of processors then take a look at the [batching section](#batching). ## [](#examples)Examples ### [](#stop-ignoring-me-taz)Stop ignoring me Taz Here we have a config where I generate animal noises and send them to Taz via HTTP. Taz has a tendency to stop his servers whenever I dispatch my animals upon him, and therefore these HTTP requests sometimes fail. However, I have the retry processor and with this super power I can specify a back off policy and it will ensure that for each animal noise the HTTP processor is attempted until either it succeeds or my Redpanda Connect instance is stopped. I even go as far as to zero-out the maximum elapsed time field, which means that for each animal noise I will wait indefinitely, because I really really want Taz to receive every single animal noise that he is entitled to. ```yaml input: generate: interval: 1s mapping: 'root.noise = [ "woof", "meow", "moo", "quack" ].index(random_int(min: 0, max: 3))' pipeline: processors: - retry: backoff: initial_interval: 100ms max_interval: 5s max_elapsed_time: 0s processors: - http: url: 'http://example.com/try/not/to/dox/taz' verb: POST output: # Drop everything because it's junk data, I don't want it lol drop: {} ``` ## [](#fields)Fields ### [](#backoff)`backoff` Determine time intervals and cut offs for retry attempts. **Type**: `object` ### [](#backoff-initial_interval)`backoff.initial_interval` The initial period to wait between retry attempts. The retry interval increases for each failed attempt, up to the `backoff.max_interval` value. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `500ms` ```yaml # Examples: initial_interval: 50ms # --- initial_interval: 1s ``` ### [](#backoff-max_elapsed_time)`backoff.max_elapsed_time` The maximum overall period of time to spend on retry attempts before the request is aborted. Setting this value to a zeroed duration (such as `0s`) will result in unbounded retries. **Type**: `string` **Default**: `1m` ```yaml # Examples: max_elapsed_time: 1m # --- max_elapsed_time: 1h ``` ### [](#backoff-max_interval)`backoff.max_interval` The maximum period to wait between retry attempts **Type**: `string` **Default**: `10s` ```yaml # Examples: max_interval: 5s # --- max_interval: 1m ``` ### [](#max_retries)`max_retries` The maximum number of retry attempts before the request is aborted. Setting this value to `0` will result in unbounded number of retries. **Type**: `int` **Default**: `0` ### [](#parallel)`parallel` When processing batches of messages these batches are ignored and the processors apply to each message sequentially. However, when this field is set to `true` each message will be processed in parallel. Caution should be made to ensure that batch sizes do not surpass a point where this would cause resource (CPU, memory, API limits) contention. **Type**: `bool` **Default**: `false` ### [](#processors)`processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to execute on each message. **Type**: `array` ## [](#batching)Batching When messages are batched the child processors of a retry are executed for each individual message in isolation, performed serially by default but in parallel when the field [`parallel`](#parallel) is set to `true`. This is an intentional limitation of the retry processor and is done in order to ensure that errors are correctly associated with a given input message. Otherwise, the archiving, expansion, grouping, filtering and so on of the child processors could obfuscate this relationship. If the target behavior of your retried processors is "batch aware", in that you wish to perform some processing across the entire batch of messages and repeat it in the event of errors, you can use an [`archive` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive/) to collapse the batch into an individual message. Then, within these child processors either perform your batch aware processing on the archive, or use an [`unarchive` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/unarchive/) in order to expand the single message back out into a batch. For example, if the retry processor were being used to wrap an HTTP request where the payload data is a batch archived into a JSON array it should look something like this: ```yaml pipeline: processors: - archive: format: json_array - retry: processors: - http: url: example.com/nope verb: POST - unarchive: format: json_array ``` --- # Page 234: schema_registry_decode **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/schema_registry_decode.md --- # schema_registry_decode > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: schema_registry_decode latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/schema_registry_decode page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/schema_registry_decode.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/schema_registry_decode.adoc description: Automatically decodes and validates messages with schemas from a Confluent Schema Registry service. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Automatically decodes and validates messages with schemas from a Confluent Schema Registry service. This processor uses the [Franz Kafka Schema Registry client](https://github.com/twmb/franz-go/tree/master/pkg/sr). #### Common ```yml processors: label: "" schema_registry_decode: avro: raw_unions: "" # No default (optional) preserve_logical_types: false translate_kafka_connect_types: false mapping: "" # No default (optional) store_schema_metadata: "" # No default (optional) protobuf: use_proto_names: false use_enum_numbers: false emit_unpopulated: false emit_default_values: false serialize_to_json: true json: coerce_data: false cache_duration: 10m url: "" # No default (required) default_schema_id: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" schema_registry_decode: avro: raw_unions: "" # No default (optional) preserve_logical_types: false translate_kafka_connect_types: false mapping: "" # No default (optional) store_schema_metadata: "" # No default (optional) protobuf: use_proto_names: false use_enum_numbers: false emit_unpopulated: false emit_default_values: false serialize_to_json: true json: coerce_data: false cache_duration: 10m url: "" # No default (required) default_schema_id: "" # No default (optional) oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} tls: skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] ``` Decodes messages automatically from a schema stored within a [Confluent Schema Registry service](https://docs.confluent.io/platform/current/schema-registry/index.html) by extracting a schema ID from the message and obtaining the associated schema from the registry. If a message fails to match against the schema then it will remain unchanged and the error can be caught using [error-handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). Avro, Protobuf and JSON schemas are supported, all are capable of expanding from schema references as of v4.22.0. ## [](#avro-json-format)Avro JSON format By default, this processor expects documents formatted as [Avro JSON](https://avro.apache.org/docs/current/specification/) when decoding with Avro schemas. In this format, the value of a union is encoded in JSON as follows: - If the union’s type is `null`, it is encoded as a JSON `null`. - Otherwise, the union is encoded as a JSON object with one name/value pair. The name is the type’s name, and the value is the recursively-encoded value. The user-specified name is used for Avro’s named types (record, fixed, or enum). For other types, the type name is used. For example, the union schema `["null","string","Transaction"]`, where `Transaction` is a record name, would encode: - `null` as a JSON `null` - The string `"a"` as `{"string": "a"}` - A `Transaction` instance as `{"Transaction": {…​}}`, where `{…​}` indicates the JSON encoding of a `Transaction` instance Alternatively, you can create documents in [standard/raw JSON format](https://pkg.go.dev/github.com/linkedin/goavro/v2#NewCodecForStandardJSONFull) by setting the field [`avro.raw_unions`](#avro-raw_unions) to `true`. ## [](#protobuf-format)Protobuf format This processor decodes Protobuf messages to JSON documents. For more information about the JSON mapping of Protobuf messages, see the [Protocol Buffers documentation](https://developers.google.com/protocol-buffers/docs/proto3#json). ## [](#metadata)Metadata This processor also adds the following metadata to each outgoing message: schema\_id: the ID of the schema in the schema registry that was associated with the message. ## [](#fields)Fields ### [](#avro)`avro` Configuration for how to decode schemas that are of type AVRO. **Type**: `object` ### [](#avro-mapping)`avro.mapping` Define a custom mapping to apply to the JSON representation of Avro schemas. You can use mappings to convert custom types emitted by other tools, such as Debezium, into standard Avro types. **Type**: `string` ```yaml # Examples: mapping: |- map isDebeziumTimestampType { root = this.type == "long" && this."connect.name" == "io.debezium.time.Timestamp" && !this.exists("logicalType") } map debeziumTimestampToAvroTimestamp { let mapped_fields = this.fields.or([]).map_each(item -> item.apply("debeziumTimestampToAvroTimestamp")) root = match { this.type == "record" => this.assign({"fields": $mapped_fields}) this.type.type() == "array" => this.assign({"type": this.type.map_each(item -> item.apply("debeziumTimestampToAvroTimestamp"))}) # Add a logical type so that it's decoded as a timestamp instead of a long. this.type.type() == "object" && this.type.apply("isDebeziumTimestampType") => this.merge({"type":{"logicalType": "timestamp-millis"}}) _ => this } } root = this.apply("debeziumTimestampToAvroTimestamp") ``` ### [](#avro-preserve_logical_types)`avro.preserve_logical_types` Choose whether to: - Transform logical types into their primitive type (default). For example, decimals become raw bytes and timestamps become plain integers. - Preserve logical types. Set to `true` to preserve logical types. **Type**: `bool` **Default**: `false` ### [](#avro-raw_unions)`avro.raw_unions` Whether Avro messages should be decoded into normal JSON (JSON that meets the expectations of regular internet JSON) rather than [Avro JSON](https://avro.apache.org/docs/current/specification/). If set to `false`, Avro messages are decoded as [Avro JSON](https://pkg.go.dev/github.com/linkedin/goavro/v2#NewCodec). For example, the union schema `["null","string","Transaction"]`, where `Transaction` is a record name, would be decoded as: - A `null` as a JSON `null` - The string `"a"` as `{"string": "a"}` - A `Transaction` instance as `{"Transaction": {…​}}`, where `{…​}` indicates the JSON encoding of a `Transaction` instance. If set to `true`, Avro messages are decoded as [standard JSON](https://pkg.go.dev/github.com/linkedin/goavro/v2#NewCodecForStandardJSONFull). For example, the same union schema `["null","string","Transaction"]` is decoded as: - A `null` as JSON `null` - The string `"a"` as `"a"` - A `Transaction` instance as `{…​}`, where `{…​}` indicates the JSON encoding of a `Transaction` instance. For more details on the difference between standard JSON and Avro JSON, see the [comment in Goavro](https://github.com/linkedin/goavro/blob/5ec5a5ee7ec82e16e6e2b438d610e1cab2588393/union.go#L224-L249) and the [underlying library used for Avro serialization](https://github.com/linkedin/goavro). **Type**: `bool` ### [](#avro-store_schema_metadata)`avro.store_schema_metadata` Optionally store the schema used to decode messages as a metadata field under the given name. This field can later be referenced in other components such as a `parquet_encode` processor in order to automatically infer their schema. **Type**: `string` ### [](#avro-translate_kafka_connect_types)`avro.translate_kafka_connect_types` Only valid if preserve\_logical\_types is true. This decodes various Kafka Connect types into their bloblang equivalents when not representable by standard logical types according to the Avro standard. Types that are currently translated: | Type Name | Bloblang Type | Description | | --- | --- | --- | | io.debezium.time.Date | timestamp | Date without time (days since epoch) | | io.debezium.time.Timestamp | timestamp | Timestamp without timezone (milliseconds since epoch) | | io.debezium.time.MicroTimestamp | timestamp | Timestamp with microsecond precision | | io.debezium.time.NanoTimestamp | timestamp | Timestamp with nanosecond precision | | io.debezium.time.ZonedTimestamp | timestamp | Timestamp with timezone (ISO-8601 format) | | io.debezium.time.Year | timestamp at January 1st at 00:00:00 | Year value | | io.debezium.time.Time | timestamp at the unix epoch | Time without date (milliseconds past midnight) | | io.debezium.time.MicroTime | timestamp at the unix epoch | Time with microsecond precision | | io.debezium.time.NanoTime | timestamp at the unix epoch | Time with nanosecond precision | **Type**: `bool` **Default**: `false` ### [](#basic_auth)`basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#cache_duration)`cache_duration` The duration after which a cached schema is considered stale and is removed from the cache. **Type**: `string` **Default**: `10m` ```yaml # Examples: cache_duration: 1h # --- cache_duration: 5m ``` ### [](#default_schema_id)`default_schema_id` This schema ID is used when a message’s schema header cannot be read (`ErrBadHeader`). If this value is not set, schema header errors are returned. This configuration does not work with protobuf schemas. > 💡 **TIP** > > You can also use the [`with_schema_registry_header`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#with_schema_registry_header) bloblang function to add a schema ID to messages. **Type**: `int` ### [](#json)`json` Configuration for how to decode schemas that are of type JSON. **Type**: `object` ### [](#json-coerce_data)`json.coerce_data` Whether decoded values should be coerced to match the types declared in the JSON Schema. By default JSON Schema decoding only validates the message and leaves it untouched, which means numbers are later interpreted as floating point (`double`) and date-time values as strings. When set to `true` the decoder rebuilds the message so that values match the schema: `integer` fields become 64-bit integers, `number` fields stay floating point, `string` fields with `format: date-time` become timestamps, and any `default` values declared in the schema are applied to absent fields. This is useful for downstream components that infer their schema from the decoded values, such as the `iceberg` outputs, which will then create `bigint` columns for integer fields rather than `double`. Note that, unlike the default behavior, this is no longer a read-only operation: the message contents are transformed. Because coercion is stricter than validation, a message that passes validation may still fail coercion (for example an integer that overflows a 64-bit value, or a `date-time` string that is not valid RFC 3339), in which case the error can be caught using [error handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). **Type**: `bool` **Default**: `false` ### [](#jwt)`jwt` (beta) Configure JSON Web Token (JWT) authentication. This feature is in beta and may change in future releases. JWT tokens provide secure, stateless authentication between services. **Type**: `object` ### [](#jwt-claims)`jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` Additional key-value pairs to include in the JWT header (optional). These headers provide extra metadata for JWT processing. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` Path to a file containing the PEM-encoded private key using PKCS#1 or PKCS#8 format. The private key must be compatible with the algorithm specified in the `signing_method` field. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` The cryptographic algorithm used to sign the JWT token. Supported algorithms include RS256, RS384, RS512, and EdDSA. This algorithm must be compatible with the private key specified in the `private_key_file` field. **Type**: `string` **Default**: `""` ### [](#oauth)`oauth` Configure OAuth version 1.0 authentication for secure API access. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#protobuf)`protobuf` Configuration for how to decode schemas that are of type PROTOBUF. **Type**: `object` ### [](#protobuf-emit_default_values)`protobuf.emit_default_values` Whether to emit default-valued primitive fields, empty lists, and empty maps. emit\_unpopulated takes precedence over emit\_default\_values **Type**: `bool` **Default**: `false` ### [](#protobuf-emit_unpopulated)`protobuf.emit_unpopulated` Whether to emit unpopulated fields. It does not emit unpopulated oneof fields or unpopulated extension fields. **Type**: `bool` **Default**: `false` ### [](#protobuf-serialize_to_json)`protobuf.serialize_to_json` If messages should be serialized to JSON bytes. If false then the message is kept in decoded form, which means that 64 bit integers are not converted to strings and types for bytes and google.protobuf.Timestamp are preserved (as they are not serialized to JSON strings). **Type**: `bool` **Default**: `true` ### [](#protobuf-use_enum_numbers)`protobuf.use_enum_numbers` Emits enum values as numbers. **Type**: `bool` **Default**: `false` ### [](#protobuf-use_proto_names)`protobuf.use_proto_names` Use proto field name instead of lowerCamelCase name. **Type**: `bool` **Default**: `false` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The base URL of the schema registry service. **Type**: `string` --- # Page 235: schema_registry_encode **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/schema_registry_encode.md --- # schema_registry_encode > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: schema_registry_encode latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/schema_registry_encode page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/schema_registry_encode.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/schema_registry_encode.adoc description: Automatically encodes and validates messages with schemas from a Confluent Schema Registry service. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Automatically encodes and validates messages with schemas from a Confluent Schema Registry service. This processor uses the [Franz Kafka Schema Registry client](https://github.com/twmb/franz-go/tree/master/pkg/sr). #### Common ```yml processors: label: "" schema_registry_encode: url: "" # No default (required) subject: "" # No default (required) refresh_period: 10m schema_metadata: "" format: "" # No default (optional) avro: raw_json: "" # No default (optional) record_name: "" namespace: "" ``` #### Advanced ```yml processors: label: "" schema_registry_encode: url: "" # No default (required) subject: "" # No default (required) refresh_period: 10m schema_metadata: "" format: "" # No default (optional) normalize: true avro: raw_json: "" # No default (optional) record_name: "" namespace: "" oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} tls: skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] ``` Encodes messages automatically from schemas obtains from a [Confluent Schema Registry service](https://docs.confluent.io/platform/current/schema-registry/index.html) by polling the service for the latest schema version for target subjects. If a message fails to encode under the schema then it will remain unchanged and the error can be caught using [error-handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). Avro, Protobuf and JSON schemas are supported, all are capable of expanding from schema references as of v4.22.0. ## [](#avro-json-format)Avro JSON format By default, this processor expects documents formatted as [Avro JSON](https://avro.apache.org/docs/current/specification/) when encoding with Avro schemas. In this format, the value of a union is encoded in JSON as follows: - If the union’s type is `null`, it is encoded as a JSON `null`. - Otherwise, the union is encoded as a JSON object with one name/value pair. The name is the type’s name, and the value is the recursively-encoded value. The user-specified name is used for Avro’s named types (record, fixed, or enum). For other types, the type name is used. For example, the union schema `["null","string","Transaction"]`, where `Transaction` is a record name, would encode: - A `null` as a JSON `null` - The string `"a"` as `{"string": "a"}` - A `Transaction` instance as `{"Transaction": {…​}}`, where `{…​}` indicates the JSON encoding of a `Transaction` instance Alternatively, you can consume documents in [standard/raw JSON format](https://pkg.go.dev/github.com/linkedin/goavro/v2#NewCodecForStandardJSONFull) by setting the field [`avro_raw_json`](#avro_raw_json) to `true`. ### [](#known-issues)Known issues Important! There is an outstanding issue in the [avro serializing library](https://github.com/linkedin/goavro) that Redpanda Connect uses which means it [doesn’t encode logical types correctly](https://github.com/linkedin/goavro/issues/252). It’s still possible to encode logical types that are in-line with the spec if `avro_raw_json` is set to true, though now of course non-logical types will not be in-line with the spec. ## [](#protobuf-format)Protobuf format This processor encodes Protobuf messages either from any format parsed within Redpanda Connect (encoded as JSON by default), or from raw JSON documents. For more information about the JSON mapping of Protobuf messages, see the [Protocol Buffers documentation](https://developers.google.com/protocol-buffers/docs/proto3#json). ### [](#multiple-message-support)Multiple message support When a target subject presents a Protobuf schema that contains multiple messages it becomes ambiguous which message definition a given input data should be encoded against. In such scenarios Redpanda Connect will attempt to encode the data against each of them and select the first to successfully match against the data, this process currently **ignores all nested message definitions**. In order to speed up this exhaustive search the last known successful message will be attempted first for each subsequent input. We will be considering alternative approaches in future so please [get in touch](https://redpanda.com/slack) with thoughts and feedback. ## [](#fields)Fields ### [](#avro)`avro` Configuration for Avro encoding. **Type**: `object` ### [](#avro-namespace)`avro.namespace` The Avro namespace for the root record type when encoding from a common schema (schema\_metadata mode). **Type**: `string` **Default**: `""` ### [](#avro-raw_json)`avro.raw_json` Whether messages encoded in Avro format should be parsed as normal JSON rather than Avro JSON. Overrides the deprecated top-level `avro_raw_json` when set. **Type**: `bool` ### [](#avro-record_name)`avro.record_name` The name to use for the root Avro record type when encoding from a common schema (schema\_metadata mode). If empty, derived from the subject. **Type**: `string` **Default**: `""` ### [](#basic_auth)`basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#basic_auth-enabled)`basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#basic_auth-password)`basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#basic_auth-username)`basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#format)`format` The encoding format to use when converting a common schema from metadata. Required when `schema_metadata` is set. **Type**: `string` **Options**: `avro`, `json_schema` ### [](#jwt)`jwt` (beta) Configure JSON Web Token (JWT) authentication. This feature is in beta and may change in future releases. JWT tokens provide secure, stateless authentication between services. **Type**: `object` ### [](#jwt-claims)`jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#jwt-enabled)`jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#jwt-headers)`jwt.headers` Additional key-value pairs to include in the JWT header (optional). These headers provide extra metadata for JWT processing. **Type**: `object` **Default**: `{}` ### [](#jwt-private_key_file)`jwt.private_key_file` Path to a file containing the PEM-encoded private key using PKCS#1 or PKCS#8 format. The private key must be compatible with the algorithm specified in the `signing_method` field. **Type**: `string` **Default**: `""` ### [](#jwt-signing_method)`jwt.signing_method` The cryptographic algorithm used to sign the JWT token. Supported algorithms include RS256, RS384, RS512, and EdDSA. This algorithm must be compatible with the private key specified in the `private_key_file` field. **Type**: `string` **Default**: `""` ### [](#normalize)`normalize` Whether to normalize the schema before registering with the schema registry (schema\_metadata mode only). **Type**: `bool` **Default**: `true` ### [](#oauth)`oauth` Configure OAuth version 1.0 authentication for secure API access. **Type**: `object` ### [](#oauth-access_token)`oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#oauth-access_token_secret)`oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_key)`oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#oauth-consumer_secret)`oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#oauth-enabled)`oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#refresh_period)`refresh_period` The period after which a schema is refreshed for each subject, this is done by polling the schema registry service. **Type**: `string` **Default**: `10m` ```yaml # Examples: refresh_period: 60s # --- refresh_period: 1h ``` ### [](#schema_metadata)`schema_metadata` When set, the processor reads a schema in benthos common schema format from this metadata key on each message, converts it to the format specified by `format`, registers it with the schema registry under the configured subject, and encodes the message. When empty (the default), the processor pulls the latest schema from the registry instead. **Type**: `string` **Default**: `""` ### [](#subject)`subject` The schema subject to derive schemas from. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ```yaml # Examples: subject: foo # --- subject: ${! meta("kafka_topic") } ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#url)`url` The base URL of the schema registry service. **Type**: `string` --- # Page 236: select_parts **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/select_parts.md --- # select_parts > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: select_parts latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/select_parts page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/select_parts.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/select_parts.adoc description: Cherry pick a set of messages from a batch by their index. Indexes larger than the number of messages are simply ignored. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Cherry pick a set of messages from a batch by their index. Indexes larger than the number of messages are simply ignored. ```yml # Config fields, showing default values label: "" select_parts: parts: [] ``` The selected parts are added to the new message batch in the same order as the selection array. E.g. with 'parts' set to \[ 2, 0, 1 \] and the message parts \[ '0', '1', '2', '3' \], the output will be \[ '2', '0', '1' \]. If none of the selected parts exist in the input batch (resulting in an empty output message) the batch is dropped entirely. Message indexes can be negative, and if so the part will be selected from the end counting backwards starting from -1. E.g. if index = -1 then the selected part will be the last part of the message, if index = -2 then the part before the last element with be selected, and so on. This processor is only applicable to [batched messages](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#parts)`parts[]` An array of message indexes of a batch. Indexes can be negative, and if so the part will be selected from the end counting backwards starting from -1. **Type**: `array` **Default**: `[]` --- # Page 237: slack_thread **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/slack_thread.md --- # slack_thread > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: slack_thread latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/slack_thread page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/slack_thread.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/slack_thread.adoc page-git-created-date: "2025-05-02" page-git-modified-date: "2026-05-26" --- Reads a Slack thread using the Slack API method [conversations.replies](https://api.slack.com/methods/conversations.replies). ```yml # Common configuration fields, showing default values label: "" slack_thread: bot_token: "" # No default (required) channel_id: "" # No default (required) thread_ts: "" # No default (required) ``` ## [](#fields)Fields ### [](#bot_token)`bot_token` Your Slack bot user’s OAuth token, which must have the correct permissions to read messages from the Slack channel specified in `channel_id`. **Type**: `string` ### [](#channel_id)`channel_id` The encoded ID of the Slack channel from which to read threads. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` ### [](#thread_ts)`thread_ts` The timestamp of the parent message of the thread you want to read. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 238: sleep **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sleep.md --- # sleep > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sleep latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/sleep page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/sleep.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/sleep.adoc description: Sleep for a period of time specified as a duration string for each message. This processor will interpolate functions within the duration field, you can find a list of functions here. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Sleep for a period of time specified as a duration string for each message. This processor will interpolate functions within the `duration` field, you can find a list of functions [here](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). ```yml # Config fields, showing default values label: "" sleep: duration: "" # No default (required) ``` ## [](#fields)Fields ### [](#duration)`duration` The duration of time to sleep for each execution. This field supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). **Type**: `string` --- # Page 239: split **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/split.md --- # split > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: split latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/split page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/split.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/split.adoc description: Breaks message batches (synonymous with multiple part messages) into smaller batches. The size of the resulting batches are determined either by a discrete size or, if the field byte_size is non-zero, then by total size in bytes (which ever limit is reached first). page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Breaks message batches (synonymous with multiple part messages) into smaller batches. The size of the resulting batches are determined either by a discrete size or, if the field `byte_size` is non-zero, then by total size in bytes (which ever limit is reached first). ```yml # Config fields, showing default values label: "" split: size: 1 byte_size: 0 ``` This processor is for breaking batches down into smaller ones. In order to break a single message out into multiple messages use the [`unarchive` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/unarchive/). If there is a remainder of messages after splitting a batch the remainder is also sent as a single batch. For example, if your target size was 10, and the processor received a batch of 95 message parts, the result would be 9 batches of 10 messages followed by a batch of 5 messages. ## [](#fields)Fields ### [](#byte_size)`byte_size` An optional target of total message bytes. **Type**: `int` **Default**: `0` ### [](#size)`size` The target number of messages. **Type**: `int` **Default**: `1` --- # Page 240: sql_insert **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sql_insert.md --- # sql_insert > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql_insert latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/sql_insert page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/sql_insert.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/sql_insert.adoc description: Inserts rows into an SQL database for each message, and leaves the message unchanged. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Inserts rows into an SQL database for each message, and leaves the message unchanged. #### Common ```yml processors: label: "" sql_insert: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) args_mapping: "" # No default (required) ``` #### Advanced ```yml processors: label: "" sql_insert: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) args_mapping: "" # No default (required) prefix: "" # No default (optional) suffix: "" # No default (optional) options: [] # No default (optional) init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) ``` If the insert fails to execute then the message will still remain unchanged and the error can be caught using [error handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ## [](#examples)Examples ### [](#table-insert-mysql)Table Insert (MySQL) Here we insert rows into a database by populating the columns id, name and topic with values extracted from messages and metadata: ```yaml pipeline: processors: - sql_insert: driver: mysql dsn: foouser:foopassword@tcp(localhost:3306)/foodb table: footable columns: [ id, name, topic ] args_mapping: | root = [ this.user.id, this.user.name, meta("kafka_topic"), ] ``` ## [](#dynamic-sql-operations)Dynamic SQL operations The `table` and `columns` fields are static strings that do not support Bloblang interpolation. For dynamic table names, dynamic column lists, DELETE operations, or any other SQL that `sql_insert` cannot express, use the [`sql_raw` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sql_raw/) instead. To use Bloblang interpolation inside ``sql_raw’s `query`` field, you must enable `unsafe_dynamic_query: true`. > ⚠️ **CAUTION** > > Interpolating unsanitized values into a query can introduce SQL injection risks. Always validate or sanitize the interpolated value beforehand. ## [](#fields)Fields ### [](#args_mapping)`args_mapping` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of columns specified. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#columns)`columns[]` A list of columns to insert. **Type**: `array` ```yaml # Examples: columns: - foo - bar - baz ``` ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#options)`options[]` A list of keyword options to add before the INTO clause of the query. **Type**: `array` ```yaml # Examples: options: - DELAYED - IGNORE ``` ### [](#prefix)`prefix` An optional prefix to prepend to the insert query (before INSERT). **Type**: `string` ### [](#suffix)`suffix` An optional suffix to append to the insert query. **Type**: `string` ```yaml # Examples: suffix: ON CONFLICT (name) DO NOTHING ``` ### [](#table)`table` The table to insert to. **Type**: `string` ```yaml # Examples: table: foo ``` --- # Page 241: sql_raw **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sql_raw.md --- # sql_raw > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql_raw latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/sql_raw page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/sql_raw.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/sql_raw.adoc description: Runs an arbitrary SQL query against a database and (optionally) returns the result as an array of objects, one for each row returned. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Runs an arbitrary SQL query against a database and (optionally) returns the result as an array of objects, one for each row returned. #### Common ```yml processors: label: "" sql_raw: driver: "" # No default (required) dsn: "" # No default (required) query: "" # No default (optional) args_mapping: "" # No default (optional) exec_only: "" # No default (optional) queries: [] # No default (optional) ``` #### Advanced ```yml processors: label: "" sql_raw: driver: "" # No default (required) dsn: "" # No default (required) query: "" # No default (optional) unsafe_dynamic_query: false args_mapping: "" # No default (optional) exec_only: "" # No default (optional) queries: [] # No default (optional) init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) ``` If the query fails to execute then the message will remain unchanged and the error can be caught using [error handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). For some scenarios where you might use this processor, see [Examples](#examples). ## [](#fields)Fields ### [](#args_mapping)`args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that includes the same number of values in an array as the placeholder arguments in the [`query`](#query) field. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#exec_only)`exec_only` Whether to discard the [`query`](#query) result. Set to `true` to leave the message contents unchanged, which is useful when you are executing inserts, updates, and so on. By default, the message contents are kept for the last query executed, and previous queries don’t change the results. **Type**: `bool` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#queries)`queries[]` A list of database statements to run in addition to your main [`query`](#query). If you specify multiple queries, they are executed within a single transaction. For more information, see [Examples](#examples). **Type**: `array` ### [](#queries-args_mapping)`queries[].args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of placeholder arguments in the field `query`. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#queries-exec_only)`queries[].exec_only` Whether the query result should be discarded. When set to `true` the message contents will remain unchanged, which is useful in cases where you are executing inserts, updates, etc. By default this is true for the last query, and previous queries don’t change the results. If set to true for any query but the last one, the subsequent `args_mappings` input is overwritten. **Type**: `bool` ### [](#queries-query)`queries[].query` The query to execute. The style of placeholder to use depends on the driver, some drivers require question marks (`?`) whereas others expect incrementing dollar signs (`$1`, `$2`, and so on) or colons (`:1`, `:2` and so on). The style to use is outlined in this table: | Driver | Placeholder Style | |---|---| | `clickhouse` | Dollar sign | | `mysql` | Question mark | | `postgres` | Dollar sign | | `pgx` | Dollar sign | | `mssql` | Question mark | | `sqlite` | Question mark | | `oracle` | Colon | | `snowflake` | Question mark | | `trino` | Question mark | | `gocosmos` | Colon | **Type**: `string` ### [](#query)`query` The query to execute. You must include the correct placeholders for the specified database driver. Some drivers use question marks (`?`), whereas others expect incrementing dollar signs (`$1`, `$2`, and so on) or colons (`:1`, `:2`, and so on). | Driver | Placeholder Style | | --- | --- | | clickhouse | Dollar sign ($) | | gocosmos | Colon (:) | | mysql | Question mark (?) | | mssql | Question mark (?) | | oracle | Colon (:) | | postgres | Dollar sign ($) | | snowflake | Question mark (?) | | spanner | Question mark (?) | | sqlite | Question mark (?) | | trino | Question mark (?) | **Type**: `string` ```yaml # Examples: query: INSERT INTO footable (foo, bar, baz) VALUES (?, ?, ?); # --- query: SELECT * FROM footable WHERE user_id = $1; ``` ### [](#unsafe_dynamic_query)`unsafe_dynamic_query` Whether to enable [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries) in the query. Great care should be made to ensure your queries are defended against injection attacks. **Type**: `bool` **Default**: `false` ## [](#examples)Examples ### [](#table-insert-mysql)Table Insert (MySQL) The following example inserts rows into the table footable with the columns foo, bar and baz populated with values extracted from messages. ```yaml pipeline: processors: - sql_raw: driver: mysql dsn: foouser:foopassword@tcp(localhost:3306)/foodb query: "INSERT INTO footable (foo, bar, baz) VALUES (?, ?, ?);" args_mapping: '[ document.foo, document.bar, meta("kafka_topic") ]' exec_only: true ``` ### [](#table-query-postgresql)Table Query (PostgreSQL) Here we query a database for columns of footable that share a `user_id` with the message field `user.id`. A [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) is used in order to insert the resulting array into the original message at the path `foo_rows`. ```yaml pipeline: processors: - branch: processors: - sql_raw: driver: postgres dsn: postgres://foouser:foopass@localhost:5432/testdb?sslmode=disable query: "SELECT * FROM footable WHERE user_id = $1;" args_mapping: '[ this.user.id ]' result_map: 'root.foo_rows = this' ``` ### [](#dynamically-creating-tables-postgresql)Dynamically Creating Tables (PostgreSQL) Here we query a database for columns of footable that share a `user_id` with the message field `user.id`. A [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) is used in order to insert the resulting array into the original message at the path `foo_rows`. ```yaml pipeline: processors: - mapping: | root = this # Prevent SQL injection when using unsafe_dynamic_query meta table_name = "\"" + metadata("table_name").replace_all("\"", "\"\"") + "\"" - sql_raw: driver: postgres dsn: postgres://localhost/postgres unsafe_dynamic_query: true queries: - query: | CREATE TABLE IF NOT EXISTS ${!metadata("table_name")} (id varchar primary key, document jsonb); - query: | INSERT INTO ${!metadata("table_name")} (id, document) VALUES ($1, $2) ON CONFLICT (id) DO UPDATE SET document = EXCLUDED.document; args_mapping: | root = [ this.id, this.document.string() ] ``` --- # Page 242: sql_select **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sql_select.md --- # sql_select > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sql_select latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/sql_select page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/sql_select.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/sql_select.adoc description: Runs an SQL select query against a database and returns the result as an array of objects, one for each row returned, containing a key for each column queried and its value. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Runs an SQL select query against a database and returns the result as an array of objects, one for each row returned, containing a key for each column queried and its value. #### Common ```yml processors: label: "" sql_select: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) where: "" # No default (optional) args_mapping: "" # No default (optional) ``` #### Advanced ```yml processors: label: "" sql_select: driver: "" # No default (required) dsn: "" # No default (required) table: "" # No default (required) columns: [] # No default (required) where: "" # No default (optional) args_mapping: "" # No default (optional) prefix: "" # No default (optional) suffix: "" # No default (optional) init_files: [] # No default (optional) init_statement: "" # No default (optional) conn_max_idle_time: "" # No default (optional) conn_max_life_time: "" # No default (optional) conn_max_idle: 2 conn_max_open: "" # No default (optional) ``` If the query fails to execute then the message will remain unchanged and the error can be caught using [error handling methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ## [](#examples)Examples ### [](#table-query-postgresql)Table Query (PostgreSQL) Here we query a database for columns of footable that share a `user_id` with the message `user.id`. A [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) is used in order to insert the resulting array into the original message at the path `foo_rows`: ```yaml pipeline: processors: - branch: processors: - sql_select: driver: postgres dsn: postgres://foouser:foopass@localhost:5432/testdb?sslmode=disable table: footable columns: [ '*' ] where: user_id = ? args_mapping: '[ this.user.id ]' result_map: 'root.foo_rows = this' ``` ## [](#fields)Fields ### [](#args_mapping)`args_mapping` An optional [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) which should evaluate to an array of values matching in size to the number of placeholder arguments in the field `where`. **Type**: `string` ```yaml # Examples: args_mapping: root = [ this.cat.meow, this.doc.woofs[0] ] # --- args_mapping: root = [ meta("user.id") ] ``` ### [](#columns)`columns[]` A list of columns to query. **Type**: `array` ```yaml # Examples: columns: - "*" # --- columns: - foo - bar - baz ``` ### [](#conn_max_idle)`conn_max_idle` An optional maximum number of connections in the idle connection pool. If conn\_max\_open is greater than 0 but less than the new conn\_max\_idle, then the new conn\_max\_idle will be reduced to match the conn\_max\_open limit. If `value ⇐ 0`, no idle connections are retained. The default max idle connections is currently 2. This may change in a future release. **Type**: `int` **Default**: `2` ### [](#conn_max_idle_time)`conn_max_idle_time` An optional maximum amount of time a connection may be idle. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections idle time. **Type**: `string` ### [](#conn_max_life_time)`conn_max_life_time` An optional maximum amount of time a connection may be reused. Expired connections may be closed lazily before reuse. If `value ⇐ 0`, connections are not closed due to a connections age. **Type**: `string` ### [](#conn_max_open)`conn_max_open` An optional maximum number of open connections to the database. If conn\_max\_idle is greater than 0 and the new conn\_max\_open is less than conn\_max\_idle, then conn\_max\_idle will be reduced to match the new conn\_max\_open limit. If `value ⇐ 0`, then there is no limit on the number of open connections. The default is 0 (unlimited). **Type**: `int` ### [](#driver)`driver` A database [driver](#drivers) to use. **Type**: `string` **Options**: `mysql`, `postgres`, `pgx`, `clickhouse`, `mssql`, `sqlite`, `oracle`, `snowflake`, `trino`, `gocosmos`, `spanner`, `databricks` ### [](#dsn)`dsn` A Data Source Name to identify the target database. #### [](#drivers)Drivers The following is a list of supported drivers, their placeholder style, and their respective DSN formats: | Driver | Data Source Name Format | | --- | --- | | clickhouse | clickhouse://[username[:password]@][netloc][:port]/dbname[?param1=value1&…​¶mN=valueN] | | mysql | [username[:password]@][protocol[(address)]]/dbname[?param1=value1&…​¶mN=valueN] | | postgres and pgx | postgres://[user[:password]@][netloc][:port][/dbname][?param1=value1&…​] | | mssql | sqlserver://[user[:password]@][netloc][:port][?database=dbname¶m1=value1&…​] | | sqlite | file:/path/to/filename.db[?param&=value1&…​] | | oracle | oracle://[username[:password]@][netloc][:port]/service_name?server=server2&server=server3 | | snowflake | username[:password]@account_identifier/dbname/schemaname[?param1=value&…​¶mN=valueN] | | trino | http[s]://user[:pass]@host[:port][?parameters] | | gocosmos | AccountEndpoint=;AccountKey=[;TimeoutMs=][;Version=][;DefaultDb/Db=][;AutoId=][;InsecureSkipVerify=] | | spanner | projects/[PROJECT]/instances/[INSTANCE]/databases/[DATABASE] | | databricks | token:@:/ | Please note that the `postgres` and `pgx` drivers enforce SSL by default, you can override this with the parameter `sslmode=disable` if required. The `pgx` driver is an alternative to the standard `postgres` (pq) driver and comes with extra functionality such as support for array insertion. The `snowflake` driver supports multiple DSN formats. Please consult [the docs](https://pkg.go.dev/github.com/snowflakedb/gosnowflake#hdr-Connection_String) for more details. For [key pair authentication](https://docs.snowflake.com/en/user-guide/key-pair-auth.html#configuring-key-pair-authentication), the DSN has the following format: `@//?warehouse=&role=&authenticator=snowflake_jwt&privateKey=`, where the value for the `privateKey` parameter can be constructed from an unencrypted RSA private key file `rsa_key.p8` using `openssl enc -d -base64 -in rsa_key.p8 | basenc --base64url -w0` (you can use `gbasenc` instead of `basenc` on OSX if you install `coreutils` via Homebrew). If you have a password-encrypted private key, you can decrypt it using `openssl pkcs8 -in rsa_key_encrypted.p8 -out rsa_key.p8`. Also, make sure fields such as the username are URL-encoded. The [`gocosmos`](https://pkg.go.dev/github.com/microsoft/gocosmos) driver is still experimental, but it has support for [hierarchical partition keys](https://learn.microsoft.com/en-us/azure/cosmos-db/hierarchical-partition-keys) as well as [cross-partition queries](https://learn.microsoft.com/en-us/azure/cosmos-db/nosql/how-to-query-container#cross-partition-query). Please refer to the [SQL notes](https://github.com/microsoft/gocosmos/blob/main/SQL.md) for details. **Type**: `string` ```yaml # Examples: dsn: clickhouse://username:password@host1:9000,host2:9000/database?dial_timeout=200ms&max_execution_time=60 # --- dsn: foouser:foopassword@tcp(localhost:3306)/foodb # --- dsn: postgres://foouser:foopass@localhost:5432/foodb?sslmode=disable # --- dsn: oracle://foouser:foopass@localhost:1521/service_name # --- dsn: token:dapi1234567890ab@dbc-a1b2345c-d6e7.cloud.databricks.com:443/sql/1.0/warehouses/abc123def456 ``` ### [](#init_files)`init_files[]` An optional list of file paths containing SQL statements to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Glob patterns are supported, including super globs (double star). Care should be taken to ensure that the statements are idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If a statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `array` ```yaml # Examples: init_files: - ./init/*.sql # --- init_files: - ./foo.sql - ./bar.sql ``` ### [](#init_statement)`init_statement` An optional SQL statement to execute immediately upon the first connection to the target database. This is a useful way to initialise tables before processing data. Care should be taken to ensure that the statement is idempotent, and therefore would not cause issues when run multiple times after service restarts. If both `init_statement` and `init_files` are specified the `init_statement` is executed _after_ the `init_files`. If the statement fails for any reason a warning log will be emitted but the operation of this component will not be stopped. **Type**: `string` ```yaml # Examples: init_statement: |- CREATE TABLE IF NOT EXISTS some_table ( foo varchar(50) not null, bar integer, baz varchar(50), primary key (foo) ) WITHOUT ROWID; ``` ### [](#prefix)`prefix` An optional prefix to prepend to the query (before SELECT). **Type**: `string` ### [](#suffix)`suffix` An optional suffix to append to the select query. **Type**: `string` ### [](#table)`table` The table to query. **Type**: `string` ```yaml # Examples: table: foo ``` ### [](#where)`where` An optional where clause to add. Placeholder arguments are populated with the `args_mapping` field. Placeholders should always be question marks, and will automatically be converted to dollar syntax when the postgres or clickhouse drivers are used. **Type**: `string` ```yaml # Examples: where: meow = ? and woof = ? # --- where: user_id = ? ``` --- # Page 243: string_split **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/string_split.md --- # string_split > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: string_split latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/string_split page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/string_split.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/string_split.adoc description: Splits a string by a delimiter into an array. Generally, using bloblang's `split` method is preferred. In some high performance use cases this processor can be faster than the equivalent bloblang if there is no additional logic. page-git-created-date: "2026-04-08" page-git-modified-date: "2026-08-11" --- Splits a string by a delimiter into an array. Generally, using bloblang’s `split` method is preferred. In some high performance use cases this processor can be faster than the equivalent bloblang if there is no additional logic. #### Common ```yml processors: label: "" string_split: delimiter: empty_as_null: false ``` #### Advanced ```yml processors: label: "" string_split: delimiter: emit_bytes: false empty_as_null: false ``` ## [](#fields)Fields ### [](#delimiter)`delimiter` The delimiter to split the string by. **Type**: `string` **Default**: \` \` ### [](#emit_bytes)`emit_bytes` When true, the output will be bloblang bytes instead of strings. **Type**: `bool` **Default**: `false` ### [](#empty_as_null)`empty_as_null` When true, empty strings resulting from the split are converted to null. **Type**: `bool` **Default**: `false` --- # Page 244: switch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/switch.md --- # switch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: switch latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/switch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/switch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/switch.adoc description: Conditionally processes messages based on their contents. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Conditionally processes messages based on their contents. ```yml # Config fields, showing default values label: "" switch: [] # No default (required) ``` For each switch case a [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) is checked and, if the result is true (or the check is empty) the child processors are executed on the message. ## [](#fields)Fields ### [](#check)`check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether a message should have the processors of this case executed on it. If left empty the case always passes. If the check mapping throws an error the message will be flagged [as having failed](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/) and will not be tested against any other cases. **Type**: `string` **Default**: `""` ```yaml # Examples: check: this.type == "foo" # --- check: this.contents.urls.contains("https://benthos.dev/") ``` ### [](#continue)`continue` Indicates whether, if this case passes for a message, the next case should also be tested. Unlike `fallthrough`, which skips the next case’s check, `continue` will evaluate the next case’s condition before executing. **Type**: `bool` **Default**: `false` ### [](#fallthrough)`fallthrough` Indicates whether, if this case passes for a message, the next case should also be executed without checking its condition. **Type**: `bool` **Default**: `false` ### [](#processors)`processors[]` A list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to execute on a message. **Type**: `array` **Default**: `[]` ## [](#examples)Examples ### [](#ignore-george)Ignore George We have a system where we’re counting a metric for all messages that pass through our system. However, occasionally we get messages from George that we don’t care about. For George’s messages we want to instead emit a metric that gauges how angry he is about being ignored and then we drop it. ```yaml pipeline: processors: - switch: - check: this.user.name.first != "George" processors: - metric: type: counter name: MessagesWeCareAbout - processors: - metric: type: gauge name: GeorgesAnger value: ${! json("user.anger") } - mapping: root = deleted() ``` ## [](#batching)Batching When a switch processor executes on a [batch of messages](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/) they are checked individually and can be matched independently against cases. During processing the messages matched against a case are processed as a batch, although the ordering of messages during case processing cannot be guaranteed to match the order as received. At the end of switch processing the resulting batch will follow the same ordering as the batch was received. If any child processors have split or otherwise grouped messages this grouping will be lost as the result of a switch is always a single batch. In order to perform conditional grouping and/or splitting use the [`group_by` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/group_by/). --- # Page 245: sync_response **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sync_response.md --- # sync_response > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: sync_response latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/sync_response page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/sync_response.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/sync_response.adoc description: Adds the payload in its current state as a synchronous response to the input source, where it is dealt with according to that specific input type. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Adds the payload in its current state as a synchronous response to the input source, where it is dealt with according to that specific input type. ```yml # Config fields, showing default values label: "" sync_response: {} ``` For most inputs this mechanism is ignored entirely, in which case the sync response is dropped without penalty. It is therefore safe to use this processor even when combining input types that might not have support for sync responses. --- # Page 246: text_chunker **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/text_chunker.md --- # text_chunker > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: text_chunker latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/text_chunker page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/text_chunker.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/text_chunker.adoc page-git-created-date: "2025-05-02" page-git-modified-date: "2026-05-26" --- Breaks down text-based message content into manageable chunks using a configurable strategy. This processor is ideal for creating vector embeddings of large text documents. #### Common ```yml processors: label: "" text_chunker: strategy: "" # No default (required) chunk_size: 512 chunk_overlap: 100 separators: - "\n\n" - "\n" - " " - "" length_measure: runes include_code_blocks: false keep_reference_links: false ``` #### Advanced ```yml processors: label: "" text_chunker: strategy: "" # No default (required) chunk_size: 512 chunk_overlap: 100 separators: - "\n\n" - "\n" - " " - "" length_measure: runes token_encoding: "" # No default (optional) allowed_special: [] disallowed_special: - "all" include_code_blocks: false keep_reference_links: false ``` ## [](#fields)Fields ### [](#allowed_special)`allowed_special[]` A list of special tokens to include in the output from this processor. **Type**: `array` **Default**: `[]` ### [](#chunk_overlap)`chunk_overlap` The number of characters duplicated in adjacent chunks of text. **Type**: `int` **Default**: `100` ### [](#chunk_size)`chunk_size` The maximum size of each chunk, using the selected [`length_measure`](#length_measure). **Type**: `int` **Default**: `512` ### [](#disallowed_special)`disallowed_special[]` A list of special tokens to exclude from the output of this processor. **Type**: `array` **Default**: ```yaml - "all" ``` ### [](#include_code_blocks)`include_code_blocks` When set to `true`, this processor includes code blocks in the output. **Type**: `bool` **Default**: `false` ### [](#keep_reference_links)`keep_reference_links` When set to `true`, this processor includes reference links in the output. **Type**: `bool` **Default**: `false` ### [](#length_measure)`length_measure` Choose a method to measure the length of a string. **Type**: `string` **Default**: `runes` | Option | Summary | | --- | --- | | graphemes | Use unicode graphemes to determine the length of a string. | | runes | Use the number of codepoints to determine the length of a string. | | token | Use the number of tokens (using the token_encoding tokenizer) to determine the length of a string. | | utf8 | Determine the length of text using the number of utf8 bytes. | ### [](#separators)`separators[]` A list of strings to use as separators between chunks when the [`recursive_character` strategy option](#strategy) is specified. By default, the following separators are tried in turn until one is successful: - Double newlines (\` `) - Single newlines (` ``) - Spaces (`" “,”"``) **Type**: `array` **Default**: ```yaml - "\n\n" - "\n" - " " - "" ``` ### [](#strategy)`strategy` Choose a strategy for breaking content down into chunks. **Type**: `string` | Option | Summary | | --- | --- | | markdown | Split text by markdown headers. | | recursive_character | Split text recursively by characters (defined in separators). | | token | Split text by tokens. | ### [](#token_encoding)`token_encoding` The type of encoding to use for tokenization. **Type**: `string` ```yaml # Examples: token_encoding: cl100k_base # --- token_encoding: r50k_base ``` --- # Page 247: try_catch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/try_catch.md --- # try_catch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: try_catch latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/try_catch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/try_catch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/try_catch.adoc description: Executes a list of child `processors` on each message and, if any of them fail, executes a separate list of `catch` processors to recover from or react to the error. page-git-created-date: "2026-07-09" page-git-modified-date: "2026-08-11" --- Executes a list of child `processors` on each message and, if any of them fail, executes a separate list of `catch` processors to recover from or react to the error. This processor combines the behavior of the [`try`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/try/) and [`catch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/catch/) processors into a single block with an explicit recovery path. Because it contains both the fallible step and its recovery within a single processor, it is the recommended way to handle expected errors when strict error handling (`error_handling.strict`) is enabled. Each message of a batch is processed individually. The `processors` field is executed with "try" semantics: as soon as a processor fails for a given message the remaining `processors` are skipped for that message. Any message that failed is then routed to the `catch` processors. Before they run, the failure is moved off the message: it is stored as a structured object in a metadata field (see `error_metadata`, `error` by default) and the message’s failure flag is **cleared**. The error is therefore available to recovery logic as an ordinary variable rather than as a message property. The object contains: - `what`: the error message. - `name`: the name of the component that failed (when known). - `label`: the label of the component that failed (when set). - `path`: the dot-path of the component that failed (when known). So a recovery mapping reads the failure with, for example, `@error.what` (equivalent to `meta("error").what`). Because the flag is cleared, the `catch` processors run under the normal error semantics, including strict, so a _new_ failure raised while recovering is treated as a fresh error and is not silently tolerated. > 📝 **NOTE** > > Because the failure flag is cleared before the `catch` processors run, the [`error`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#error) and `error_source_*` functions do not report the original failure within the `catch` block; use the metadata object instead. An empty or omitted `catch` simply records the error in metadata and clears the flag (the failure is swallowed). ```yaml pipeline: processors: - try_catch: processors: - resource: foo - resource: bar catch: - mutation: 'root = "failed to process: " + @error.what' ``` In the example above, if either `foo` or `bar` fails for a message then the `mutation` is applied to that message, replacing its contents with a description of the error (read from the metadata object), and the message continues downstream without a failure flag. More information about error handling can be found in [Error Handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). #### Common ```yml processors: label: "" try_catch: processors: [] catch: [] error_metadata: error ``` #### Advanced ```yml processors: label: "" try_catch: processors: [] catch: [] error_metadata: error ``` ## [](#fields)Fields ### [](#catch)`catch[]` A list of processors to execute on each message that failed one of the `processors` above. The message is no longer flagged as failed when these run; the error is available as an object in the metadata field named by `error_metadata` (e.g. `@error.what`). When omitted or empty the error is recorded in metadata and the flag is cleared (the failure is swallowed). **Type**: `array` **Default**: `[]` ### [](#error_metadata)`error_metadata` The metadata key under which the caught error is stored, as an object with a `what` field (the error message) plus `name`, `label` and `path` fields describing the component that failed, before the `catch` processors are executed. **Type**: `string` **Default**: `error` ### [](#processors)`processors[]` A list of processors to execute on each message. If a processor fails for a given message the remaining processors in this list are skipped for that message, and the message is routed to the `catch` processors. **Type**: `array` **Default**: `[]` --- # Page 248: try **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/try.md --- # try > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: try latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/try page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/try.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/try.adoc description: Executes a list of child processors on messages only if no prior processors have failed (or the errors have been cleared). page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a list of child processors on messages only if no prior processors have failed (or the errors have been cleared). ```yml # Config fields, showing default values label: "" try: [] ``` This processor behaves similarly to the [`for_each`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/for_each/) processor, where a list of child processors are applied to individual messages of a batch. However, if a message has failed any prior processor (before or during the try block) then that message will skip all following processors. For example, with the following config: ```yaml pipeline: processors: - resource: foo - try: - resource: bar - resource: baz - resource: buz ``` If the processor `bar` fails for a particular message, that message will skip the processors `baz` and `buz`. Similarly, if `bar` succeeds but `baz` does not then `buz` will be skipped. If the processor `foo` fails for a message then none of `bar`, `baz` or `buz` are executed on that message. This processor is useful for when child processors depend on the successful output of previous processors. This processor can be followed with a [catch](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/catch/) processor for defining child processors to be applied only to failed messages. More information about error handing can be found in [Error Handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ## [](#nest-within-a-catch-block)Nest within a catch block In some cases it might be useful to nest a try block within a catch block, since the [`catch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/catch/) only clears errors _after_ executing its child processors this means a nested try processor will not execute unless the errors are explicitly cleared beforehand. This can be done by inserting an empty catch block before the try block like as follows: ```yaml pipeline: processors: - resource: foo - catch: - log: level: ERROR message: "Foo failed due to: ${! error() }" - catch: [] # Clear prior error - try: - resource: bar - resource: baz ``` --- # Page 249: unarchive **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/unarchive.md --- # unarchive > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: unarchive latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/unarchive page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/unarchive.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/unarchive.adoc description: Unarchives messages according to the selected archive format into multiple messages within a batch. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Unarchives messages according to the selected archive format into multiple messages within a [batch](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ```yml # Config fields, showing default values label: "" unarchive: format: "" # No default (required) ``` When a message is unarchived the new messages replace the original message in the batch. Messages that are selected but fail to unarchive (invalid format) will remain unchanged in the message batch but will be flagged as having failed, allowing you to [error handle them](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). ## [](#metadata)Metadata The metadata found on the messages handled by this processor will be copied into the resulting messages. For the unarchive formats that contain file information (tar, zip), a metadata field is also added to each message called `archive_filename` with the extracted filename. ## [](#fields)Fields ### [](#format)`format` The unarchiving format to apply. **Type**: `string` | Option | Summary | | --- | --- | | binary | Extract messages from a binary blob format. | | csv | Attempt to parse the message as a csv file (header required) and for each row in the file expands its contents into a json object in a new message. | | csv:x | Attempt to parse the message as a csv file (header required) and for each row in the file expands its contents into a json object in a new message using a custom delimiter. The custom delimiter must be a single character, e.g. the format "csv:\t" would consume a tab delimited file. | | json_array | Attempt to parse a message as a JSON array, and extract each element into its own message. | | json_documents | Attempt to parse a message as a stream of concatenated JSON documents. Each parsed document is expanded into a new message. | | json_map | Attempt to parse the message as a JSON map and for each element of the map expands its contents into a new message. A metadata field is added to each message called archive_key with the relevant key from the top-level map. | | lines | Extract the lines of a message each into their own message. | | tar | Extract messages from a unix standard tape archive. | | zip | Extract messages from a zip file. | --- # Page 250: while **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/while.md --- # while > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: while latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/while page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/while.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/while.adoc description: A processor that checks a Bloblang query against each batch of messages and executes child processors on them for as long as the query resolves to true. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A processor that checks a [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) against each batch of messages and executes child processors on them for as long as the query resolves to true. #### Common ```yml processors: label: "" while: at_least_once: false check: "" processors: [] # No default (required) ``` #### Advanced ```yml processors: label: "" while: at_least_once: false max_loops: 0 check: "" processors: [] # No default (required) ``` The field `at_least_once`, if true, ensures that the child processors are always executed at least one time (like a do .. while loop.) The field `max_loops`, if greater than zero, caps the number of loops for a message batch to this value. If following a loop execution the number of messages in a batch is reduced to zero the loop is exited regardless of the condition result. If following a loop execution there are more than 1 message batches the query is checked against the first batch only. The conditions of this processor are applied across entire message batches. You can find out more about batching [in this doc](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/). ## [](#fields)Fields ### [](#at_least_once)`at_least_once` Whether to always run the child processors at least one time. **Type**: `bool` **Default**: `false` ### [](#check)`check` A [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that should return a boolean value indicating whether the while loop should execute again. **Type**: `string` **Default**: `""` ```yaml # Examples: check: errored() # --- check: this.urls.unprocessed.length() > 0 ``` ### [](#max_loops)`max_loops` An optional maximum number of loops to execute. Helps protect against accidentally creating infinite loops. **Type**: `int` **Default**: `0` ### [](#processors)`processors[]` A list of child processors to execute on each loop. **Type**: `array` --- # Page 251: workflow **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/workflow.md --- # workflow > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: workflow latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/workflow page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/workflow.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/workflow.adoc description: Executes a topology of branch processors, performing them in parallel where possible. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Executes a topology of [`branch` processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/), performing them in parallel where possible. ### Common ```yml processors: label: "" workflow: meta_path: meta.workflow order: [] branches: request_map: "" processors: [] # No default (required) result_map: "" ``` ### Advanced ```yml processors: label: "" workflow: meta_path: meta.workflow order: [] branch_resources: [] branches: request_map: "" processors: [] # No default (required) result_map: "" ``` ## [](#why-use-a-workflow)Why use a workflow ### [](#performance)Performance Most of the time the best way to compose processors is also the simplest, just configure them in series. This is because processors are often CPU bound, low-latency, and you can gain vertical scaling by increasing the number of processor pipeline threads, allowing Redpanda Connect to process [multiple messages in parallel](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/processing_pipelines/). However, some processors, such as [`aws_lambda`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/aws_lambda/) and [`cache`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/cache/), interact with external services and therefore spend most of their time waiting for a response. These processors tend to be high-latency and low CPU activity, which causes messages to process slowly. When a processing pipeline contains multiple network processors that aren’t dependent on each other we can benefit from performing these processors in parallel for each individual message, reducing the overall message processing latency. ### [](#simplifying-processor-topology)Simplifying processor topology A workflow is often expressed as a [DAG](https://en.wikipedia.org/wiki/Directed_acyclic_graph) of processing stages, where each stage can result in N possible next stages, until finally the flow ends at an exit node. For example, if we had processing stages A, B, C and D, where stage A could result in either stage B or C being next, always followed by D, it might look something like this: ```text /--> B --\ A --| |--> D \--> C --/ ``` This flow would be easy to express in a standard Redpanda Connect config, we could simply use a [`switch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/switch/) to route to either B or C depending on a condition on the result of A. However, this method of flow control quickly becomes unfeasible as the DAG gets more complicated, imagine expressing this flow using switch processors: ```text /--> B -------------|--> D / / A --| /--> E --| \--> C --| \ \----------|--> F ``` And imagine doing so knowing that the diagram is subject to change over time. Yikes! Instead, with a workflow we can either trust it to automatically resolve the DAG or express it manually as simply as `order: [ [ A ], [ B, C ], [ E ], [ D, F ] ]`, and the conditional logic for determining if a stage is executed is defined as part of the branch itself. ## [](#examples)Examples ### [](#automatic-ordering)Automatic Ordering When the field `order` is omitted a best attempt is made to determine a dependency tree between branches based on their request and result mappings. In the following example the branches foo and bar will be executed first in parallel, and afterwards the branch baz will be executed. ```yaml pipeline: processors: - workflow: meta_path: meta.workflow branches: foo: request_map: 'root = ""' processors: - http: url: TODO result_map: 'root.foo = this' bar: request_map: 'root = this.body' processors: - aws_lambda: function: TODO result_map: 'root.bar = this' baz: request_map: | root.fooid = this.foo.id root.barstuff = this.bar.content processors: - cache: resource: TODO operator: set key: ${! json("fooid") } value: ${! json("barstuff") } ``` ### [](#conditional-branches)Conditional Branches Branches of a workflow are skipped when the `request_map` assigns `deleted()` to the root. In this example the branch A is executed when the document type is "foo", and branch B otherwise. Branch C is executed afterwards and is skipped unless either A or B successfully provided a result at `tmp.result`. ```yaml pipeline: processors: - workflow: branches: A: request_map: | root = if this.document.type != "foo" { deleted() } processors: - http: url: TODO result_map: 'root.tmp.result = this' B: request_map: | root = if this.document.type == "foo" { deleted() } processors: - aws_lambda: function: TODO result_map: 'root.tmp.result = this' C: request_map: | root = if this.tmp.result != null { deleted() } processors: - http: url: TODO_SOMEWHERE_ELSE result_map: 'root.tmp.result = this' ``` ### [](#resources)Resources The `order` field can be used in order to refer to [branch processor resources](#resources), this can sometimes make your pipeline configuration cleaner, as well as allowing you to reuse branch configurations in order places. It’s also possible to mix and match branches configured within the workflow and configured as resources. ```yaml pipeline: processors: - workflow: order: [ [ foo, bar ], [ baz ] ] branches: bar: request_map: 'root = this.body' processors: - aws_lambda: function: TODO result_map: 'root.bar = this' processor_resources: - label: foo branch: request_map: 'root = ""' processors: - http: url: TODO result_map: 'root.foo = this' - label: baz branch: request_map: | root.fooid = this.foo.id root.barstuff = this.bar.content processors: - cache: resource: TODO operator: set key: ${! json("fooid") } value: ${! json("barstuff") } ``` ## [](#fields)Fields ### [](#branch_resources)`branch_resources[]` An optional list of [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) names that are configured as [Resources](#resources). These resources will be included in the workflow with any branches configured inline within the [`branches`](#branches) field. The order and parallelism in which branches are executed is automatically resolved based on the mappings of each branch. When using resources with an explicit order it is not necessary to list resources in this field. **Type**: `array` **Default**: `[]` ### [](#branches)`branches` An object of named [`branch` processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) that make up the workflow. The order and parallelism in which branches are executed can either be made explicit with the field `order`, or if omitted an attempt is made to automatically resolve an ordering based on the mappings of each branch. **Type**: `object` **Default**: `{}` ### [](#branches-processors)`branches.processors[]` A list of processors to apply to mapped requests. When processing message batches the resulting batch must match the size and ordering of the input batch, therefore filtering, grouping should not be performed within these processors. **Type**: `array` ### [](#branches-request_map)`branches.request_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that describes how to create a request payload suitable for the child processors of this branch. If left empty then the branch will begin with an exact copy of the origin message (including metadata). **Type**: `string` **Default**: `""` ```yaml # Examples: request_map: |- root = { "id": this.doc.id, "content": this.doc.body.text } # --- request_map: |- root = if this.type == "foo" { this.foo.request } else { deleted() } ``` ### [](#branches-result_map)`branches.result_map` A [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) that describes how the resulting messages from branched processing should be mapped back into the original payload. If left empty the origin message will remain unchanged (including metadata). **Type**: `string` **Default**: `""` ```yaml # Examples: result_map: |- meta foo_code = metadata("code") root.foo_result = this # --- result_map: |- meta = metadata() root.bar.body = this.body root.bar.id = this.user.id # --- result_map: root.raw_result = content().string() # --- result_map: |- root.enrichments.foo = if metadata("request_failed") != null { throw(metadata("request_failed")) } else { this } # --- result_map: |- # Retain only the updated metadata fields which were present in the origin message meta = metadata().filter(v -> @.get(v.key) != null) ``` ### [](#meta_path)`meta_path` A [dot path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) indicating where to store and reference [structured metadata](#structured-metadata) about the workflow execution. **Type**: `string` **Default**: `meta.workflow` ### [](#order)`order` An explicit declaration of branch ordered tiers, which describes the order in which parallel tiers of branches should be executed. Branches should be identified by the name as they are configured in the field `branches`. It’s also possible to specify branch processors configured [as a resource](#resources). **Type**: `array>` **Default**: `[]` ```yaml # Examples: order: - - foo - bar - - baz # --- order: - - foo - - bar - - baz ``` ## [](#structured-metadata)Structured metadata When the field `meta_path` is non-empty the workflow processor creates an object describing which workflows were successful, skipped or failed for each message and stores the object within the message at the end. The object is of the following form: ```json { "succeeded": [ "foo" ], "skipped": [ "bar" ], "failed": { "baz": "the error message from the branch" } } ``` If a message already has a meta object at the given path when it is processed then the object is used in order to determine which branches have already been performed on the message (or skipped) and can therefore be skipped on this run. This is a useful pattern when replaying messages that have failed some branches previously. For example, given the above example object the branches foo and bar would automatically be skipped, and baz would be reattempted. The previous meta object will also be preserved in the field `.previous` when the new meta object is written, preserving a full record of all workflow executions. If a field `.apply` exists in the meta object for a message and is an array then it will be used as an explicit list of stages to apply, all other stages will be skipped. ## [](#error-handling)Error handling The recommended approach to handle failures within a workflow is to query against the [structured metadata](#structured-metadata) it provides, as it provides granular information about exactly which branches failed and which ones succeeded and therefore aren’t necessary to perform again. For example, if our meta object is stored at the path `meta.workflow` and we wanted to check whether a message has failed for any branch we can do that using a [Bloblang query](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) like `this.meta.workflow.failed.length() | 0 > 0`, or to check whether a specific branch failed we can use `this.exists("meta.workflow.failed.foo")`. However, if structured metadata is disabled by setting the field `meta_path` to empty then the workflow processor instead adds a general error flag to messages when any executed branch fails. In this case it’s possible to handle failures using [standard error handling patterns](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). --- # Page 252: xml **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/xml.md --- # xml > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: xml latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/processors/xml page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/processors/xml.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/processors/xml.adoc description: Parses messages as an XML document, performs a mutation on the data, and then overwrites the previous contents with the new value. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Parses messages as an XML document, performs a mutation on the data, and then overwrites the previous contents with the new value. ```yml # Config fields, showing default values label: "" xml: operator: "" cast: false ``` ## [](#operators)Operators ### [](#to_json)`to_json` Converts an XML document into a JSON structure, where elements appear as keys of an object according to the following rules: - If an element contains attributes they are parsed by prefixing a hyphen, `-`, to the attribute label. - If the element is a simple element and has attributes, the element value is given the key `#text`. - XML comments, directives, and process instructions are ignored. - When elements are repeated the resulting JSON value is an array. - XML namespaces are stripped from element and attribute names, and namespace declarations (`xmlns`) are omitted. For example, given the following XML: ```xml This is a title This is a description foo1 foo2 foo3 ``` The resulting JSON structure would look like this: ```json { "root":{ "title":"This is a title", "description":{ "#text":"This is a description", "-tone":"boring" }, "elements":[ {"#text":"foo1","-id":"1"}, {"#text":"foo2","-id":"2"}, "foo3" ] } } ``` With cast set to true, the resulting JSON structure would look like this: ```json { "root":{ "title":"This is a title", "description":{ "#text":"This is a description", "-tone":"boring" }, "elements":[ {"#text":"foo1","-id":1}, {"#text":"foo2","-id":2}, "foo3" ] } } ``` ## [](#fields)Fields ### [](#cast)`cast` Whether to try to cast values that are numbers and booleans to the right type. Default: all values are strings. **Type**: `bool` **Default**: `false` ### [](#operator)`operator` An XML [operation](#operators) to apply to messages. **Type**: `string` **Default**: `""` **Options**: `to_json` --- # Page 253: Rate Limits **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/about.md --- # Rate Limits > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Rate Limits latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/rate_limits/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/rate_limits/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/rate_limits/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A rate limit is a strategy for limiting the usage of a shared resource across parallel components in a Redpanda Connect instance, or potentially across multiple instances. They are configured as a resource: ```yaml rate_limit_resources: - label: foobar local: count: 500 interval: 1s ``` And most components that hit external services have a field `rate_limit` for specifying a rate limit resource to use, identified by the `label` field. For example, if we wanted to use our `foobar` rate limit with a `http_client` input it would look like this: ```yaml input: http_client: url: TODO verb: GET rate_limit: foobar ``` By using a rate limit in this way we can guarantee that our input will only poll our HTTP source at the rate of 500 requests per second. Some components don’t have a `rate_limit` field but we might still wish to throttle them by a rate limit, in which case we can use the [`rate_limit` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/rate_limit/) that applies back pressure to a processing pipeline when the limit is reached. --- # Page 254: local **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/local.md --- # local > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: local latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/rate_limits/local page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/rate_limits/local.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/rate_limits/local.adoc description: A simple X every Y rate limit that can be shared across components within a pipeline. It does not support distributed rate limiting across instances. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- The local rate limit is a simple X every Y type rate limit that can be shared across any number of components within the pipeline but does not support distributed rate limits across multiple running instances of Benthos. ```yml # Config fields, showing default values label: "" local: count: 1000 interval: 1s ``` ## [](#fields)Fields ### [](#count)`count` The maximum number of requests to allow for a given period of time. **Type**: `int` **Default**: `1000` ### [](#interval)`interval` The time window to limit requests by. **Type**: `string` **Default**: `"1s"` --- # Page 255: redis **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/rate_limits/redis.md --- # redis > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redis latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/rate_limits/redis page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/rate_limits/redis.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/rate_limits/redis.adoc description: A token bucket rate limit backed by Redis, shared across all Redpanda Connect instances that use the same Redis instance. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- A rate limit implementation using Redis. It works by using a simple token bucket algorithm to limit the number of requests to a given count within a given time period. The rate limit is shared across all instances of Redpanda Connect that use the same Redis instance, which must all have a consistent count and interval. #### Common ```yml # Common config fields, showing default values label: "" redis: url: redis://:6379 # No default (required) count: 1000 interval: 1s key: "" # No default (required) ``` #### Advanced ```yml # All config fields, showing default values label: "" redis: url: redis://:6379 # No default (required) kind: simple master: "" tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] count: 1000 interval: 1s key: "" # No default (required) ``` ## [](#fields)Fields ### [](#url)`url` The URL of the target Redis server. Database is optional and is supplied as the URL path. **Type**: `string` ```yml # Examples url: redis://:6379 url: redis://localhost:6379 url: redis://foousername:foopassword@redisplace:6379 url: redis://:foopassword@redisplace:6379 url: redis://localhost:6379/1 url: redis://localhost:6379/1,redis://localhost:6380/1 ``` ### [](#kind)`kind` Specifies a simple, cluster-aware, or failover-aware redis client. **Type**: `string` **Default**: `"simple"` Options: `simple` , `cluster` , `failover` . ### [](#master)`master` Name of the redis master when `kind` is `failover` **Type**: `string` **Default**: `""` ```yml # Examples master: mymaster ``` ### [](#tls)`tls` Custom TLS settings can be used to override system defaults. **Troubleshooting** Some cloud hosted instances of Redis (such as Azure Cache) might need some hand holding in order to establish stable connections. Unfortunately, it is often the case that TLS issues will manifest as generic error messages such as "i/o timeout". If you’re using TLS and are seeing connectivity problems consider setting `enable_renegotiation` to `true`, and ensuring that the server supports at least TLS version 1.2. **Type**: `object` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yml # Examples root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yml # Examples root_cas_file: ./root_cas.pem ``` ### [](#tls-client_certs)`tls.client_certs` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yml # Examples client_certs: - cert: foo key: bar client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yml # Examples password: foo password: ${KEY_PASSWORD} ``` ### [](#count)`count` The maximum number of messages to allow for a given period of time. **Type**: `int` **Default**: `1000` ### [](#interval)`interval` The time window to limit requests by. **Type**: `string` **Default**: `"1s"` ### [](#key)`key` The key to use for the rate limit. **Type**: `string` --- # Page 256: redpanda **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/redpanda/about.md --- # redpanda > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/redpanda/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/redpanda/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/redpanda/about.adoc page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- The Redpanda Connect configuration service allows you to: - Configure Redpanda cluster credentials in a single configuration block, which is referenced by multiple components in data pipeline. For more information, see the [Pipeline example](#pipeline-example). - Send logs and status updates to topics on a Redpanda cluster, in addition to the [default logger](https://docs.redpanda.com/connect/components/logger/about/). The `redpanda` namespace contains the configuration of this service. #### Common ```yml # Common configuration fields, showing default values redpanda: seed_brokers: [] # No default (optional) pipeline_id: "" logs_topic: "" logs_level: info status_topic: "" ``` #### Advanced ```yml # All configuration fields, showing default values redpanda: seed_brokers: [] # No default (optional) client_id: benthos tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 5m request_timeout_overhead: 10s conn_idle_timeout: 20s pipeline_id: "" logs_topic: "" logs_level: info status_topic: "" partitioner: "" # No default (optional) idempotent_write: true compression: "" # No default (optional) timeout: 10s max_message_bytes: 1MB broker_write_max_bytes: 100MB allow_auto_topic_creation: true ``` ## [](#pipeline-example)Pipeline example This data pipeline reads data from `topic_A` and `topic_B` on a Redpanda cluster, and then writes the data to `topic_C` on the same cluster. The cluster details are configured within the `redpanda` configuration block, so you only need to configure them once. This is a useful feature when you have multiple inputs and outputs in the same data pipeline that need to connect to the same cluster. ```none input: redpanda_common: topics: [ topic_A, topic_B ] output: redpanda_common: topic: topic_C key: ${! @id } redpanda: seed_brokers: [ "127.0.0.1:9092" ] tls: enabled: true sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ## [](#fields)Fields ### [](#seed_brokers)`seed_brokers` A list of broker addresses to connect to in order. Use commas to separate multiple addresses in a single list item. **Type**: `array` ```yml # Examples seed_brokers: - localhost:9092 seed_brokers: - foo:9092 - bar:9092 seed_brokers: - foo:9092,bar:9092 ``` ### [](#client_id)`client_id` An identifier for the client connection. **Type**: `string` **Default**: `benthos` ### [](#tls)`tls` Override system defaults with custom TLS settings. **Type**: `object` ### [](#tls-enabled)`tls.enabled` Whether custom TLS settings are enabled. **Type**: `bool` **Default**: `false` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. **Type**: `bool` **Default**: `false` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent trusted root certificate, through possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yml # Example root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent trusted root certificate, through possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yml # Example root_cas_file: ./root_cas.pem ``` ### [](#tls-client_certs)`tls.client_certs` A list of client certificates to use. For each certificate, specify either the fields `cert` and `key` or `cert_file` and `key_file`. **Type**: `array` **Default**: `[]` ```yml # Examples client_certs: - cert: foo key: bar client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` The plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` The plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` The plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. > ⚠️ **WARNING** > > The `pbeWithMD5AndDES-CBC` algorithm does not authenticate ciphertext, and is vulnerable to padding oracle attacks which may allow an attacker to recover the plain text password. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yml # Examples password: foo password: ${KEY_PASSWORD} ``` ### [](#sasl)`sasl` Specify one or more methods or mechanisms of SASL authentication. They are tried in order. If the broker supports the first SASL mechanism, all connections use it. If the first mechanism fails, the client picks the first supported mechanism. If the broker does not support any client mechanisms, all connections fail. **Type**: `array` ```yml # Example sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM-based authentication as specified by the aws-msk-iam-auth Java library. | | OAUTHBEARER | OAuth Bearer-based authentication. | | PLAIN | Plain text authentication. | | SCRAM-SHA-256 | SCRAM-based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM-based authentication as specified in RFC5802. | | none | Disable SASL authentication | ### [](#sasl-username)`sasl[].username` A username for `PLAIN` or `SCRAM-*` authentication. **Type**: `string` **Default**: `""` ### [](#sasl-password)`sasl[].password` A password for `PLAIN` or `SCRAM-*` authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s `OAUTHBEARER` authentication. **Type**: `string` **Default**: `""` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to `OAUTHBEARER` authentication requests. **Type**: `object` ### [](#sasl-aws)`sasl[].aws` AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` **Default**: `""` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Specify a custom endpoint for the AWS API. **Type**: `string` **Default**: `""` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Manually configure the AWS credentials to use (optional). For more information, see the [Amazon Web Services guide](https://docs.redpanda.com/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` The profile from `~/.aws/credentials` to use. **Type**: `string` **Default**: `""` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of the AWS credentials to use. **Type**: `string` **Default**: `""` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the AWS credentials in use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the AWS credentials in use. This is a required value for short-term credentials. **Type**: `string` **Default**: `""` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume an [IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` **Default**: `false` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` The role ARN to assume. **Type**: `string` **Default**: `""` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to use when assuming a role. **Type**: `string` **Default**: `""` ### [](#metadata_max_age)`metadata_max_age` The maximum period of time after which metadata is refreshed. **Type**: `string` **Default**: `5m` ### [](#request_timeout_overhead)`request_timeout_overhead` Grants an additional buffer or overhead to requests that have timeout fields defined. This field is based on the behavior of Apache Kafka’s `request.timeout.ms` parameter, but with the option to extend the timeout deadline. **Type**: `string` **Default**: `10s` ### [](#conn_idle_timeout)`conn_idle_timeout` Define how long connections can remain idle before they are closed. **Type**: `string` ### [](#pipeline_id)`pipeline_id` The ID of a Redpanda Connect data pipeline (optional). When specified, the pipeline ID is written to all logs and status updates sent to the configured topics. **Type**: `string` **Default**: `""` ### [](#logs_topic)`logs_topic` The topic that logs are sent to. **Type**: `string` **Default**: `""` ```yml # Example logs_topic: __redpanda.connect.logs ``` ### [](#logs_level)`logs_level` The logging level of logs sent to Redpanda. **Type**: `string` **Default**: `info` **Options**: `debug`, `info`, `warn`, `error` ### [](#status_topic)`status_topic` The topic that status updates are sent to. When configured, Redpanda Connect emits status events to this internal topic as pipelines start, run, and stop. **Type**: `string` **Default**: `""` ```yml # Example status_topic: __redpanda.connect.status ``` For monitoring and troubleshooting guidance, see [Monitor Pipeline Status](https://docs.redpanda.com/connect/guides/monitor-pipeline-status/). #### [](#status-event-types)Status event types Status events are emitted during the pipeline lifecycle. All events include `pipeline_id`, `instance_id`, and `timestamp` fields. **TYPE\_INITIALIZING** (value: 1) Emitted when a pipeline instance has successfully parsed its configuration and is attempting to start. This is the first event sent after pipeline startup. Example: ```json { "type": "TYPE_INITIALIZING", "pipeline_id": "my-pipeline", "instance_id": "abc123xyz", "timestamp": 1717890000 } ``` **TYPE\_CONNECTION\_HEALTHY** (value: 2) Emitted every 30 seconds (heartbeat) when all pipeline connections (inputs and outputs) are active and functioning normally. Example: ```json { "type": "TYPE_CONNECTION_HEALTHY", "pipeline_id": "my-pipeline", "instance_id": "abc123xyz", "timestamp": 1717890030 } ``` **TYPE\_CONNECTION\_ERROR** (value: 3) Emitted every 30 seconds (during the same heartbeat cycle) when one or more connections are inactive or experiencing errors. The event includes detailed error information for each failing connection: - `path`: The configuration path of the connector (see [field paths](https://docs.redpanda.com/connect/configuration/field_paths/)) - `label`: Optional label assigned to the connector - `message`: The error message describing the connection failure Example: ```json { "type": "TYPE_CONNECTION_ERROR", "pipeline_id": "my-pipeline", "instance_id": "abc123xyz", "timestamp": 1717890060, "connection_errors": [ { "path": "input.kafka_franz", "label": "primary-input", "message": "kafka: connection refused" } ] } ``` **TYPE\_EXITING** (value: 4) Emitted when a pipeline instance is shutting down, either gracefully or due to an error. If shutdown was caused by an error, the event includes an `exit_error` object with the error message. Example (graceful shutdown): ```json { "type": "TYPE_EXITING", "pipeline_id": "my-pipeline", "instance_id": "abc123xyz", "timestamp": 1717890090 } ``` Example (error shutdown): ```json { "type": "TYPE_EXITING", "pipeline_id": "my-pipeline", "instance_id": "abc123xyz", "timestamp": 1717890090, "exit_error": { "message": "failed to process message: invalid format" } } ``` #### [](#message-format)Message format Status events are written to the configured topic as Protocol Buffer JSON messages: - **Topic**: The value specified in `status_topic` - **Key**: The `pipeline_id` value (enables tracking all events for a specific pipeline) - **Value**: JSON-encoded status event - **Encoding**: Protocol Buffer JSON (protojson) For the full protobuf schema specification, see the [status.proto definition](https://github.com/redpanda-data/connect/blob/main/proto/redpanda/api/connect/v1alpha1/status.proto). ### [](#partitioner)`partitioner` Override the default murmur2 hashing partitioner. **Type**: `string` | Option | Summary | | --- | --- | | least_backup | Chooses the least backed up partition. The partition with the fewest buffered records. Partitions are selected per batch. | | manual | Manually select a partition for each message. You must also specify a value for the partition field. | | murmur2_hash | Kafka’s default hash algorithm that uses a 32-bit murmur2 hash of the key to compute the partition for the record. | | round_robin | Does a round robin of messages through all available partitions. This algorithm has lower throughput and causes higher CPU load on brokers, but is useful if you want to ensure an even distribution of records to partitions. | ### [](#idempotent_write)`idempotent_write` Enable the idempotent write producer option. This requires the `IDEMPOTENT_WRITE` permission on `CLUSTER`. Disable this option if the `IDEMPOTENT_WRITE` permission is not available. **Type**: `bool` **Default**: `true` ### [](#compression)`compression` Set an explicit compression type (optional). The default preference is to use `snappy` when the broker supports it. Otherwise, use `none`. **Type**: `string` Options: `lz4` , `snappy` , `gzip` , `none` , `zstd` ### [](#timeout)`timeout` The maximum period of time allowed for sending log or status update messages before a request is abandoned and a retry attempted. **Type**: `string` **Default**: `10s` ### [](#max_message_bytes)`max_message_bytes` The maximum size of an individual message in bytes. Messages larger than this value are rejected. This field is equivalent to Kafka’s `max.message.bytes`. **Type**: `string` **Default**: `1MB` ```yml # Examples max_message_bytes: 100MB max_message_bytes: 50mib ``` ### [](#broker_write_max_bytes)`broker_write_max_bytes` The upper bound for the number of bytes written to a broker connection in a single write. This field corresponds to Kafka’s `socket.request.max.bytes`. **Type**: `string` **Default**: `"100MB"` ```yml # Examples broker_write_max_bytes: 128MB broker_write_max_bytes: 50mib ``` --- # Page 257: Scanners **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/about.md --- # Scanners > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Scanners latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- For such inputs it’s necessary to define a mechanism by which the stream of source bytes can be chopped into smaller logical messages, processed and outputted as a continuous process whilst the stream is being read, as this dramatically reduces the memory usage of Redpanda Connect as a whole and results in a more fluid flow of data. The way in which we define this chopping mechanism is through scanners, configured as a field on each input that requires one. For example, if we wished to consume files line-by-line, which each individual line being processed as a discrete message, we could use the [`lines` scanner](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/lines/) with our `file` input: ## Common ```yaml input: file: paths: [ "./*.txt" ] scanner: lines: {} ``` ## Advanced ```yaml # Instead of newlines, use a custom delimiter: input: file: paths: [ "./*.txt" ] scanner: lines: custom_delimiter: "---END---" max_buffer_size: 100_000_000 # 100MB line buffer ``` A scanner is a plugin similar to any other core Redpanda Connect component (inputs, processors, outputs, etc), which means it’s possible to define your own scanners that can be utilized by inputs that need them. --- # Page 258: avro **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/avro.md --- # avro > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: avro latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/avro page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/avro.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/avro.adoc description: Consume a stream of Avro OCF datum. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consume a stream of Avro OCF datum. #### Common ```yml scanners: avro: {} ``` #### Advanced ```yml scanners: avro: raw_json: false ``` ## [](#avro-json-format)Avro JSON format This scanner creates documents formatted as [Avro JSON](https://avro.apache.org/docs/current/specification/) when decoding with Avro schemas. In this format, the value of a union is encoded in JSON as follows: - If the union’s type is `null`, it is encoded as a JSON `null`. - Otherwise, the union is encoded as a JSON object with one name/value pair. The `"name"` is the type’s name and the `"value"` is the recursively encoded value. For Avro’s named types (record, fixed or enum), the user-specified name is used. For other types, the type name is used. For example, the union schema `["null","string","Transaction"]`, where `Transaction` is a record name, would encode: - The `null` as a JSON `null` - The string `"a"` as `{"string": "a"}` - A `Transaction` instance as `{"Transaction": {…​}}`, where `{…​}` indicates the JSON encoding of a `Transaction` instance Alternatively, you can create documents in [standard/raw JSON format](https://pkg.go.dev/github.com/linkedin/goavro/v2#NewCodecForStandardJSONFull) by setting the field [`raw_json`](#raw_json) to `true`. ## [](#metadata)Metadata This scanner emits the following metadata for each message: - The `@avro_schema` field: The canonical Avro schema. - The `@avro_schema_fingerprint` field: The schema ID or fingerprint. ## [](#fields)Fields ### [](#raw_json)`raw_json` Whether to decode messages into normal JSON rather than [Avro JSON](https://avro.apache.org/docs/current/specification/_print/#json-encoding). When true, this unwraps union values (bare values instead of {"type": value} wrappers). **Type**: `bool` **Default**: `false` --- # Page 259: chunker **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/chunker.md --- # chunker > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: chunker latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/chunker page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/chunker.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/chunker.adoc description: Split an input stream into chunks of a given number of bytes. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Split an input stream into chunks of a given number of bytes. ```yml # Config fields, showing default values chunker: size: 0 # No default (required) ``` ## [](#fields)Fields ### [](#size)`size` The size of each chunk in bytes. **Type**: `int` --- # Page 260: csv **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/csv.md --- # csv > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: csv latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/csv page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/csv.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/csv.adoc description: Consume comma-separated values row by row, including support for custom delimiters. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consume comma-separated values row by row, including support for custom delimiters. ```yml # Config fields, showing default values csv: custom_delimiter: "" # No default (optional) parse_header_row: true lazy_quotes: false continue_on_error: false ``` ## [](#metadata)Metadata This scanner adds the following metadata to each message: - `csv_row` The index of each row, beginning at 0. ## [](#fields)Fields ### [](#continue_on_error)`continue_on_error` If a row fails to parse due to any error emit an empty message marked with the error and then continue consuming subsequent rows when possible. This can sometimes be useful in situations where input data contains individual rows which are malformed. However, when a row encounters a parsing error it is impossible to guarantee that following rows are valid, as this indicates that the input data is unreliable and could potentially emit misaligned rows. **Type**: `bool` **Default**: `false` ### [](#custom_delimiter)`custom_delimiter` Use a provided custom delimiter instead of the default comma. **Type**: `string` ### [](#lazy_quotes)`lazy_quotes` If set to `true`, a quote may appear in an unquoted field and a non-doubled quote may appear in a quoted field. **Type**: `bool` **Default**: `false` ### [](#parse_header_row)`parse_header_row` Whether to reference the first row as a header row. If set to true the output structure for messages will be an object where field keys are determined by the header row. Otherwise, each message will consist of an array of values from the corresponding CSV row. **Type**: `bool` **Default**: `true` --- # Page 261: decompress **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/decompress.md --- # decompress > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: decompress latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/decompress page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/decompress.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/decompress.adoc description: Decompress the stream of bytes according to an algorithm, before feeding it into a child scanner. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Decompress the stream of bytes according to an algorithm, before feeding it into a child scanner. ```yml # Config fields, showing default values decompress: algorithm: "" # No default (required) into: to_the_end: {} ``` ## [](#fields)Fields ### [](#algorithm)`algorithm` One of `gzip`, `pgzip`, `zlib`, `bzip2`, `flate`, `snappy`, `lz4`, `zstd`. **Type**: `string` ### [](#into)`into` The child scanner to feed the decompressed stream into. **Type**: `scanner` **Default**: ```yaml to_the_end: {} ``` --- # Page 262: json_array **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/json_array.md --- # json_array > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: json_array latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/json_array page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/json_array.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/json_array.adoc description: Consumes a stream of one or more JSON elements within a top level array. page-git-created-date: "2025-09-26" page-git-modified-date: "2026-08-11" --- Consumes a stream of one or more JSON elements within a top level array. This scanner is useful for: - Processing exports from systems that generate a JSON array as the top-level JSON structure (for example, logs, bulk exports, etc). - Efficiently breaking up large files with many objects into individual events/messages. Suppose you have a file `events.json`: `events.json` ```json [ {"event": "login", "user": "alice"}, {"event": "logout", "user": "bob"}, {"event": "purchase", "user": "carol", "amount": 42} ] ``` The configuration to process this file is: ```yaml input: file: paths: [ "./events.json" ] scanner: json_array: {} ``` Result: Each event in the array is processed as a separate message. ## [](#requirements)Requirements The `json_array` scanner expects the input to be a single JSON array, where each array element is a JSON object or value. ## [](#fields)Fields The `json_array` scanner has no required fields. You declare it as `{}` in your config. ```yaml json_array: {} ``` --- # Page 263: json_documents **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/json_documents.md --- # json_documents > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: json_documents latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/json_documents page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/json_documents.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/json_documents.adoc description: Consumes a stream of one or more JSON documents. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consumes a stream of one or more JSON documents. ```yml # Config fields, showing default values json_documents: {} ``` --- # Page 264: lines **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/lines.md --- # lines > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: lines latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/lines page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/lines.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/lines.adoc description: Split an input stream into a message per line of data. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Split an input stream into a message per line of data. ```yml # Config fields, showing default values lines: custom_delimiter: "" # No default (optional) max_buffer_size: 65536 omit_empty: false ``` ## [](#fields)Fields ### [](#custom_delimiter)`custom_delimiter` Use a provided custom delimiter for detecting the end of a line rather than a single line break. **Type**: `string` ### [](#max_buffer_size)`max_buffer_size` Set the maximum buffer size for storing line data, this limits the maximum size that a line can be without causing an error. **Type**: `int` **Default**: `65536` ### [](#omit_empty)`omit_empty` Omit empty lines. **Type**: `bool` **Default**: `false` --- # Page 265: re_match **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/re_match.md --- # re_match > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: re_match latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/re_match page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/re_match.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/re_match.adoc description: Split an input stream into segments matching against a regular expression. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Split an input stream into segments matching against a regular expression. ```yml # Config fields, showing default values re_match: pattern: (?m)^\d\d:\d\d:\d\d # No default (required) max_buffer_size: 65536 ``` ## [](#fields)Fields ### [](#max_buffer_size)`max_buffer_size` Set the maximum buffer size for storing line data, this limits the maximum size that a message can be without causing an error. **Type**: `int` **Default**: `65536` ### [](#pattern)`pattern` The pattern to match against. **Type**: `string` ```yaml # Examples: pattern: (?m)^\d\d:\d\d:\d\d ``` --- # Page 266: skip_bom **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/skip_bom.md --- # skip_bom > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: skip_bom latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/skip_bom page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/skip_bom.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/skip_bom.adoc description: Skip one or more byte order marks for each opened child scanner. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Skip one or more byte order marks for each opened child scanner. ```yml # Config fields, showing default values skip_bom: into: to_the_end: {} ``` ## [](#fields)Fields ### [](#into)`into` The child scanner to feed the resulting stream into. **Type**: `scanner` **Default**: ```yaml to_the_end: {} ``` --- # Page 267: switch **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/switch.md --- # switch > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: switch latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/switch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/switch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/switch.adoc description: Select a child scanner dynamically for source data based on factors such as the filename. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Select a child scanner dynamically for source data based on factors such as the filename. ```yml # Config fields, showing default values switch: [] # No default (required) ``` This scanner outlines a list of potential child scanner candidates to be chosen, and for each source of data the first candidate to pass will be selected. A candidate without any conditions acts as a catch-all and will pass for every source, it is recommended to always have a catch-all scanner at the end of your list. If a given source of data does not pass a candidate an error is returned and the data is rejected. ## [](#fields)Fields ### [](#re_match_name)`re_match_name` A regular expression to test against the name of each source of data fed into the scanner (filename or equivalent). If this pattern matches the child scanner is selected. **Type**: `string` ### [](#scanner)`scanner` The scanner to activate if this candidate passes. **Type**: `scanner` ## [](#examples)Examples ### [](#switch-based-on-file-name)Switch based on file name In this example a file input chooses a scanner based on the extension of each file ```yaml input: file: paths: [ ./data/* ] scanner: switch: - re_match_name: '\.avro$' scanner: { avro: {} } - re_match_name: '\.csv$' scanner: { csv: {} } - re_match_name: '\.csv.gz$' scanner: decompress: algorithm: gzip into: csv: {} - re_match_name: '\.tar$' scanner: { tar: {} } - re_match_name: '\.tar.gz$' scanner: decompress: algorithm: gzip into: tar: {} - scanner: { to_the_end: {} } ``` --- # Page 268: tar **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/tar.md --- # tar > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: tar latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/tar page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/tar.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/tar.adoc description: Consume a tar archive file by file. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Consume a tar archive file by file. ```yml # Config fields, showing default values tar: {} ``` ## [](#metadata)Metadata This scanner adds the following metadata to each message: - `tar_name` --- # Page 269: to_the_end **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/scanners/to_the_end.md --- # to_the_end > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: to_the_end latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/scanners/to_the_end page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/scanners/to_the_end.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/scanners/to_the_end.adoc description: Read the input stream all the way until the end and deliver it as a single message. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Read the input stream all the way until the end and deliver it as a single message. ```yml # Config fields, showing default values to_the_end: {} ``` > ⚠️ **CAUTION** > > Some sources of data may not have a logical end, therefore caution should be made to exclusively use this scanner when the end of an input stream is clearly defined (and well within memory). --- # Page 270: Tracers **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/about.md --- # Tracers > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Tracers latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/tracers/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/tracers/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/tracers/about.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- When a tracer is configured all messages will be allocated a root span during ingestion that represents their journey through a Redpanda Connect pipeline. Many Redpanda Connect processors create spans, and so tracing is a great way to analyse the pathways of individual messages as they progress through a Redpanda Connect instance. Some inputs, such as `http_server` and `http_client`, are capable of extracting a root span from the source of the message (HTTP headers). This is a work in progress and should eventually expand so that all inputs have a way of doing so. Other inputs, such as `kafka` can be configured to extract a root span by using the `extract_tracing_map` field. A tracer config section looks like this: ```yaml tracer: jaeger: agent_address: localhost:6831 sampler_type: const sampler_param: 1 ``` > ⚠️ **CAUTION** > > Although the configuration spec of this component is stable the format of spans, tags and logs created by Redpanda Connect is subject to change as it is tuned for improvement. --- # Page 271: gcp_cloudtrace **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/gcp_cloudtrace.md --- # gcp_cloudtrace > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: gcp_cloudtrace latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/tracers/gcp_cloudtrace page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/tracers/gcp_cloudtrace.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/tracers/gcp_cloudtrace.adoc description: Send tracing events to Google Cloud Trace. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Send tracing events to a [Google Cloud Trace](https://cloud.google.com/trace). #### Common ```yml tracers: gcp_cloudtrace: project: "" # No default (required) sampling_ratio: 1 flush_interval: "" # No default (optional) ``` #### Advanced ```yml tracers: gcp_cloudtrace: project: "" # No default (required) sampling_ratio: 1 tags: {} flush_interval: "" # No default (optional) ``` ## [](#fields)Fields ### [](#flush_interval)`flush_interval` The period of time between each flush of tracing spans. **Type**: `string` ### [](#project)`project` The google project with Cloud Trace API enabled. If this is omitted then the Google Cloud SDK will attempt auto-detect it from the environment. **Type**: `string` ### [](#sampling_ratio)`sampling_ratio` Sets the ratio of traces to sample. Tuning the sampling ratio is recommended for high-volume production workloads. **Type**: `float` **Default**: `1` ```yaml # Examples: sampling_ratio: 1 ``` ### [](#tags)`tags` A map of tags to add to tracing spans. **Type**: `object` **Default**: `{}` --- # Page 272: none **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/none.md --- # none > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: none latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/tracers/none page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/tracers/none.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/tracers/none.adoc description: Do not send tracing events anywhere. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Do not send tracing events anywhere. ```yml # Config fields, showing default values tracer: none: {} ``` --- # Page 273: open_telemetry_collector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/open_telemetry_collector.md --- # open_telemetry_collector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: open_telemetry_collector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/tracers/open_telemetry_collector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/tracers/open_telemetry_collector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/tracers/open_telemetry_collector.adoc description: Send tracing events to an Open Telemetry collector. page-git-created-date: "2026-05-28" page-git-modified-date: "2026-08-11" --- Send tracing events to an [Open Telemetry collector](https://opentelemetry.io/docs/collector/). #### Common ```yml tracers: open_telemetry_collector: service: benthos http: [] # No default (required) grpc: [] # No default (required) sampling: enabled: false ratio: "" # No default (optional) ``` #### Advanced ```yml tracers: open_telemetry_collector: service: benthos http: [] # No default (required) grpc: [] # No default (required) tags: {} sampling: enabled: false ratio: "" # No default (optional) ``` ## [](#fields)Fields ### [](#grpc)`grpc[]` A list of grpc collectors. **Type**: `array` ### [](#grpc-address)`grpc[].address` The endpoint of a collector to send events to. **Type**: `string` ```yaml # Examples: address: localhost:4317 ``` ### [](#grpc-secure)`grpc[].secure` Connect to the collector with client transport security **Type**: `bool` **Default**: `false` ### [](#http)`http[]` A list of http collectors. **Type**: `array` ### [](#http-address)`http[].address` The endpoint of a collector to send events to. **Type**: `string` ```yaml # Examples: address: localhost:4318 ``` ### [](#http-secure)`http[].secure` Connect to the collector over HTTPS **Type**: `bool` **Default**: `false` ### [](#sampling)`sampling` Settings for trace sampling. Sampling is recommended for high-volume production workloads. **Type**: `object` ### [](#sampling-enabled)`sampling.enabled` Whether to enable sampling. **Type**: `bool` **Default**: `false` ### [](#sampling-ratio)`sampling.ratio` Sets the ratio of traces to sample. **Type**: `float` ```yaml # Examples: ratio: 0.85 # --- ratio: 0.5 ``` ### [](#service)`service` The name of the service in traces. **Type**: `string` **Default**: `benthos` ### [](#tags)`tags` A map of tags to add to all exported spans and metrics. **Type**: `object` **Default**: `{}` --- # Page 274: redpanda **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/redpanda.md --- # redpanda > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: redpanda latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/components/tracers/redpanda page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/components/tracers/redpanda.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/components/tracers/redpanda.adoc description: Send tracing events to a Redpanda topic. page-git-created-date: "2025-12-03" page-git-modified-date: "2026-08-11" --- Export distributed tracing data to a Redpanda topic, enabling you to monitor and debug your Redpanda Connect pipelines. Traces are exported in OpenTelemetry format as JSON, allowing integration with observability platforms like Jaeger, Grafana Tempo, or custom trace consumers. #### Common ```yml tracers: redpanda: seed_brokers: [] # No default (required) topic: otel-traces format: json schema_registry: url: "" # No default (optional) tls: skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} service: redpanda-connect sampling: enabled: false ratio: "" # No default (optional) ``` #### Advanced ```yml tracers: redpanda: seed_brokers: [] # No default (required) client_id: redpanda-connect tls: enabled: false skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] sasl: [] # No default (optional) metadata_max_age: 1m request_timeout_overhead: 10s conn_idle_timeout: 20s tcp: connect_timeout: 0s keep_alive: idle: 15s interval: 15s count: 9 tcp_user_timeout: 0s partitioner: "" # No default (optional) idempotent_write: true acks: all compression: "" # No default (optional) allow_auto_topic_creation: true timeout: 10s max_message_bytes: 1MiB broker_write_max_bytes: 100MiB max_buffered_records: 10000 max_buffered_bytes: 0 max_in_flight_requests: 1 record_retries: 0 record_delivery_timeout: 0s topic: otel-traces format: json schema_registry: url: "" # No default (optional) tls: skip_cert_verify: false enable_renegotiation: false root_cas: "" root_cas_file: "" client_certs: [] oauth2: enabled: false client_key: "" client_secret: "" token_url: "" scopes: [] endpoint_params: {} oauth: enabled: false consumer_key: "" consumer_secret: "" access_token: "" access_token_secret: "" basic_auth: enabled: false username: "" password: "" jwt: enabled: false private_key_file: "" signing_method: "" claims: {} headers: {} service: redpanda-connect tags: {} sampling: enabled: false ratio: "" # No default (optional) ``` This tracer automatically captures trace spans as messages flow through your pipeline, recording timing information, component metadata, and error details. Use this to: - **Track message flow** through complex pipelines with multiple processors. - **Identify performance bottlenecks** by analyzing span durations. - **Debug failures** by examining trace context and error details. - **Monitor pipeline health** across distributed Redpanda Connect instances. - **Correlate activity** across multiple services using trace IDs. The tracer writes to a dedicated Redpanda topic that can be consumed by trace analysis tools. Configure sampling to control trace volume in high-throughput environments. ## [](#fields)Fields ### [](#acks)`acks` The number of acknowledgements the leader broker must receive from ISR brokers before responding to the produce request. When `idempotent_write` is enabled this must be set to `all`. **Type**: `string` **Default**: `all` | Option | Summary | | --- | --- | | all | Wait for all in-sync replicas to acknowledge (acks=-1). Required when idempotent_write is enabled. | | leader | Wait for the leader broker to acknowledge (acks=1). Messages are lost if the leader fails before replication. | | none | Do not wait for any acknowledgement (acks=0). Highest throughput but messages may be lost. | ### [](#allow_auto_topic_creation)`allow_auto_topic_creation` Whether to automatically create the trace topic if it doesn’t exist. If false, the topic must be created manually before starting the tracer. **Type**: `bool` **Default**: `true` ### [](#broker_write_max_bytes)`broker_write_max_bytes` The maximum number of bytes this output can write to a broker connection in a single write. This field corresponds to Kafka’s `socket.request.max.bytes`. **Type**: `string` **Default**: `100MiB` ```yaml # Examples: broker_write_max_bytes: 128MB # --- broker_write_max_bytes: 50mib ``` ### [](#client_id)`client_id` An identifier for the client connection. This appears in broker logs and metrics to help identify which Redpanda Connect instance is sending traces. **Type**: `string` **Default**: `redpanda-connect` ### [](#compression)`compression` Compression codec to use for trace messages. Options include `gzip`, `snappy`, `lz4`, `zstd`, or none. Compression can reduce network bandwidth and storage costs. **Type**: `string` **Options**: `lz4`, `snappy`, `gzip`, `none`, `zstd` ### [](#conn_idle_timeout)`conn_idle_timeout` The maximum duration that connections can remain idle before they are automatically closed. This field accepts Go duration format strings such as `100ms`, `1s`, or `5s`. **Type**: `string` **Default**: `20s` ### [](#format)`format` The format for trace data. Currently only `json` is supported, which exports OpenTelemetry spans as JSON messages. **Type**: `string` **Default**: `json` | Option | Summary | | --- | --- | | json | Emit in JSON Format | | protobuf | Emit in Protobuf Format | | schema-registry-json | Emit in JSON Format with Schema Registry encoding | | schema-registry-protobuf | Emit in Protobuf Format with Schema Registry encoding | ### [](#idempotent_write)`idempotent_write` Enable idempotent writes to prevent duplicate trace messages in case of retries. Recommended for production environments. **Type**: `bool` **Default**: `true` ### [](#max_buffered_bytes)`max_buffered_bytes` The maximum number of bytes the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered. Set to `0` to disable the byte-level limit (only `max_buffered_records` applies). This limit is checked after `max_buffered_records`. **Type**: `string` **Default**: `0` ```yaml # Examples: max_buffered_bytes: 256MB # --- max_buffered_bytes: 50mib ``` ### [](#max_buffered_records)`max_buffered_records` The maximum number of records the client will buffer in memory before blocking. When this limit is reached, `Produce()` calls will block until buffered records are delivered and space frees up. Increase this value for high-throughput pipelines to avoid back-pressure stalls. **Type**: `int` **Default**: `10000` ### [](#max_in_flight_requests)`max_in_flight_requests` The maximum number of produce requests in flight per broker connection. While `idempotent_write` is enabled (the default) this must be `1`, as the client relies on a single in-flight request per broker to guarantee ordering. To use higher values you must set `idempotent_write` to `false`, which allows requests to be pipelined for throughput but may cause duplicate and out-of-order delivery on retries. Note that this is distinct from the output’s `max_in_flight` field, which counts message batches being written in parallel rather than produce requests on the wire. A value of `1` is not a throughput ceiling: records from concurrent writes are coalesced into fewer, larger produce requests. **Type**: `int` **Default**: `1` ### [](#max_message_bytes)`max_message_bytes` The maximum size of individual trace messages. Traces exceeding this size will be truncated or dropped. **Type**: `string` **Default**: `1MiB` ```yaml # Examples: max_message_bytes: 100MB # --- max_message_bytes: 50mib ``` ### [](#metadata_max_age)`metadata_max_age` The maximum age of cached cluster metadata before it is refreshed. Reducing this value can help detect cluster changes faster but increases metadata requests. **Type**: `string` **Default**: `1m` ### [](#partitioner)`partitioner` Override the default partitioner for trace messages. By default, traces are distributed across partitions for load balancing. **Type**: `string` | Option | Summary | | --- | --- | | least_backup | Chooses the least backed up partition (the partition with the fewest amount of buffered records). Partitions are selected per batch. | | manual | Manually select a partition for each message, requires the field partition to be specified. | | murmur2_hash | Kafka’s default hash algorithm that uses a 32-bit murmur2 hash of the key to compute which partition the record will be on. | | round_robin | Round-robin’s messages through all available partitions. This algorithm has lower throughput and causes higher CPU load on brokers, but can be useful if you want to ensure an even distribution of records to partitions. | ### [](#record_delivery_timeout)`record_delivery_timeout` The maximum time a record can sit in the producer buffer before it is failed, roughly equivalent to Kafka’s `delivery.timeout.ms`. This is evaluated before writing a request or after a produce response. When a record times out, all records in the same partition are also failed. Set to `0s` for no timeout (the default). With `idempotent_write` enabled, timeouts are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `string` **Default**: `0s` ### [](#record_retries)`record_retries` The maximum number of times a record produce is retried on failure before the record is failed. When a record fails, all records buffered in the same partition are also failed to preserve gapless ordering. Set to `0` for unlimited retries (the default). With `idempotent_write` enabled, retries are only enforced when safe to do so without creating invalid sequence numbers. **Type**: `int` **Default**: `0` ### [](#request_timeout_overhead)`request_timeout_overhead` Additional time to apply as overhead when calculating request deadlines. This buffer helps prevent premature timeouts. **Type**: `string` **Default**: `10s` ### [](#sampling)`sampling` Configure trace sampling to control the volume of trace data. Sampling is essential for high-throughput pipelines to prevent trace data from overwhelming your observability infrastructure. **Type**: `object` ### [](#sampling-enabled)`sampling.enabled` Whether to enable trace sampling. When disabled, all traces are exported. When enabled, traces are sampled according to the configured ratio. **Type**: `bool` **Default**: `false` ### [](#sampling-ratio)`sampling.ratio` The sampling ratio as a decimal between 0 and 1. For example, `0.1` samples 10% of traces, `0.01` samples 1%. Lower ratios reduce trace volume and overhead. For high-throughput production systems, start with 0.01-0.1 and adjust based on your needs. **Type**: `float` ```yaml # Examples: ratio: 0.05 # --- ratio: 0.85 # --- ratio: 0.5 ``` ### [](#sasl)`sasl[]` Specify one or more methods or mechanisms of SASL authentication, which are attempted in order. If the broker supports the first SASL mechanism, all connections use it. If the first mechanism fails, the client picks the first supported mechanism. If the broker does not support any client mechanisms, all connections fail. **Type**: `array` ```yaml # Examples: sasl: - mechanism: SCRAM-SHA-512 password: bar username: foo ``` ### [](#sasl-aws)`sasl[].aws` Contains AWS specific fields for when the `mechanism` is set to `AWS_MSK_IAM`. **Type**: `object` ### [](#sasl-aws-credentials)`sasl[].aws.credentials` Optional manual configuration of AWS credentials to use. More information can be found in [Amazon Web Services](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/). **Type**: `object` ### [](#sasl-aws-credentials-from_ec2_role)`sasl[].aws.credentials.from_ec2_role` Use the credentials of a host EC2 machine configured to assume [an IAM role associated with the instance](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_switch-role-ec2.html). **Type**: `bool` ### [](#sasl-aws-credentials-id)`sasl[].aws.credentials.id` The ID of credentials to use. **Type**: `string` ### [](#sasl-aws-credentials-profile)`sasl[].aws.credentials.profile` A profile from `~/.aws/credentials` to use. **Type**: `string` ### [](#sasl-aws-credentials-role)`sasl[].aws.credentials.role` A role ARN to assume. **Type**: `string` ### [](#sasl-aws-credentials-role_external_id)`sasl[].aws.credentials.role_external_id` An external ID to provide when assuming a role. **Type**: `string` ### [](#sasl-aws-credentials-secret)`sasl[].aws.credentials.secret` The secret for the credentials being used. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` ### [](#sasl-aws-credentials-token)`sasl[].aws.credentials.token` The token for the credentials being used, required when using short term credentials. **Type**: `string` ### [](#sasl-aws-endpoint)`sasl[].aws.endpoint` Allows you to specify a custom endpoint for the AWS API. **Type**: `string` ### [](#sasl-aws-region)`sasl[].aws.region` The AWS region to target. **Type**: `string` ### [](#sasl-aws-tcp)`sasl[].aws.tcp` TCP socket configuration. **Type**: `object` ### [](#sasl-aws-tcp-connect_timeout)`sasl[].aws.tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-aws-tcp-keep_alive)`sasl[].aws.tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#sasl-aws-tcp-keep_alive-count)`sasl[].aws.tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#sasl-aws-tcp-keep_alive-idle)`sasl[].aws.tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-keep_alive-interval)`sasl[].aws.tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#sasl-aws-tcp-tcp_user_timeout)`sasl[].aws.tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#sasl-extensions)`sasl[].extensions` Key/value pairs to add to OAUTHBEARER authentication requests. **Type**: `object` ### [](#sasl-mechanism)`sasl[].mechanism` The SASL mechanism to use. **Type**: `string` | Option | Summary | | --- | --- | | AWS_MSK_IAM | AWS IAM based authentication as specified by the 'aws-msk-iam-auth' java library. | | OAUTHBEARER | OAuth Bearer based authentication. | | PLAIN | Plain text authentication. | | REDPANDA_CLOUD_SERVICE_ACCOUNT | Redpanda Cloud Service Account authentication when running in Redpanda Cloud. | | SCRAM-SHA-256 | SCRAM based authentication as specified in RFC5802. | | SCRAM-SHA-512 | SCRAM based authentication as specified in RFC5802. | | none | Disable sasl authentication | ### [](#sasl-password)`sasl[].password` A password to provide for PLAIN or SCRAM-\* authentication. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#sasl-token)`sasl[].token` The token to use for a single session’s OAUTHBEARER authentication. **Type**: `string` **Default**: `""` ### [](#sasl-username)`sasl[].username` A username to provide for PLAIN or SCRAM-\* authentication. **Type**: `string` **Default**: `""` ### [](#schema_registry)`schema_registry` Schema registry information to publish schemas for tracing data along with the data. **Type**: `object` ### [](#schema_registry-basic_auth)`schema_registry.basic_auth` Allows you to specify basic authentication. **Type**: `object` ### [](#schema_registry-basic_auth-enabled)`schema_registry.basic_auth.enabled` Whether to use basic authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-basic_auth-password)`schema_registry.basic_auth.password` A password to authenticate with. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-basic_auth-username)`schema_registry.basic_auth.username` A username to authenticate as. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt)`schema_registry.jwt` (beta) Allows you to specify JWT authentication. **Type**: `object` ### [](#schema_registry-jwt-claims)`schema_registry.jwt.claims` A value used to identify the claims that issued the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-enabled)`schema_registry.jwt.enabled` Whether to use JWT authentication in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-jwt-headers)`schema_registry.jwt.headers` Add optional key/value headers to the JWT. **Type**: `object` **Default**: `{}` ### [](#schema_registry-jwt-private_key_file)`schema_registry.jwt.private_key_file` A file with the PEM encoded via PKCS1 or PKCS8 as private key. **Type**: `string` **Default**: `""` ### [](#schema_registry-jwt-signing_method)`schema_registry.jwt.signing_method` A method used to sign the token such as RS256, RS384, RS512 or EdDSA. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth)`schema_registry.oauth` Allows you to specify open authentication via OAuth version 1. **Type**: `object` ### [](#schema_registry-oauth-access_token)`schema_registry.oauth.access_token` A value used to gain access to the protected resources on behalf of the user. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-access_token_secret)`schema_registry.oauth.access_token_secret` A secret provided in order to establish ownership of a given access token. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_key)`schema_registry.oauth.consumer_key` A value used to identify the client to the service provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-consumer_secret)`schema_registry.oauth.consumer_secret` A secret used to establish ownership of the consumer key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth-enabled)`schema_registry.oauth.enabled` Whether to use OAuth version 1 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-oauth2)`schema_registry.oauth2` Allows you to specify open authentication via OAuth version 2 using the client credentials token flow. **Type**: `object` ### [](#schema_registry-oauth2-client_key)`schema_registry.oauth2.client_key` A value used to identify the client to the token provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth2-client_secret)`schema_registry.oauth2.client_secret` A secret used to establish ownership of the client key. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-oauth2-enabled)`schema_registry.oauth2.enabled` Whether to use OAuth version 2 in requests. **Type**: `bool` **Default**: `false` ### [](#schema_registry-oauth2-endpoint_params)`schema_registry.oauth2.endpoint_params` A list of optional endpoint parameters, values should be arrays of strings. **Type**: `object` **Default**: `{}` ```yaml # Examples: endpoint_params: audience: - https://example.com resource: - https://api.example.com ``` ### [](#schema_registry-oauth2-scopes)`schema_registry.oauth2.scopes[]` A list of optional requested permissions. **Type**: `array` **Default**: `[]` ### [](#schema_registry-oauth2-token_url)`schema_registry.oauth2.token_url` The URL of the token provider. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls)`schema_registry.tls` Custom TLS settings can be used to override system defaults. **Type**: `object` ### [](#schema_registry-tls-client_certs)`schema_registry.tls.client_certs[]` A list of client certificates to use. For each certificate either the fields `cert` and `key`, or `cert_file` and `key_file` should be specified, but not both. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#schema_registry-tls-client_certs-cert)`schema_registry.tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-cert_file)`schema_registry.tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key)`schema_registry.tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-key_file)`schema_registry.tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#schema_registry-tls-client_certs-password)`schema_registry.tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#schema_registry-tls-enable_renegotiation)`schema_registry.tls.enable_renegotiation` Whether to allow the remote server to repeatedly request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#schema_registry-tls-root_cas)`schema_registry.tls.root_cas` An optional root certificate authority to use. This is a string, representing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#schema_registry-tls-root_cas_file)`schema_registry.tls.root_cas_file` An optional path of a root certificate authority file to use. This is a file, often with a .pem extension, containing a certificate chain from the parent trusted root certificate, to possible intermediate signing certificates, to the host certificate. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#schema_registry-tls-skip_cert_verify)`schema_registry.tls.skip_cert_verify` Whether to skip server side certificate verification. **Type**: `bool` **Default**: `false` ### [](#schema_registry-url)`schema_registry.url` The base URL of the schema registry service. **Type**: `string` ### [](#seed_brokers)`seed_brokers[]` A list of broker addresses to connect to in order. Use commas to separate multiple addresses in a single list item. **Type**: `array` ```yaml # Examples: seed_brokers: - "localhost:9092" # --- seed_brokers: - "foo:9092" - "bar:9092" # --- seed_brokers: - "foo:9092,bar:9092" ``` ### [](#service)`service` The service name to identify this Redpanda Connect instance in traces. This appears in trace visualizations and helps correlate traces across distributed systems. Use descriptive names like `order-processor` or `analytics-pipeline`. **Type**: `string` **Default**: `redpanda-connect` ### [](#tags)`tags` Custom key-value tags to attach to all traces from this instance. Use tags to add metadata like environment (`production`, `staging`), region, version, or instance identifiers. Tags appear as resource attributes in OpenTelemetry traces. **Type**: `object` **Default**: `{}` ### [](#tcp)`tcp` Configure TCP socket-level settings to optimize network performance and reliability. These low-level controls are useful for: - **High-latency networks**: Increase `connect_timeout` to allow more time for connection establishment - **Long-lived connections**: Configure `keep_alive` settings to detect and recover from stale connections - **Unstable networks**: Tune keep-alive probes to balance between quick failure detection and avoiding false positives - **Linux systems with specific requirements**: Use `tcp_user_timeout` (Linux 2.6.37+) to control data acknowledgment timeouts Most users should keep the default values. Only modify these settings if you’re experiencing connection stability issues or have specific network requirements. **Type**: `object` ### [](#tcp-connect_timeout)`tcp.connect_timeout` Maximum amount of time a dial will wait for a connect to complete. Zero disables. **Type**: `string` **Default**: `0s` ### [](#tcp-keep_alive)`tcp.keep_alive` TCP keep-alive probe configuration. **Type**: `object` ### [](#tcp-keep_alive-count)`tcp.keep_alive.count` Maximum unanswered keep-alive probes before dropping the connection. Zero defaults to 9. **Type**: `int` **Default**: `9` ### [](#tcp-keep_alive-idle)`tcp.keep_alive.idle` Duration the connection must be idle before sending the first keep-alive probe. Zero defaults to 15s. Negative values disable keep-alive probes. **Type**: `string` **Default**: `15s` ### [](#tcp-keep_alive-interval)`tcp.keep_alive.interval` Duration between keep-alive probes. Zero defaults to 15s. **Type**: `string` **Default**: `15s` ### [](#tcp-tcp_user_timeout)`tcp.tcp_user_timeout` Maximum time to wait for acknowledgment of transmitted data before killing the connection. Linux-only (kernel 2.6.37+), ignored on other platforms. When enabled, keep\_alive.idle must be greater than this value per RFC 5482. Zero disables. **Type**: `string` **Default**: `0s` ### [](#timeout)`timeout` The maximum time to wait for trace messages to be acknowledged by the broker before considering the write failed. **Type**: `string` **Default**: `10s` ### [](#tls)`tls` Configure Transport Layer Security (TLS) settings to secure network connections. This includes options for standard TLS as well as mutual TLS (mTLS) authentication where both client and server authenticate each other using certificates. Key configuration options include `enabled` to enable TLS, `client_certs` for mTLS authentication, `root_cas`/`root_cas_file` for custom certificate authorities, and `skip_cert_verify` for development environments. **Type**: `object` ### [](#tls-client_certs)`tls.client_certs[]` A list of client certificates for mutual TLS (mTLS) authentication. Configure this field to enable mTLS, authenticating the client to the server with these certificates. You must set `tls.enabled: true` for the client certificates to take effect. **Certificate pairing rules**: For each certificate item, provide either: - Inline PEM data using both `cert` **and** `key` or - File paths using both `cert_file` **and** `key_file`. Mixing inline and file-based values within the same item is not supported. **Type**: `array` **Default**: `[]` ```yaml # Examples: client_certs: - cert: foo key: bar # --- client_certs: - cert_file: ./example.pem key_file: ./example.key ``` ### [](#tls-client_certs-cert)`tls.client_certs[].cert` A plain text certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-cert_file)`tls.client_certs[].cert_file` The path of a certificate to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key)`tls.client_certs[].key` A plain text certificate key to use. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-key_file)`tls.client_certs[].key_file` The path of a certificate key to use. **Type**: `string` **Default**: `""` ### [](#tls-client_certs-password)`tls.client_certs[].password` A plain text password for when the private key is password encrypted in PKCS#1 or PKCS#8 format. The obsolete `pbeWithMD5AndDES-CBC` algorithm is not supported for the PKCS#8 format. Because the obsolete pbeWithMD5AndDES-CBC algorithm does not authenticate the ciphertext, it is vulnerable to padding oracle attacks that can let an attacker recover the plaintext. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: password: foo # --- password: ${KEY_PASSWORD} ``` ### [](#tls-enable_renegotiation)`tls.enable_renegotiation` Whether to allow the remote server to request renegotiation. Enable this option if you’re seeing the error message `local error: tls: no renegotiation`. **Type**: `bool` **Default**: `false` ### [](#tls-enabled)`tls.enabled` Whether to use TLS for the connection to the Redpanda cluster. **Type**: `bool` **Default**: `false` ### [](#tls-root_cas)`tls.root_cas` Specify a root certificate authority to use (optional). This is a string that represents a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for inline certificate data or `root_cas_file` for file-based certificate loading. > ⚠️ **CAUTION** > > This field contains sensitive information that usually shouldn’t be added to a configuration directly. For more information, see [Manage Secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) before adding it to your configuration. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- ``` ### [](#tls-root_cas_file)`tls.root_cas_file` Specify the path to a root certificate authority file (optional). This is a file, often with a `.pem` extension, which contains a certificate chain from the parent-trusted root certificate, through possible intermediate signing certificates, to the host certificate. Use either this field for file-based certificate loading or `root_cas` for inline certificate data. **Type**: `string` **Default**: `""` ```yaml # Examples: root_cas_file: ./root_cas.pem ``` ### [](#tls-skip_cert_verify)`tls.skip_cert_verify` Whether to skip server-side certificate verification. Set to `true` only for testing environments as this reduces security by disabling certificate validation. When using self-signed certificates or in development, this may be necessary, but should never be used in production. Consider using `root_cas` or `root_cas_file` to specify trusted certificates instead of disabling verification entirely. **Type**: `bool` **Default**: `false` ### [](#topic)`topic` The Redpanda topic where trace data is written. This topic should be dedicated to traces and configured with appropriate retention policies. Default: `otel-traces` **Type**: `string` **Default**: `otel-traces` --- # Page 275: Configuration **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/about.md --- # Configuration > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configuration latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/about.adoc description: Learn about different options for configuring Redpanda Connect. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Redpanda Connect pipelines are configured in a YAML file that consists of a number of root sections, arranged like so: #### Common ```yaml input: kafka: addresses: [ TODO ] topics: [ foo, bar ] consumer_group: foogroup pipeline: processors: - mapping: | root.message = this root.meta.link_count = this.links.length() output: aws_s3: bucket: TODO path: '${! meta("kafka_topic") }/${! json("message.id") }.json' ``` #### Full ```yaml http: address: 0.0.0.0:4195 debug_endpoints: false input: kafka: addresses: [ TODO ] topics: [ foo, bar ] consumer_group: foogroup buffer: none: {} pipeline: processors: - mapping: | root.message = this root.meta.link_count = this.links.length() output: aws_s3: bucket: TODO path: '${! meta("kafka_topic") }/${! json("message.id") }.json' input_resources: [] cache_resources: [] processor_resources: [] rate_limit_resources: [] output_resources: [] logger: level: INFO static_fields: '@service': benthos metrics: prometheus: {} tracer: none: {} shutdown_timeout: 20s shutdown_delay: "" ``` Most sections represent a component type, which you can read about in more detail in [this document](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/about/). These types are hierarchical. For example, an `input` can have a list of child `processor` types attached to it, which in turn can have their own `processor` children. This is powerful but can potentially lead to large and cumbersome configuration files. This document outlines tooling provided by Redpanda Connect to help with writing and managing these more complex configuration files. ## [](#testing)Testing For guidance on how to write and run unit tests for your configuration files read [this guide](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/unit_testing/). ## [](#customizing-your-configuration)Customizing your configuration Sometimes it’s useful to write a configuration where certain fields can be defined during deployment. For this purpose Redpanda Connect supports [environment variable interpolation](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/), allowing you to set fields in your config with environment variables like so: ```yaml input: kafka: addresses: - ${KAFKA_BROKER:localhost:9092} topics: - ${KAFKA_TOPIC:default-topic} ``` This is very useful for sharing configuration files across different deployment environments. ## [](#labels)Labels Labels are unique, user-defined identifiers used throughout Redpanda Connect configurations. They serve two purposes: - **Reference:** Allow different parts of your pipeline to refer to specific components or resources. - **Readability:** Make your configuration more understandable for humans, especially in complex deployments. You can assign labels to most pipeline components, including resources, inputs, outputs, processors, and entire pipelines. Using clear, descriptive labels improves both maintainability and clarity. Labels are commonly applied to the following components: ### [](#resources)Resources Labels identify [reusable resources](#reuse) such as processors, caches, and rate limiters, making them easy to reference elsewhere in your pipeline. ```yaml processor_resources: - label: my-transformer # Processor resource label mapping: 'root = content().uppercase()' cache_resources: - label: user-cache # Cache resource label memory: default_ttl: 300s rate_limit_resources: - label: api-limiter # Rate limiter resource label local: count: 100 interval: 1m ``` ### [](#component-labeling-for-clarity)Component labeling for clarity You can also use labels on inputs, outputs, processors, and other components to improve the human-readability of your configuration and make troubleshooting easier. For example: ```yaml input: label: ingest_api http_server: {} pipeline: label: user_data_ingest processors: - label: sanitize_fields mapping: 'root = this.trim()' - resource: my-transformer ``` ## [](#label-naming-requirements)Label naming requirements Labels must meet the following criteria: - **Length**: 3-128 characters - **Allowed characters**: Alphanumeric, hyphens, and underscores (`A-Za-z0-9-_`) - **Case sensitivity**: Labels are case-sensitive Example valid labels my-processor data\_transformer\_01 UserAnalytics-v2 Example invalid labels ab // Too short (less than 3 characters) my.processor // Invalid character: period my processor // Invalid character: space ## [](#reuse)Reusing configuration snippets Sometimes it’s necessary to use a rather large component multiple times. Instead of copy/pasting the configuration or using YAML anchors you can define your component as a resource. In the following example we want to make an HTTP request with our payloads. Occasionally the payload might get rejected due to garbage within its contents, and so we catch these rejected requests, attempt to "cleanse" the contents and try to make the same HTTP request again. Since the HTTP request component is quite large (and likely to change over time) we make sure to avoid duplicating it by defining it as a resource `get_foo`: ```yaml pipeline: processors: - resource: get_foo - catch: - mapping: | root = this root.content = this.content.strip_html() - resource: get_foo processor_resources: - label: get_foo http: url: http://example.com/foo verb: POST headers: SomeThing: "set-to-this" SomeThingElse: "set-to-something-else" ``` ## [](#shutting-down)Shutting down Under normal operating conditions, the Redpanda Connect process will shut down when there are no more messages produced by inputs and the final message has been processed. The shutdown procedure can also be initiated by sending the process a interrupt (`SIGINT`) or termination (`SIGTERM`) signal. There are two top-level configuration options that control the shutdown behavior: `shutdown_timeout` and `shutdown_delay`. ### [](#shutdown-delay)Shutdown delay The `shutdown_delay` option can be used to delay the start of the shutdown procedure. This is useful for pipelines that need a short grace period to have their metrics and traces scraped. While the shutdown delay is in effect, the HTTP metrics endpoint continues to be available for scraping and any active tracers are free to flush remaining traces. The shutdown delay can be interrupted by sending the Redpanda Connect process a second OS interrupt or termination signal. ### [](#shutdown-timeout)Shutdown timeout The `shutdown_timeout` option sets a hard deadline for Redpanda Connect process to gracefully terminate. If this duration is exceeded then the process is forcefully terminated and any messages that were in-flight will be dropped. This option takes effect after the `shutdown_delay` duration has passed if that is enabled. --- # Page 276: Message Batching **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching.md --- # Message Batching > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Message Batching latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/batching page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/batching.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/batching.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Redpanda Connect is able to join sources and sinks with sometimes conflicting batching behaviors without sacrificing its strong delivery guarantees. It’s also able to perform powerful [processing functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/windowed_processing/) across batches of messages such as grouping, archiving and reduction. Therefore, batching within Redpanda Connect is a mechanism that serves multiple purposes: 1. [Performance (throughput)](#performance) 2. [Grouped message processing](#grouped-message-processing) 3. [Compatibility (mixing multi and single part message protocols)](#compatibility) ## [](#performance)Performance For most users the only benefit of batching messages is improving throughput over your output protocol. For some protocols this can happen in the background and requires no configuration from you. However, if an output has a `batching` configuration block this means it benefits from batching and requires you to specify how you’d like your batches to be formed by configuring a [batching policy](#batch-policy): ```yaml output: kafka: addresses: [ todo:9092 ] topic: benthos_stream # Either send batches when they reach 10 messages or when 100ms has passed # since the last batch. batching: count: 10 period: 100ms ``` However, a small number of inputs such as [`kafka`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/kafka/) must be consumed sequentially (in this case by partition) and therefore benefit from specifying your batch policy at the input level instead: ```yaml input: kafka: addresses: [ todo:9092 ] topics: [ benthos_input_stream ] batching: count: 10 period: 100ms output: kafka: addresses: [ todo:9092 ] topic: benthos_stream ``` Inputs that behave this way are documented as such and have a `batching` configuration block. Sometimes you may prefer to create your batches before processing in order to benefit from [batch wide processing](#grouped-message-processing), in which case if your input doesn’t already support [a batch policy](#batch-policy) you can instead use a [`broker`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/broker/), which also allows you to combine inputs with a single batch policy: ```yaml input: broker: inputs: - resource: foo - resource: bar batching: count: 50 period: 500ms ``` This also works the same with [output brokers](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/broker/). ## [](#grouped-message-processing)Grouped message processing And some processors such as [`while`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/while/) are executed once across a whole batch, you can avoid this behavior with the [`for_each` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/for_each/): ```yaml pipeline: processors: - for_each: - while: at_least_once: true max_loops: 0 check: errored() processors: - catch: [] # Wipe any previous error - resource: foo # Attempt this processor until success ``` There’s a vast number of processors that specialise in operations across batches such as [grouping](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/group_by/) and [archiving](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive/). For example, the following processors group a batch of messages according to a metadata field and compresses them into separate `.tar.gz` archives: ```yaml pipeline: processors: - group_by_value: value: ${! meta("kafka_partition") } - archive: format: tar - compress: algorithm: gzip output: aws_s3: bucket: TODO path: docs/${! meta("kafka_partition") }/${! count("files") }-${! timestamp_unix_nano() }.tar.gz ``` For more examples of batched (or windowed) processing check out [this document](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/windowed_processing/). ## [](#compatibility)Compatibility Redpanda Connect is able to read and write over protocols that support multiple part messages, and all payloads travelling through Redpanda Connect are represented as a multiple part message. Therefore, all components within Redpanda Connect are able to work with multiple parts in a message as standard. When messages reach an output that _doesn’t_ support multiple parts the message is broken down into an individual message per part, and then one of two behaviors happen depending on the output. If the output supports batch sending messages then the collection of messages are sent as a single batch. Otherwise, Redpanda Connect falls back to sending the messages sequentially in multiple, individual requests. This behavior means that not only can multiple part message protocols be easily matched with single part protocols, but also the concept of multiple part messages and message batches are interchangeable within Redpanda Connect. ### [](#shrinking-batches)Shrinking batches A message batch (or multiple part message) can be broken down into smaller batches using the [`split`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/split/) processor: ```yaml input: # Consume messages that arrive in three parts. resource: foo processors: # Drop the third part - select_parts: parts: [ 0, 1 ] # Then break our message parts into individual messages - split: size: 1 ``` This is also useful when your input source creates batches that are too large for your output protocol: ```yaml input: aws_s3: bucket: todo pipeline: processors: - decompress: algorithm: gzip - unarchive: format: tar # Limit batch sizes to 5MB - split: byte_size: 5_000_000 ``` ## [](#batch-policy)Batch policy When an input or output component has a config field `batching` that means it supports a batch policy. This is a mechanism that allows you to configure exactly how your batching should work on messages before they are routed to the input or output it’s associated with. Batches are considered complete and will be flushed downstream when either of the following conditions are met: - The `byte_size` field is non-zero and the total size of the batch in bytes matches or exceeds it (disregarding metadata.) - The `count` field is non-zero and the total number of messages in the batch matches or exceeds it. - A message added to the batch causes the [`check`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) to return to `true`. - The `period` field is non-empty and the time since the last batch exceeds its value. This allows you to combine conditions: ```yaml output: kafka: addresses: [ todo:9092 ] topic: benthos_stream # Either send batches when they reach 10 messages or when 100ms has passed # since the last batch. batching: count: 10 period: 100ms ``` > ⚠️ **CAUTION** > > A batch policy has the capability to _create_ batches, but not to break them down. If your configured pipeline is processing messages that are batched _before_ they reach the batch policy then they may circumvent the conditions you’ve specified here, resulting in sizes you aren’t expecting. If you are affected by this limitation then consider breaking the batches down with a [`split` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/split/) before they reach the batch policy. ### [](#post-batch-processing)Post-batch processing A batch policy also has a field `processors` which allows you to define an optional list of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) to apply to each batch before it is flushed. This is a good place to aggregate or archive the batch into a compatible format for an output: ```yaml output: http_client: url: http://localhost:4195/post batching: count: 10 processors: - archive: format: lines ``` The above config will batch up messages and then merge them into a line delimited format before sending it over HTTP. This is an easier format to parse than the default which would have been [rfc1342](https://www.w3.org/Protocols/rfc1341/7_2_Multipart.html). During shutdown any remaining messages waiting for a batch to complete will be flushed down the pipeline. --- # Page 277: Contextual Variables **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/contextual-variables.md --- # Contextual Variables > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Contextual Variables latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/contextual-variables page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/contextual-variables.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/contextual-variables.adoc description: Learn about the advantages of using contextual variables, and how to add them to your data pipelines. page-git-created-date: "2025-01-09" page-git-modified-date: "2025-08-08" --- Learn about the advantages of using contextual variables, and how to add them to your data pipelines. ## [](#understanding-contextual-variables)Understanding contextual variables Contextual variables provide an easy way to access information about the environment in which a data pipeline is running and the pipeline itself. You can add any of the following contextual variables to your pipeline configurations: | Contextual variable name | Description | | --- | --- | | ${REDPANDA_BROKERS} | The bootstrap server address of the cluster on which the data pipeline is running. | | ${REDPANDA_ID} | The ID of the cluster on which the data pipeline is running. | | ${REDPANDA_REGION} | The cloud region where the data pipeline is deployed. | | ${REDPANDA_PIPELINE_ID} | The ID of the data pipeline that is currently running. | | ${REDPANDA_PIPELINE_NAME} | The display name of the data pipeline that is currently running. | | ${REDPANDA_SCHEMA_REGISTRY_URL} | The URL of the Schema Registry associated with the cluster on which the data pipeline is running. | Contextual variables are automatically set at runtime, which means that you can reuse them across multiple pipelines and development environments. For example, if you add the contextual variable `${REDPANDA_ID}` to a pipeline configuration, it’s always set to the ID of the cluster on which the data pipeline is running, whether the pipeline is in your development, user acceptance testing, or production environment. This increases the portability of pipeline configurations and reduces maintenance overheads. You can also use contextual variables to improve data traceability. See the [Example pipeline configuration](#example-pipeline-configuration) for full details. ## [](#add-contextual-variable-to-a-data-pipeline)Add contextual variable to a data pipeline Add a contextual variable to any pipeline configuration using the notation `${CONTEXTUAL_VARIABLE_NAME}`, for example: ```yaml output: kafka_franz: seed_brokers: - ${REDPANDA_BROKERS} ``` ### [](#example-pipeline-configuration)Example pipeline configuration For improved data traceability, the following pipeline configuration adds the data pipeline display name (`${REDPANDA_PIPELINE_NAME}`) and ID (`${REDPANDA_PIPELINE_ID}`) to all messages that are processed. The configuration also uses the `$REDPANDA_BROKERS` contextual variable to automatically populate the bootstrap server address of the cluster on which the pipeline is run, which allows Redpanda Connect to write updated messages to the `data` topic defined in the `kafka_franz` output. ```yaml input: generate: mapping: | root.data = "test message" interval: 10s pipeline: processors: - bloblang: | root = this root.source = "${REDPANDA_PIPELINE_NAME}" root.source_id = "${REDPANDA_PIPELINE_ID}" output: kafka_franz: seed_brokers: - ${REDPANDA_BROKERS} topic: data tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: cluster-username password: cluster-password ``` ## [](#suggested-reading)Suggested reading - Learn how to [add secrets to your pipeline](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/). - Try one of our [Redpanda Connect cookbooks](https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/). - Choose [connectors for your use case](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/about/). --- # Page 278: Error Handling **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling.md --- # Error Handling > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Error Handling latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/error_handling page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/error_handling.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/error_handling.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Redpanda Connect supports a range of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/), such as `http` and `aws_lambda`, that may fail when retry attempts are exhausted. By default, when a processor fails, the message data continues through the pipeline mostly unchanged, except for the addition of a metadata flag, which you can use for handling errors. To make processing errors terminal instead, see [Strict error handling](#strict-error-handling). This topic explains some common error-handling patterns, including dropping messages, recovering them with more processing, and routing them to a dead-letter queue. It also shows how to combine these approaches, where appropriate. ## [](#strict-error-handling)Strict error handling To make processing errors terminal instead of relying on error flags, enable strict error handling with the top-level `error_handling` configuration block: ```yaml error_handling: strict: true ``` When `error_handling.strict` is enabled: - A processing error is terminal for the affected message. The message skips the remaining processors in its pipeline. - The failed message is rejected (nacked) at the output rather than written. - A standalone `catch` processor does not recover failed messages, because a failed message short-circuits past it. To recover from an expected error under strict mode, wrap the fallible step and its recovery logic in a [`try_catch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/try_catch/). To retry transient failures until they succeed, use a [`retry` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/retry/). > 📝 **NOTE** > > Strict error handling will become the default and only behavior in the next major version of Redpanda Connect. The rest of this topic describes error-handling patterns for the default behavior. ## [](#abandon-on-failure)Abandon on failure You can use the [`try` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/try/) to define a list of processors that are executed in sequence. If a processor fails for a particular message, that message skips the remaining processors. For example: - If `processor_1` fails to process a message, that message skips `processor_2` and `processor_3`. - If a message is processed by `processor_1`, but `processor_2` fails, that message skips `processor_3`, and so on. ```yaml pipeline: processors: - try: - resource: processor_1 - resource: processor_2 # Skip if processor_1 fails - resource: processor_3 # Skip if processor_1 or processor_2 fails ``` ## [](#recover-failed-messages)Recover failed messages You can also route failed messages through defined processing steps using a [`catch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/catch/). For example, if `processor_1` fails to process a message, it is rerouted to `processor_2`. ```yaml pipeline: processors: - resource: processor_1 # Processor that might fail - catch: - resource: processor_2 # Processes rerouted messages ``` After messages complete all processing steps defined in the `catch` block, failure flags are removed and they are treated like regular messages. To keep failure flags in messages, you can simulate a `catch` block using a [`switch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/switch/): ```yaml pipeline: processors: - resource: processor_1 # Processor that might fail - switch: - check: errored() processors: - resource: processor_2 # Processes rerouted messages ``` ## [](#logging-errors)Logging errors When an error occurs, there may be useful information stored in the error flag. You can use [`error`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#error) Bloblang function interpolations to write this information to logs. You can also add the following Bloblang functions to expose additional details about the processor that triggered the error. - [`error_source_label`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#error_source_label) - [`error_source_name`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#error_source_name) - [`error_source_path`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#error_source_path) For example, this configuration catches processor failures and writes the following information to logs: - The label of the processor (`${!error_source_label()}`) that failed - The cause of the failure (`${!error()}`) ```yaml pipeline: processors: - try: - resource: processor_1 # Processor that might fail - resource: processor_2 # Processor that might fail - resource: processor_3 # Processor that might fail - catch: - log: message: "Processor ${!error_source_label()} failed due to: ${!error()}" ``` You could also add an error message to the message payload: ```yaml pipeline: processors: - resource: processor_1 # Processor that might fail - resource: processor_2 # Processor that might fail - resource: processor_3 # Processor that might fail - catch: - mapping: | root = this root.meta.error = error() ``` ## [](#attempt-until-success)Attempt until success To process a particular message until it is successful, try using a [`retry`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/retry/) processor: ```yaml pipeline: processors: - retry: backoff: initial_interval: 1s max_interval: 5s max_elapsed_time: 30s processors: # Retries this processor until the message is processed, or the maximum elapsed time is reached. - resource: processor_1 ``` ## [](#drop-failed-messages)Drop failed messages To filter out any failed messages from your pipeline, you can use a [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/): ```yaml pipeline: processors: - mapping: root = if errored() { deleted() } ``` The mapping uses the error flag to identify any failed messages in a batch and drops the messages, which propagates acknowledgements (also known as "acks") upstream to the pipeline’s input. ## [](#reject-messages)Reject messages Some inputs, such as `nats`, `gcp_pubsub`, and `amqp_1`, support nacking (rejecting) messages. Rather than delivering unprocessed messages to your output, you can use the [`reject_errored` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/reject_errored/) to perform a nack (or rejection) on them: ```yaml output: reject_errored: resource: processor_1 # Only non-errored messages go here ``` ## [](#route-to-a-dead-letter-queue)Route to a dead-letter queue You can also route failed messages to a different output by nesting the [`reject_errored` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/reject_errored/) within a [`fallback` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/fallback/) ```yaml output: fallback: - reject_errored: resource: processor_1 # Only non-errored messages go here - resource: processor_2 # Only errored messages, or delivery failures to processor_1, go here ``` If you want to route data differently based on the type of error message, you can use a [`switch` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/switch/): ```yaml output: switch: cases: # Capture specifically cat-related errors - check: errored() && error().contains("meow") output: resource: processor_1 # Capture all other errors - check: errored() output: resource: processor_2 # Finally, route all successfully processed messages here - output: resource: processor_3 ``` Finally, you can attach additional metadata when routing messages to the dead-letter queue, such as the error message. This can be done by running a series of [processors](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/about/) before sending the data to the final [output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about/). ```yaml output: fallback: - reject_errored: resource: processor_1 # Only non-errored messages go here - processors: - mutation: | root.error = @fallback_error # Adds the error message before sending the message to the dead-letter queue output resource: processor_2 # Only errored messages, or delivery failures to processor_1, go here ``` --- # Page 279: Field Paths **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths.md --- # Field Paths > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Field Paths latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/field_paths page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/field_paths.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/field_paths.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- Many components within Redpanda Connect allow you to target certain fields using a JSON dot path. The syntax of a path within Redpanda Connect is similar to [JSON Pointers](https://tools.ietf.org/html/rfc6901), except with dot separators instead of slashes (and no leading dot.) When a path is used to set a value any path segment that does not yet exist in the structure is created as an object. For example, if we had the following JSON structure: ```json { "foo": { "bar": 21 } } ``` The query path `foo.bar` would return `21`. The characters `~` (%x7E) and `.` (%x2E) have special meaning in Redpanda Connect paths. Therefore `~` needs to be encoded as `~0` and `.` needs to be encoded as `~1` when these characters appear within a key. For example, if we had the following JSON structure: ```json { "foo.foo": { "bar~bo": { "": { "baz": 22 } } } } ``` The query path `foo~1foo.bar~0bo..baz` would return `22`. ## [](#arrays)Arrays When Redpanda Connect encounters an array while traversing a JSON structure it requires the next path segment to be either an integer of an existing index or, depending on whether the path is used to query or set the target value, the character `*` or `-` respectively. For example, if we had the following JSON structure: ```json { "foo": [ 0, 1, { "bar": 23 } ] } ``` The query path `foo.2.bar` would return `23`. ### [](#querying)Querying When a query reaches an array the character `*` indicates that the query should return the value of the remaining path from each array element (within an array.) ### [](#setting)Setting When an array is reached the character `-` indicates that a new element should be appended to the end of the existing elements, if this character is not the final segment of the path then an object is created. --- # Page 280: Interpolation **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation.md --- # Interpolation > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Interpolation latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/interpolation page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/interpolation.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/interpolation.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- > 📝 **NOTE** > > Environment variables are not currently supported in Redpanda Connect in Redpanda Cloud, but you can use [contextual variables](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/contextual-variables/) to access information about the environment in which a data pipeline is running, and the pipeline itself. Redpanda Connect allows you to dynamically set config fields with environment variables anywhere within a config file using the syntax `${}` (or `${:}` in order to specify a default value). This is useful for setting environment specific fields such as addresses: ```yaml input: kafka: addresses: [ "${BROKERS}" ] consumer_group: redpanda_connect_consumer topics: [ "haha_business" ] ``` ```sh BROKERS="foo:9092,bar:9092" rpk connect run ./config.yaml ``` If a literal string is required that matches this pattern (`${foo}`) you can escape it with double brackets. For example, the string `${{foo}}` is read as the literal `${foo}`. ## [](#undefined-variables)Undefined variables When an environment variable interpolation is found within a config, does not have a default value specified, and the environment variable is not defined a linting error will be reported. In order to avoid this it is possible to specify environment variable interpolations with an explicit empty default value by adding the colon without a following value, i.e. `${FOO:}` would be equivalent to `${FOO}` and would not trigger a linting error should `FOO` not be defined. ## [](#yaml-tags)YAML tags By default, Redpanda Connect interpolates environment variables as strings. You can use [YAML tags](https://yaml.org/spec/1.2.2/#24-tags) to interpret values as another scalar type, such as integers. ```yaml output: redpanda: # ... batching: count: !!int ${BATCHING_COUNT:500} period: "${BATCHING_PERIOD:1s}" ``` Redpanda Connect supports the [core schema tags](https://yaml.org/spec/1.2.2/#103-core-schema) for scalar types: - `null` - `bool` - `int` - `float` - `str` (default) ## [](#bloblang-queries)Bloblang queries Some Redpanda Connect fields also support [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) function interpolations, which are much more powerful expressions that allow you to query the contents of messages and perform arithmetic. The syntax of a function interpolation is `${!}`, where the contents are a bloblang query (the right-hand-side of a bloblang map) including a range of [functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#functions). For example, with the following config: ```yaml output: kafka: addresses: [ "TODO:6379" ] topic: 'dope-${! json("topic") }' ``` A message with the contents `{"topic":"foo","message":"hello world"}` would be routed to the Kafka topic `dope-foo`. If a literal string is required that matches this pattern (`${!foo}`) then, similar to environment variables, you can escape it with double brackets. For example, the string `${{!foo}}` would be read as the literal `${!foo}`. Bloblang supports arithmetic, boolean operators, coalesce and mapping expressions. For more in-depth details about the language [check out the docs](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/). ## [](#examples)Examples ### [](#reference-metadata)Reference metadata A common usecase for interpolated functions is dynamic routing at the output level using metadata: ```yaml output: kafka: addresses: [ TODO ] topic: ${! meta("output_topic") } key: ${! meta("key") } ``` ### [](#coalesce-and-mapping)Coalesce and mapping Bloblang supports coalesce and mapping, which makes it easy to extract values from slightly varying data structures: ```yaml pipeline: processors: - cache: resource: foocache operator: set key: '${! json().message.(foo | bar).id }' value: '${! content() }' ``` Here’s a map of inputs to resulting values: {"foo":{"a":{"baz":"from\_a"},"c":{"baz":"from\_c"}}} -> from\_a {"foo":{"b":{"baz":"from\_b"},"c":{"baz":"from\_c"}}} -> from\_b {"foo":{"b":null,"c":{"baz":"from\_c"}}} -> from\_c ### [](#delayed-processing)Delayed processing We have a stream of JSON documents each with a unix timestamp field `doc.received_at` which is set when our platform receives it. We wish to only process messages an hour _after_ they were received. We can achieve this by running the `sleep` processor using an interpolation function to calculate the seconds needed to wait for: ```yaml pipeline: processors: - sleep: duration: '${! 3600 - ( timestamp_unix() - json("doc.created_at").number() ) }s' ``` If the calculated result is less than or equal to zero the processor does not sleep at all. If the value of `doc.created_at` is a string then our method `.number()` will attempt to parse it into a number. --- # Page 281: Metadata **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/metadata.md --- # Metadata > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Metadata latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/metadata page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/metadata.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/metadata.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- In Redpanda Connect each message has raw contents and metadata, which is a map of key/value pairs representing an arbitrary amount of complementary data. When an input protocol supports attributes or metadata they will automatically be added to your messages, refer to the respective input documentation for a list of metadata keys. When an output supports attributes or metadata any metadata key/value pairs in a message will be sent (subject to service limits). ## [](#editing-metadata)Editing metadata Redpanda Connect allows you to add and remove metadata using the [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/). For example, you can do something like this in your pipeline: ```yaml pipeline: processors: - mapping: | # Remove all existing metadata from messages meta = deleted() # Add a new metadata field `time` from the contents of a JSON # field `event.timestamp` meta time = event.timestamp ``` You can also use [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) to delete individual metadata keys with: ```bloblang meta foo = deleted() ``` Or do more interesting things like remove all metadata keys with a certain prefix: ```bloblang meta = @.filter(kv -> !kv.key.has_prefix("kafka_")) ``` ## [](#using-metadata)Using metadata Metadata values can be referenced in any field that supports [interpolation functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/). For example, you can route messages to Kafka topics using interpolation of metadata keys: ```yaml output: kafka: addresses: [ TODO ] topic: ${! meta("target_topic") } ``` Redpanda Connect also allows you to conditionally process messages based on their metadata with the [`switch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/switch/): ```yaml pipeline: processors: - switch: - check: '@doc_type == "nested"' processors: - sql_insert: driver: mysql dsn: foouser:foopassword@tcp(localhost:3306)/foodb table: footable columns: [ foo, bar, baz ] args_mapping: | root = [ this.document.foo, this.document.bar, @kafka_topic, ] # In: {"document":{"foo":"value1","bar":"value2"}} ``` ## [](#restricting-metadata)Restricting metadata Outputs that support metadata, headers or some other variant of enriched fields on messages will attempt to send all metadata key/value pairs by default. However, sometimes it’s useful to refer to metadata fields at the output level even though we do not wish to send them with our data. In this case it’s possible to restrict the metadata keys that are sent with the field `metadata.exclude_prefixes` within the respective output config. For example, if we were sending messages to kafka using a metadata key `target_topic` to determine the topic but we wished to prevent that metadata key from being sent as a header we could use the following configuration: ```yaml output: kafka: addresses: [ TODO ] topic: ${! meta("target_topic") } metadata: exclude_prefixes: - target_topic ``` And when the list of metadata keys that we do _not_ want to send is large it can be helpful to use a [Bloblang mapping](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) in order to give all of these "private" keys a common prefix: ```yaml pipeline: processors: # Has an explicit list of public metadata keys, and everything else is given # an underscore prefix. - mapping: | let allowed_meta = [ "foo", "bar", "baz", ] meta = @.map_each_key(key -> if !$allowed_meta.contains(key) { "_" + key }) output: kafka: addresses: [ TODO ] topic: ${! meta("_target_topic") } metadata: exclude_prefixes: [ "_" ] ``` --- # Page 282: Monitor Data Pipelines on BYOC and Dedicated Clusters **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/monitor-connect.md --- # Monitor Data Pipelines on BYOC and Dedicated Clusters > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Monitor Data Pipelines on BYOC and Dedicated Clusters latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/monitor-connect page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/monitor-connect.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/monitor-connect.adoc description: Configure Prometheus monitoring of your data pipelines on BYOC clusters. page-git-created-date: "2024-09-09" page-git-modified-date: "2024-12-03" --- You can configure monitoring on BYOC and Dedicated clusters to understand the behavior, health, and performance of your data pipelines. Redpanda Connect automatically exports [detailed metrics for each component of your data pipeline](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/) to a Prometheus endpoint, along with metrics for all other cluster services. You don’t need to update the configuration of your pipeline. ## [](#configure-prometheus)Configure Prometheus To monitor a BYOC cluster in [Prometheus](https://prometheus.io/): 1. On the Redpanda Cloud **Overview** page for your cluster, under **How to connect**, click the **Prometheus** tab. 2. Click the copy icon next to **Prometheus YAML** to copy the contents to your clipboard. The YAML contains the Prometheus scrape target configuration, as well as authentication, for the cluster. ```yaml - job_name: redpandaCloud-sample static_configs: - targets: - console-..fmc.cloud.redpanda.com metrics_path: /api/cloud/prometheus/public_metrics basic_auth: username: prometheus password: "" scheme: https ``` 3. Save the YAML configuration to Prometheus replacing the following placeholders: - `.`: ID and identifier from the **HTTPS endpoint**. - ``: Copy and paste the onscreen Prometheus password. Metrics from Redpanda endpoints are scraped into Prometheus. The metrics for each data pipeline are labelled by pipeline ID. ## [](#use-redpanda-monitoring-examples)Use Redpanda monitoring examples For hands-on learning, Redpanda provides a repository with examples of monitoring Redpanda with Prometheus and Grafana: [redpanda-data/observability](https://github.com/redpanda-data/observability/tree/main/cloud). ![Example Redpanda Connect Dashboard^](https://docs.redpanda.com/cloud-data-platform/shared/_images/redpanda_connect_dashboard.png) It includes [an example Grafana dashboard for Redpanda Connect](https://github.com/redpanda-data/observability/blob/main/grafana-dashboards/Redpanda-Connect-Dashboard.json) and a [sandbox environment](https://github.com/redpanda-data/observability#sandbox-environment) in which you launch a Dockerized Redpanda cluster and create a custom workload to monitor with dashboards. --- # Page 283: Process Pipelines **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/processing_pipelines.md --- # Process Pipelines > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Process Pipelines latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/processing_pipelines page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/processing_pipelines.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/processing_pipelines.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- If you have processors that are heavy on CPU and aren’t specific to a certain input or output they are best suited for the pipeline section. It is advantageous to use the pipeline section as it allows you to set an explicit number of parallel threads of execution: ```yaml input: resource: foo pipeline: threads: 4 processors: - mapping: | root = this fans = fans.map_each(match { this.obsession > 0.5 => this _ => deleted() }) output: resource: bar ``` If the field `threads` is set to `-1` (the default) it will automatically match the number of logical CPUs available. By default almost all Redpanda Connect sources will utilize as many processing threads as have been configured, which makes horizontal scaling easy. --- # Page 284: Manage Pipeline Resources on Clusters **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/resource-management.md --- # Manage Pipeline Resources on Clusters > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Pipeline Resources on Clusters latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/resource-management page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/resource-management.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/resource-management.adoc description: Learn how to set an initial resource limit for a standard data pipeline (excluding Ollama AI components) and how to manually scale the pipeline’s resources to improve performance. page-git-created-date: "2024-12-18" page-git-modified-date: "2026-05-26" --- Learn how to set an initial resource limit for a standard data pipeline (excluding Ollama AI components) and how to manually scale the pipeline’s resources to improve performance. ## [](#prerequisites)Prerequisites - A running Redpanda Cloud cluster. - An estimate of the throughput of your data pipeline. You can get some basic statistics by running your data pipeline locally using the [`benchmark` processor](https://docs.redpanda.com/connect/components/processors/benchmark/). ### [](#understanding-compute-units)Understanding compute units A compute unit allocates a specific amount of server resources (CPU and memory) to a data pipeline to handle message throughput. By default, each pipeline is allocated one compute unit, which includes 0.1 CPU (100 milliCPU or `100m`) and 400 MB (`400M`) of memory. For sizing purposes, one compute unit supports an estimated message throughput of 1 MB/s. However, actual performance depends on the complexity of a pipeline, including the components it contains and the processing it does. You can allocate a maximum of 72 compute units per pipeline. You can add compute units in increments of one up to 15 compute units. Beyond this, scaling options increase to 33 and then to 72 compute units. This scaling strategy is based on the number of machine cores required to provision resources, which scale from two to four, and then to eight cores. Server resources are charged at an [hourly rate in compute unit hours (compute/hour)](https://docs.redpanda.com/cloud-data-platform/billing/billing/#redpanda-connect-pipeline-metrics). | Number of compute units | CPU | Memory | | --- | --- | --- | | 1 | 0.1 CPU (100m) | 400 MB (400M) | | 2 | 0.2 CPU (200m) | 800 MB (800M) | | 3 | 0.3 CPU (300m) | 1.2 GB (1200M) | | 4 | 0.4 CPU (400m) | 1.6 GB (1600M) | | 5 | 0.5 CPU (500m) | 2.0 GB (2000M) | | 6 | 0.6 CPU (600m) | 2.4 GB (2400M) | | 7 | 0.7 CPU (700m) | 2.8 GB (2800M) | | 8 | 0.8 CPU (800m) | 3.2 GB (3200M) | | 9 | 0.9 CPU (900m) | 3.6 GB (3600M) | | 10 | 1.0 CPU (1000m) | 4.0 GB (4000M) | | 11 | 1.1 CPU (1100m) | 4.4 GB (4400M) | | 12 | 1.2 CPU (1200m) | 4.8 GB (4800M) | | 13 | 1.3 CPU (1300m) | 5.2 GB (5200M) | | 14 | 1.4 CPU (1400m) | 5.6 GB (5600M) | | 15 | 1.5 CPU (1500m) | 6.0 GB (6000M) | | 33 | 3.3 CPU (3300m) | 13.2 GB (13200M) | | 72 | 7.2 CPU (7200m) | 28.8 GB (28800M) | > 📝 **NOTE** > > A GPU machine is automatically assigned to each pipeline that contains embedded Ollama AI components. By default, GPU-enabled pipelines are allocated eight compute units. For larger workloads, you can scale them up to a maximum of 30 compute units. ### [](#set-an-initial-resource-limit)Set an initial resource limit When you create a data pipeline, you can allocate a fixed amount of server resources to it using compute units. > 📝 **NOTE** > > If your pipeline reaches the CPU limit, it becomes throttled, which reduces the data processing rate. If it reaches the memory limit, the pipeline restarts. To set an initial resource limit: 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. On the **Clusters** page, select the cluster where you want to add a pipeline. 3. Go to the **Connect** page. 4. Select the **Redpanda Connect** tab. 5. Click **Create pipeline**. 6. Enter details for your pipeline, including a short name and description. 7. For **Compute units**, leave the default **1** compute unit to experiment with pipelines that create low message volumes. For higher throughputs, you can allocate a maximum of 72 compute units. 8. For **Configuration**, paste your pipeline configuration and click **Create** to run it. ### [](#scale-resources)Scale resources View the server resources allocated to a data pipeline, and manually scale those resources to improve performance or decrease resource consumption. To view resources already allocated to a data pipeline: #### Cloud UI 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Go to the cluster where the pipeline is set up. 3. On the **Connect** page, select your pipeline and look at the value for **Resources**. - CPU resources are displayed first, in milliCPU. For example, `1` compute unit is `100m` or 0.1 CPU. - Memory is displayed next in megabytes. For example, `1` compute unit is `400M` or 400 MB. #### Data Plane API 1. [Authenticate and get the base URL](https://docs.redpanda.com/api/doc/cloud-dataplane/topic/topic-quickstart) for the Data Plane API. 2. Make a request to [`GET /v1/redpanda-connect/pipelines`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_listpipelines), which lists details of all pipelines on your cluster by ID. - Memory (`memory_shares`) is displayed in megabytes. For example, `1` compute unit is `400M` or 400 MB. - CPU resources (`cpu_shares`) are displayed in milliCPU. For example, `1` compute unit is `100m` or 0.1 CPU. To scale the resources for a pipeline: #### Cloud UI 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Go to the cluster where the pipeline is set up. 3. On the **Connect** page, select your pipeline and click **Edit**. 4. For **Compute units**, update the number of compute units. You can allocate a maximum of 72 compute units per pipeline. 5. Click **Update** to apply your changes. The specified resources are available immediately. #### Data Plane API You can only update CPU resources using the Data Plane API. For every 0.1 CPU that you allocate, Redpanda Cloud automatically reserves 400 MB of memory for the exclusive use of the pipeline. 1. [Authenticate and get the base URL](https://docs.redpanda.com/api/doc/cloud-dataplane/topic/topic-quickstart) for the Data Plane API, if you haven’t already. 2. Make a request to [`GET /v1/redpanda-connect/pipelines/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_getpipeline), including the ID of the pipeline you want to update. You’ll use the returned values in the next step. 3. Now make a request to [`PUT /v1/redpanda-connect/pipelines/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_updatepipeline), to update the pipeline resources: - Reuse the values returned by your `GET` request to populate the request body. - Replace the `cpu_shares` value with the resources you want to allocate, and enter any valid value for `memory_shares`. This example allocates 0.2 CPU or 200 milliCPU to a data pipeline. For `cpu_shares`, `0.1` CPU is the minimum allocation. ```bash curl -X PUT "https:///v1/redpanda-connect/pipelines/xxx..." \ -H 'accept: application/json'\ -H 'authorization: Bearer xxx...' \ -H "content-type: application/json" \ -d '{ "config_yaml": "input:\n generate:\n interval: 1s\n mapping: |\n root.id = uuid_v4()\n root.user.name = fake(\"name\")\n root.user.email = fake(\"email\")\n root.content = fake(\"paragraph\")\n\npipeline:\n processors:\n - mutation: |\n root.title = \"PRIVATE AND CONFIDENTIAL\"\n\noutput:\n kafka_franz:\n seed_brokers:\n - seed-j888.byoc.prd.cloud.redpanda.com:9092\n sasl:\n mechanism: SCRAM-SHA-256\n password: password\n username: connect\n topic: processed-emails\n tls:\n enabled: true\n", "description": "Email processor", "display_name": "emailprocessor-pipeline", "resources": { "memory_shares": "800M", "cpu_shares": "200m" } }' ``` A successful response shows the updated resource allocations with the `cpu_shares` value returned in milliCPU. 4. Make a request to [`GET /v1/redpanda-connect/pipelines`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_listpipelines) to verify your pipeline resource updates. --- # Page 285: Manage Secrets **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management.md --- # Manage Secrets > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Secrets latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/secret-management page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/secret-management.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/secret-management.adoc description: Learn how to manage secrets in Redpanda Connect using the Cloud Console, Data Plane API, or Terraform, and learn how to add them to your data pipelines. page-git-created-date: "2024-12-03" page-git-modified-date: "2026-08-07" --- Learn how to manage secrets in Redpanda Connect, and how to add them to your data pipelines without exposing them. Secrets are stored in the secret management solution of your cloud provider and are retrieved when you run a pipeline configuration that references them. ## [](#manage-secrets)Manage secrets You can manage secrets from the Cloud UI or the Data Plane API. If you manage your Redpanda Cloud resources with Terraform, you can also create secrets with the `redpanda_secret` resource. See [Manage cluster secrets](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/#manage-cluster-secrets). ### [](#create-a-secret)Create a secret You can create a secret and reference it in multiple data pipelines on the same cluster. #### Cloud UI 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Go to the **Secrets Store** page. 3. Click **Create secret**. 4. For **ID**, enter a name for the secret. You cannot rename the secret once it is created. 5. For **Value**, enter the secret you need to add. 6. For **Scopes**, select Redpanda Connect. 7. Optionally, add labels to help organize your secrets. 8. Click **Create**. You can now [add the secret to your data pipeline](#add-a-secret-to-a-data-pipeline). #### Data Plane API You must use a Base64-encoded secret. 1. [Authenticate and get the base URL](https://docs.redpanda.com/api/doc/cloud-dataplane/topic/topic-quickstart) for the Data Plane API. 2. Make a request to [`POST /v1/secrets`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-secretservice_createsecret). ```bash curl -X POST "https:///v1/secrets" \ -H 'accept: application/json'\ -H 'authorization: Bearer '\ -H 'content-type: application/json' \ -d '{"id":"","scopes":["SCOPE_REDPANDA_CONNECT"],"secret_data":""}' ``` You must include the following values: - ``: The base URL for the Data Plane API. - ``: The API key you generated during authentication. - ``: The ID or name of the secret you want to add. Use only the following characters: `^[A-Z][A-Z0-9_]*$`. - ``: The Base64-encoded secret. - This scope: `"SCOPE_REDPANDA_CONNECT"`. The response returns the name of the secret and the scope `"SCOPE_REDPANDA_CONNECT"`. You can now [add the secret to your data pipeline](#add-a-secret-to-a-data-pipeline). ### [](#update-a-secret)Update a secret You can only update the secret value, not its name. > 📝 **NOTE** > > Changes to secret values do not take effect until a pipeline is restarted. #### Cloud UI 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Go to the **Secrets Store** page. 3. Find the secret you want to update, and click the edit icon. 4. Enter the new secret value or labels, and click **Update**. 5. Start and stop any pipelines that reference the secret. #### Data Plane API You must use a Base64-encoded secret. 1. [Authenticate and get the base URL](https://docs.redpanda.com/api/doc/cloud-dataplane/topic/topic-quickstart) for the Data Plane API. 2. Make a request to [`PUT /v1/secrets/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-secretservice_updatesecret). ```bash curl -X PUT "https:///v1/secrets/" \ -H 'accept: application/json'\ -H 'authorization: Bearer '\ -H 'content-type: application/json' \ -d '{"scopes":["SCOPE_REDPANDA_CONNECT"],"secret_data":""}' ``` You must include the following values: - ``: The base URL for the Data Plane API. - ``: The name of the secret you want to update. - ``: The API key you generated during authentication. - This scope: `"SCOPE_REDPANDA_CONNECT"`. - ``: Your new Base64-encoded secret. The response returns the name of the secret and the scope `"SCOPE_REDPANDA_CONNECT"`. ### [](#delete-a-secret)Delete a secret Before you delete a secret, make sure that you remove references to it from your data pipelines. > 📝 **NOTE** > > Changes do not affect pipelines that are already running. #### Cloud UI 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Go to the **Secrets Store** page. 3. Find the secret you want to remove, and click the delete icon. 4. Confirm your deletion. #### Data Plane API 1. [Authenticate and get the base URL](https://docs.redpanda.com/api/doc/cloud-dataplane/topic/topic-quickstart) for the Data Plane API. 2. Make a request to [`DELETE /v1/secrets/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-secretservice_deletesecret). ```bash curl -X DELETE "https:///v1/secrets/" \ -H 'accept: application/json'\ -H 'authorization: Bearer '\ ``` You must include the following values: - ``: The base URL for the Data Plane API. - ``: The name of the secret you want to delete. - ``: The API key you generated during authentication. ## [](#add-a-secret-to-a-data-pipeline)Add a secret to a data pipeline ### Cloud UI 1. Go to the **Connect** page, and create a pipeline (or open an existing pipeline to edit). 2. Click the **Secret** button to add a new or existing secret to the pipeline. ### Data Plane API You can add a secret to any pipeline in your cluster using the notation `${secrets.SECRET_NAME}`. For example: ```yml sasl: - mechanism: SCRAM-SHA-256 username: "user" password: "${secrets.PASSWORD}" ``` --- # Page 286: Unit Testing **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/unit_testing.md --- # Unit Testing > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Unit Testing latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/unit_testing page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/unit_testing.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/unit_testing.adoc page-git-created-date: "2024-09-09" page-git-modified-date: "2026-05-26" --- The Redpanda Connect service offers a command `rpk connect test` for running unit tests on sections of a configuration file. This makes it easy to protect your config files from regressions over time. ## [](#writing-a-test)Writing a test Let’s imagine we have a configuration file `foo.yaml` containing some processors: ```yaml input: kafka: addresses: [ TODO ] topics: [ foo, bar ] consumer_group: foogroup pipeline: processors: - mapping: '"%vend".format(content().uppercase().string())' output: aws_s3: bucket: TODO path: '${! meta("kafka_topic") }/${! json("message.id") }.json' ``` One way to write our unit tests for this config is to accompany it with a file of the same name and extension but suffixed with `_benthos_test`, which in this case would be `foo_benthos_test.yaml`. ```yml tests: - name: example test target_processors: '/pipeline/processors' environment: {} input_batch: - content: 'example content' metadata: example_key: example metadata value output_batches: - - content_equals: EXAMPLE CONTENTend metadata_equals: example_key: example metadata value ``` Under `tests` we have a list of any number of unit tests to execute for the config file. Each test is run in complete isolation, including any resources defined by the config file. Tests should be allocated a unique `name` that identifies the feature being tested. The field `target_processors` is either the label of a processor to test, or a [JSON Pointer](https://tools.ietf.org/html/rfc6901) that identifies the position of a processor, or list of processors, within the file which should be executed by the test. For example a value of `foo` would target a processor with the label `foo`, and a value of `/input/processors` would target all processors within the input section of the config. The field `environment` allows you to define an object of key/value pairs that set environment variables to be evaluated during the parsing of the target config file. These are unique to each test, allowing you to test different environment variable interpolation combinations. The field `input_batch` lists one or more messages to be fed into the targeted processors as a batch. Each message of the batch may have its raw content defined as well as metadata key/value pairs. For the common case where the messages are in JSON format, you can use `json_content` instead of `content` to specify the message structurally rather than verbatim. The field `output_batches` lists any number of batches of messages which are expected to result from the target processors. Each batch lists any number of messages, each one defining [`conditions`](#output-conditions) to describe the expected contents of the message. If the number of batches defined does not match the resulting number of batches the test will fail. If the number of messages defined in each batch does not match the number in the resulting batches the test will fail. If any condition of a message fails then the test fails. ### [](#inline-tests)Inline tests Sometimes it’s more convenient to define your tests within the config being tested. This is fine, simply add the `tests` field to the end of the config being tested. ### [](#bloblang-tests)Bloblang tests Sometimes when working with large [Bloblang mappings](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) it’s preferred to have the full mapping in a separate file to your Redpanda Connect configuration. In this case it’s possible to write unit tests that target and execute the mapping directly with the field `target_mapping`, which when specified is interpreted as either an absolute path or a path relative to the test definition file that points to a file containing only a Bloblang mapping. For example, if we were to have a file `cities.blobl` containing a mapping: ```bloblang root.Cities = this.locations. filter(loc -> loc.state == "WA"). map_each(loc -> loc.name). sort().join(", ") ``` We can accompany it with a test file `cities_test.yaml` containing a regular test definition: ```yml tests: - name: test cities mapping target_mapping: './cities.blobl' environment: {} input_batch: - content: | { "locations": [ {"name": "Seattle", "state": "WA"}, {"name": "New York", "state": "NY"}, {"name": "Bellevue", "state": "WA"}, {"name": "Olympia", "state": "WA"} ] } output_batches: - - json_equals: {"Cities": "Bellevue, Olympia, Seattle"} ``` And execute this test the same way we execute other Redpanda Connect tests (`rpk connect test ./dir/cities_test.yaml`, `rpk connect test ./dir/…​`, etc). ### [](#fragmented-tests)Fragmented tests Sometimes the number of tests you need to define in order to cover a config file is so vast that it’s necessary to split them across multiple test definition files. This is possible but Redpanda Connect still requires a way to detect the configuration file being targeted by these fragmented test definition files. In order to do this we must prefix our `target_processors` field with the path of the target relative to the definition file. The syntax of `target_processors` in this case is a full [JSON Pointer](https://tools.ietf.org/html/rfc6901) that should look something like `target.yaml#/pipeline/processors`. For example, if we saved our test definition above in an arbitrary location like `./tests/first.yaml` and wanted to target our original `foo.yaml` config file, we could do that with the following: ```yml tests: - name: example test target_processors: '../foo.yaml#/pipeline/processors' environment: {} input_batch: - content: 'example content' metadata: example_key: example metadata value output_batches: - - content_equals: EXAMPLE CONTENTend metadata_equals: example_key: example metadata value ``` ## [](#input-definitions)Input Definitions ### [](#content)`content` Sets the raw content of the message. ### [](#json_content)`json_content` ```yml json_content: foo: foo value bar: [ element1, 10 ] ``` Sets the raw content of the message to a JSON document matching the structure of the value. ### [](#file_content)`file_content` ```yml file_content: ./foo/bar.txt ``` Sets the raw content of the message by reading a file. The path of the file should be relative to the path of the test file. ### [](#metadata)`metadata` A map of key/value pairs that sets the metadata values of the message. ## [](#output-conditions)Output Conditions ### [](#bloblang)`bloblang` ```yml bloblang: 'this.age > 10 && @foo.length() > 0' ``` Executes a [Bloblang expression](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) on a message, if the result is anything other than a boolean equalling `true` the test fails. ### [](#content_equals)`content_equals` ```yml content_equals: example content ``` Checks the full raw contents of a message against a value. ### [](#content_matches)`content_matches` ```yml content_matches: "^foo [a-z]+ bar$" ``` Checks whether the full raw contents of a message matches a regular expression (re2). ### [](#metadata_equals)`metadata_equals` ```yml metadata_equals: example_key: example metadata value ``` Checks a map of metadata keys to values against the metadata stored in the message. If there is a value mismatch between a key of the condition versus the message metadata this condition will fail. ### [](#file_equals)`file_equals` ```yml file_equals: ./foo/bar.txt ``` Checks that the contents of a message matches the contents of a file. The path of the file should be relative to the path of the test file. ### [](#file_json_equals)`file_json_equals` ```yml file_json_equals: ./foo/bar.json ``` Checks that both the message and the file contents are valid JSON documents, and that they are structurally equivalent. Will ignore formatting and ordering differences. The path of the file should be relative to the path of the test file. ### [](#json_equals)`json_equals` ```yml json_equals: { "key": "value" } ``` Checks that both the message and the condition are valid JSON documents, and that they are structurally equivalent. Will ignore formatting and ordering differences. You can also structure the condition content as YAML and it will be converted to the equivalent JSON document for testing: ```yml json_equals: key: value ``` ### [](#json_contains)`json_contains` ```yml json_contains: { "key": "value" } ``` Checks that both the message and the condition are valid JSON documents, and that the message is a superset of the condition. ## [](#running-tests)Running tests Executing tests for a specific config can be done by pointing the subcommand `test` at either the config to be tested or its test definition, e.g. `rpk connect test ./config.yaml` and `rpk connect test ./config_benthos_test.yaml` are equivalent. The `test` subcommand also supports wildcard patterns e.g. `rpk connect test ./foo/*.yaml` will execute all tests within matching files. In order to walk a directory tree and execute all tests found you can use the shortcut `./…​`, e.g. `rpk connect test ./…​` will execute all tests found in the current directory, any child directories, and so on. If you want to allow components to write logs at a provided level to stdout when running the tests, you can use `rpk connect test --log `. Please consult the [logger docs](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/logger/about/) for further details. ## [](#mocking-processors)Mocking processors BETA: This feature is currently in a BETA phase, which means breaking changes could be made if a fundamental issue with the feature is found. Sometimes you’ll want to write tests for a series of processors, where one or more of them are networked (or otherwise stateful). Rather than creating and managing mocked services you can define mock versions of those processors in the test definition. For example, if we have a config with the following processors: ```yaml pipeline: processors: - mapping: 'root = "simon says: " + content()' - label: get_foobar_api http: url: http://example.com/foobar verb: GET - mapping: 'root = content().uppercase()' ``` Rather than create a fake service for the `http` processor to interact with we can define a mock in our test definition that replaces it with a [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/). Mocks are configured as a map of labels that identify a processor to replace and the config to replace it with: ```yaml tests: - name: mocks the http proc target_processors: '/pipeline/processors' mocks: get_foobar_api: mapping: 'root = content().string() + " this is some mock content"' input_batch: - content: "hello world" output_batches: - - content_equals: "SIMON SAYS: HELLO WORLD THIS IS SOME MOCK CONTENT" ``` With the above test definition the `http` processor will be swapped out for `mapping: 'root = content().string() + " this is some mock content"'`. For the purposes of mocking it is recommended that you use a [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/) that simply mutates the message in a way that you would expect the mocked processor to. > 📝 **NOTE** > > It’s not currently possible to mock components that are imported as separate resource files (using `--resource`/`-r`). It is recommended that you mock these by maintaining separate definitions for test purposes (`-r "./test/*.yaml"`). ### [](#more-granular-mocking)More granular mocking It is also possible to target specific fields within the test config by [JSON pointers](https://tools.ietf.org/html/rfc6901) as an alternative to labels. The following test definition would create the same mock as the previous: ```yaml tests: - name: mocks the http proc target_processors: '/pipeline/processors' mocks: /pipeline/processors/1: mapping: 'root = content().string() + " this is some mock content"' input_batch: - content: "hello world" output_batches: - - content_equals: "SIMON SAYS: HELLO WORLD THIS IS SOME MOCK CONTENT" ``` ## [](#fields)Fields The schema of a template file is as follows: ### [](#tests)`tests` A list of one or more unit tests to execute. **Type**: `array` ### [](#tests-name)`tests[].name` The name of the test, this should be unique and give a rough indication of what behavior is being tested. **Type**: `string` ### [](#tests-environment)`tests[].environment` An optional map of environment variables to set for the duration of the test. **Type**: `object` ### [](#tests-target_processors)`tests[].target_processors` A \[JSON Pointer\]\[json-pointer\] that identifies the specific processors which should be executed by the test. The target can either be a single processor or an array of processors. Alternatively a resource label can be used to identify a processor. It is also possible to target processors in a separate file by prefixing the target with a path relative to the test file followed by a # symbol. **Type**: `string` **Default**: `"/pipeline/processors"` ```yml # Examples target_processors: foo_processor target_processors: /pipeline/processors/0 target_processors: target.yaml#/pipeline/processors target_processors: target.yaml#/pipeline/processors ``` ### [](#tests-target_mapping)`tests[].target_mapping` A file path relative to the test definition path of a Bloblang file to execute as an alternative to testing processors with the `target_processors` field. This allows you to define unit tests for Bloblang mappings directly. **Type**: `string` **Default**: `""` ### [](#tests-mocks)`tests[].mocks` An optional map of processors to mock. Keys should contain either a label or a JSON pointer of a processor that should be mocked. Values should contain a processor definition, which will replace the mocked processor. Most of the time you’ll want to use a \[`mapping` processor\]\[processors.mapping\] here, and use it to create a result that emulates the target processor. **Type**: `object` ```yml # Examples mocks: get_foobar_api: mapping: root = content().string() + " this is some mock content" mocks: /pipeline/processors/1: mapping: root = content().string() + " this is some mock content" ``` ### [](#tests-input_batch)`tests[].input_batch` Define a batch of messages to feed into your test, specify either an `input_batch` or a series of `input_batches`. **Type**: `array` ### [](#tests-input_batch-content)`tests[].input_batch[].content` The raw content of the input message. **Type**: `string` ### [](#tests-input_batch-json_content)`tests[].input_batch[].json_content` Sets the raw content of the message to a JSON document matching the structure of the value. **Type**: `object` ```yml # Examples json_content: bar: - element1 - 10 foo: foo value ``` ### [](#tests-input_batch-file_content)`tests[].input_batch[].file_content` Sets the raw content of the message by reading a file. The path of the file should be relative to the path of the test file. **Type**: `string` ```yml # Examples file_content: ./foo/bar.txt ``` ### [](#tests-input_batch-metadata)`tests[].input_batch[].metadata` A map of metadata key/values to add to the input message. **Type**: `object` ### [](#tests-input_batches)`tests[].input_batches` Define a series of batches of messages to feed into your test, specify either an `input_batch` or a series of `input_batches`. **Type**: `two-dimensional array` ### [](#tests-input_batches-content)`tests[].input_batches[][].content` The raw content of the input message. **Type**: `string` ### [](#tests-input_batches-json_content)`tests[].input_batches[][].json_content` Sets the raw content of the message to a JSON document matching the structure of the value. **Type**: `object` ```yml # Examples json_content: bar: - element1 - 10 foo: foo value ``` ### [](#tests-input_batches-file_content)`tests[].input_batches[][].file_content` Sets the raw content of the message by reading a file. The path of the file should be relative to the path of the test file. **Type**: `string` ```yml # Examples file_content: ./foo/bar.txt ``` ### [](#tests-input_batches-metadata)`tests[].input_batches[][].metadata` A map of metadata key/values to add to the input message. **Type**: `object` ### [](#tests-output_batches)`tests[].output_batches` List of output batches. **Type**: `two-dimensional array` ### [](#tests-output_batches-bloblang)`tests[].output_batches[][].bloblang` Executes a Bloblang mapping on the output message, if the result is anything other than a boolean equalling `true` the test fails. **Type**: `string` ```yml # Examples bloblang: this.age > 10 && @foo.length() > 0 ``` ### [](#tests-output_batches-content_equals)`tests[].output_batches[][].content_equals` Checks the full raw contents of a message against a value. **Type**: `string` ### [](#tests-output_batches-content_matches)`tests[].output_batches[][].content_matches` Checks whether the full raw contents of a message matches a regular expression (re2). **Type**: `string` ```yml # Examples content_matches: ^foo [a-z]+ bar$ ``` ### [](#tests-output_batches-metadata_equals)`tests[].output_batches[][].metadata_equals` Checks a map of metadata keys to values against the metadata stored in the message. If there is a value mismatch between a key of the condition versus the message metadata this condition will fail. **Type**: `object` ```yml # Examples metadata_equals: example_key: example metadata value ``` ### [](#tests-output_batches-file_equals)`tests[].output_batches[][].file_equals` Checks that the contents of a message matches the contents of a file. The path of the file should be relative to the path of the test file. **Type**: `string` ```yml # Examples file_equals: ./foo/bar.txt ``` ### [](#tests-output_batches-file_json_equals)`tests[].output_batches[][].file_json_equals` Checks that both the message and the file contents are valid JSON documents, and that they are structurally equivalent. Will ignore formatting and ordering differences. The path of the file should be relative to the path of the test file. **Type**: `string` ```yml # Examples file_json_equals: ./foo/bar.json ``` ### [](#tests-output_batches-json_equals)`tests[].output_batches[][].json_equals` Checks that both the message and the condition are valid JSON documents, and that they are structurally equivalent. Will ignore formatting and ordering differences. **Type**: `object` ```yml # Examples json_equals: key: value ``` ### [](#tests-output_batches-json_contains)`tests[].output_batches[][].json_contains` Checks that both the message and the condition are valid JSON documents, and that the message is a superset of the condition. **Type**: `object` ```yml # Examples json_contains: key: value ``` ### [](#tests-output_batches-file_json_contains)`tests[].output_batches[][].file_json_contains` Checks that both the message and the file contents are valid JSON documents, and that the message is a superset of the condition. Will ignore formatting and ordering differences. The path of the file should be relative to the path of the test file. **Type**: `string` ```yml # Examples file_json_contains: ./foo/bar.json ``` --- # Page 287: Windowed Processing **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/windowed_processing.md --- # Windowed Processing > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Windowed Processing latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/configuration/windowed_processing page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/configuration/windowed_processing.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/configuration/windowed_processing.adoc description: Learn how to process periodic windows of messages with Redpanda Connect. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- A window is a batch of messages made with respect to time, with which we are able to perform processing that can analyze or aggregate the messages of the window. This is useful in stream processing as the dataset is never "complete", and therefore in order to perform analysis against a collection of messages we must do so by creating a continuous feed of windows (collections), where our analysis is made against each window. For example, given a stream of messages relating to cars passing through various traffic lights: ```json { "traffic_light": "cbf2eafc-806e-4067-9211-97be7e42cee3", "created_at": "2021-08-07T09:49:35Z", "registration_plate": "AB1C DEF", "passengers": 3 } ``` Windowing allows us to produce a stream of messages representing the total traffic for each light every hour: ```json { "traffic_light": "cbf2eafc-806e-4067-9211-97be7e42cee3", "created_at": "2021-08-07T10:00:00Z", "unique_cars": 15, "passengers": 43 } ``` ## [](#creating-windows)Creating windows The first step in processing windows is producing the windows themselves, this can be done by configuring a window producing buffer after your input: ### System A `system_window` buffer creates windows by following the system clock of the running machine. Windows will be created and emitted at predictable times, but this also means windows for historic data will not be emitted and therefore prevents backfills of traffic data: ```yaml input: kafka: addresses: [ TODO ] topics: [ traffic_data ] consumer_group: traffic_consumer checkpoint_limit: 1000 buffer: system_window: timestamp_mapping: root = this.created_at size: 1h allowed_lateness: 3m ``` For more information about this buffer refer to the `system_window` buffer docs. ## [](#grouping)Grouping With a window buffer chosen our stream of messages will be emitted periodically as batches of all messages that fit within each window. Since we want to analyse the window separately for each traffic light we need to expand this single batch out into one for each traffic light identifier within the window. For that purpose we have two processor options: [`group_by`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/group_by/) and [`group_by_value`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/group_by_value/). In our case we want to group by the value of the field `traffic_light` of each message, which we can do with the following: ```yaml pipeline: processors: - group_by_value: value: ${! json("traffic_light") } ``` ## [](#aggregating)Aggregating Once our window has been grouped the next step is to calculate the aggregated passenger and unique cars counts. For this purpose the Redpanda Connect [mapping language Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) comes in handy as the method [`from_all`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#from_all) executes the target function against the entire batch and returns an array of the values, allowing us to mutate the result with chained methods such as [`sum`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#sum): ```yaml pipeline: processors: - group_by_value: value: ${! json("traffic_light") } - mapping: | let is_first_message = batch_index() == 0 root.traffic_light = this.traffic_light root.created_at = @window_end_timestamp root.total_cars = if $is_first_message { json("registration_plate").from_all().unique().length() } root.passengers = if $is_first_message { json("passengers").from_all().sum() } # Only keep the first batch message containing the aggregated results. root = if ! $is_first_message { deleted() } ``` [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) is very powerful, and by using [`from`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#from) and [`from_all`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#from_all) it’s possible to perform a wide range of batch-wide processing. If you fancy a challenge try updating the above mapping to only count passengers from the first journey of each registration plate in the window (hint: the [`fold` method](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#fold) might come in handy). --- # Page 288: Redpanda Connect Quickstart **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart.md --- # Redpanda Connect Quickstart > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Connect Quickstart latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/connect-quickstart page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/connect-quickstart.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/connect-quickstart.adoc description: Learn how to quickly start building data pipelines with Redpanda Connect. page-topic-type: tutorial personas: streaming_developer learning-objective-1: Build a producer pipeline that generates and publishes data to a topic learning-objective-2: Build a consumer pipeline that reads, transforms, and logs data from a topic page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-24" --- In this quickstart, you build data pipelines to generate, transform, and handle streaming data end-to-end. You create two pipelines: one that generates dad jokes and writes them to a topic in your cluster, and another that reads those jokes and gives each one a random "cringe rating". After completing this quickstart, you will be able to: - Build a producer pipeline that generates and publishes data to a topic - Build a consumer pipeline that reads, transforms, and logs data from a topic ## [](#prerequisites)Prerequisites You must have a Redpanda Cloud account with a Serverless, Dedicated, or standard BYOC cluster. If you don’t already have an account, [sign up for a free trial](https://redpanda.com/try-redpanda/cloud-trial). > 📝 **NOTE** > > Clusters can create up to 100 pipelines. For additional pipelines, contact [Redpanda support](https://support.redpanda.com/hc/en-us/requests/new). ## [](#quickstart-pipelines)Quickstart pipelines This quickstart creates the following pipelines: - The first pipeline produces dad jokes and writes them to a topic in your cluster. - The second pipeline consumes those dad jokes and gives each one a random "cringe rating" from 1-10. The **producer pipeline** uses the following Redpanda Connect components: | Component type | Component | Purpose | | --- | --- | --- | | Input | generate | Creates jokes | | Output | redpanda | Writes messages to your topic | | Processor | log | Logs generated messages | | Processor | catch | Catches errors | The **consumer pipeline** uses the following Redpanda Connect components: | Component type | Component | Purpose | | --- | --- | --- | | Input | redpanda | Reads messages from your topic | | Output | drop | Drops the processed messages | | Processor | bloblang | Processes ratings | | Processor | log | Logs processed messages | | Processor | catch | Catches errors | ## [](#visual-editor)Explore the Visual editor When you create or edit a pipeline, the editor has two tabs: **YAML** and **Visual**. When you open an existing pipeline, a **Monitor** tab appears alongside them, and the **Visual** tab is read-only. Click **Edit pipeline** to make changes. - The **YAML** tab gives you an IDE-like experience: after you add a component, click its leaf icon in the left sidebar to open that component’s reference documentation, or use the slash command menu to insert secrets and other variables. - The **Visual** tab renders your pipeline as an interactive node diagram, so you can inspect and edit inputs, processors, outputs, and control-flow constructs (such as [`switch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/switch/), [`branch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/), and [`try`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/try/)/[`catch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/catch/)) without hand-writing YAML. Both tabs stay in sync, so you can switch between them at any time. A `switch` groups its cases into a single box, with each case labeled by its routing condition. A `try` groups the processors it protects into its own box, and a `catch` that follows it is connected with a dashed line labeled **ON ERROR**, so the error path is visually distinct from the main data flow: ![The Visual tab rendering a switch with two cases and a try/catch pair joined by a dashed ON ERROR line](https://docs.redpanda.com/cloud-data-platform/shared/_images/rpcn-visual-editor-control-flow.png) This quickstart uses the **YAML** tab, so click **YAML** if the editor opens on the **Visual** tab. ## [](#build-a-producer-pipeline)Build a producer pipeline Every pipeline requires an input and an output in a configuration file. You can select components in the left sidebar and customize the YAML in the editor. To create the producer pipeline: 1. Go to the **Connect** page for your cluster and click **Create a pipeline**. 2. Enter this name for the pipeline: `joke-generator-producer`. 3. In the left sidebar, click **Add input** and search for and select the `generate` input connector. The YAML for this connector appears in the editor. 4. Click **Add output** and search for and select the `redpanda` output connector. The YAML for this connector also appears in the editor. 5. The `redpanda` connector requires a Redpanda topic and user: 1. In the left sidebar, click **Topics +**, then click **Create topic**. Enter `dad-jokes` for the topic name and click **Create**. 2. In the left sidebar, click **Users +**, then click **Create user**. Enter `connect` for the username and click **Create**. 6. Replace the generated YAML in the editor with the following. This configuration includes the `log` and `catch` processors and the `mapping` for joke generation. Bloblang is Redpanda Connect’s scripting language used to add logic. ```yaml input: generate: interval: 5s count: 0 mapping: | let jokes = [ "Why don't scientists trust atoms? Because they make up everything!", "I'm reading a book about anti-gravity. It's impossible to put down!", "Why did the scarecrow win an award? He was outstanding in his field!", "What do you call a fake noodle? An impasta!", "Why don't eggs tell jokes? They'd crack each other up!", "I used to play piano by ear, but now I use my hands.", "What do you call a bear with no teeth? A gummy bear!", "Why did the bicycle fall over? It was two tired!", "What do you call a fish wearing a crown? A king fish!", "Why don't skeletons fight each other? They don't have the guts!", "What do you call cheese that isn't yours? Nacho cheese!", "Why can't you hear a pterodactyl using the bathroom? Because the 'p' is silent!", "What did the ocean say to the beach? Nothing, it just waved!", "Why did the math book look sad? It had too many problems!", "What do you call a sleeping bull? A bulldozer!", "How do you organize a space party? You planet!", "What's orange and sounds like a parrot? A carrot!", "Why did the coffee file a police report? It got mugged!", "What do you call a can opener that doesn't work? A can't opener!", "Why don't oysters donate to charity? Because they're shellfish!" ] let joke_index = random_int() % $jokes.length() root.joke = $jokes.index($joke_index) root.id = uuid_v4() root.timestamp = now() root.source = "dad-joke-generator" root.joke_length = root.joke.length() pipeline: processors: - log: level: INFO message: "📝 Generating joke: ${! json(\"joke\") }" - catch: - log: level: ERROR message: "❌ Error generating joke: ${! error() }" output: redpanda: seed_brokers: # Optional - ${REDPANDA_BROKERS} tls: enabled: true # Optional (default: false) client_certs: [] sasl: - mechanism: SCRAM-SHA-256 username: ${secrets.KAFKA_USER_CONNECT} password: ${secrets.KAFKA_PASSWORD_CONNECT} topic: dad-jokes # Optional ``` > 📝 **NOTE** > > - Notice the `${REDPANDA_BROKERS}` [contextual variable](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/contextual-variables/) in the configuration. This references your cluster’s bootstrap server address, so you can use it in any pipeline without hardcoding connection details. Use the slash command menu in the YAML editor or use the command palette to insert the Redpanda broker’s contextual variable. > > - Notice `${secrets.KAFKA_USER_CONNECT}` and `${secrets.KAFKA_PASSWORD_CONNECT}`. These reference secrets that you can create using the slash command menu in the YAML editor or on the **Security** page. > > - The Brave browser does not fully support code snippets. 7. Click the **Visual** tab (next to **YAML**) to see this pipeline as a diagram. ![The Visual tab rendering the producer pipeline as connected generate / log / catch / redpanda nodes](https://docs.redpanda.com/cloud-data-platform/shared/_images/rpcn-visual-editor-canvas.png) 8. Click **Save**. Your pipeline details display, and after a few seconds, the pipeline starts running. The pipeline generates jokes and writes the jokes to your Redpanda topic. ### [](#review-the-pipeline-logs)Review the pipeline logs The page loads new log messages as they come in. When Live mode is disabled, you can filter logs, for example, by level, message content, or path. The log shows activity from the past five hours. Click through the log messages to see the startup sequence. For example, you’ll see when the output becomes active: ```json { "instance_id": "d73c39bp7l8c73d7lll0", "label": "", "level": "INFO", "message": "Output type redpanda is now active", "path": "root.output", "pipeline_id": "d73a55ptub9s73agpthg", "time": "2026-03-27T17:43:02.36416142Z" } ``` ### [](#view-the-processed-messages)View the processed messages 1. Go to the **Topics** page for your cluster and select the `dad-jokes` topic. 2. Click any message to see the structure. For example: ```json { "id": "d242c355-4cee-4382-817a-190c7a115a19", "joke": "I used to play piano by ear, but now I use my hands.", "joke_length": 52, "source": "dad-joke-generator", "timestamp": "2026-03-27T15:30:38.963227997Z" } ``` ## [](#build-a-consumer-pipeline)Build a consumer pipeline This next pipeline rates the jokes that you generated in the first pipeline. To create the consumer pipeline: 1. Go back to the **Connect** page for your cluster, and click **Create a pipeline**. 2. Enter this name for the pipeline: `joke-generator-consumer`. 3. In the left sidebar, click **Add input**, and search for and select the `redpanda` input connector. 4. The `redpanda` connector requires a Redpanda topic and user: 1. In the left sidebar, click **Topics +** and select the existing topic `dad-jokes`. 2. In the left sidebar, click **Users +** and select the existing user `connect`. 3. Grant the `connect` user access to the `dad-joke-raters` consumer group: go to the **Security** page for your cluster, select the **Permissions** tab, and click **Create ACL**. Set **Principal** to `connect`, **Resource Type** to **Consumer Group**, and **Resource Name** to `dad-joke-raters`. Set **Operation** to **Read** and **Permission** to **Allow**, then click **Add ACL**. Repeat with **Operation** set to **Describe** to also grant that permission. 5. Click **Add output**, and search for and select the `drop` output connector. (For testing purposes, this output drops messages instead of forwarding them. In a real scenario you would replace the `drop` connector with your real destination.) 6. Replace the generated YAML in the editor with the following configuration, which includes the `bloblang`, `log`, and `catch` processors. > 📝 **NOTE** > > This example explicitly includes several optional configuration fields for the `redpanda` input. They’re shown here for demonstration purposes, so you can see a range of available settings. ```yaml input: redpanda: seed_brokers: # Optional - ${REDPANDA_BROKERS} client_id: benthos # Optional (default: "benthos") tls: enabled: true # Optional (default: false) client_certs: [] sasl: - mechanism: SCRAM-SHA-256 username: ${secrets.KAFKA_USER_CONNECT} password: ${secrets.KAFKA_PASSWORD_CONNECT} metadata_max_age: 5m # Optional (default: "5m") request_timeout_overhead: 10s # Optional (default: "10s") conn_idle_timeout: 20s # Optional (default: "20s") topics: # Required (mutually exclusive with regexp_topics) - dad-jokes regexp_topics: false # Optional (default: false). Mutually exclusive with topics. rebalance_timeout: 45s # Optional (default: "45s") session_timeout: 1m # Optional (default: "1m") heartbeat_interval: 3s # Optional (default: "3s") start_from_oldest: true # Optional (default: true) start_offset: earliest # Optional (default: "earliest") fetch_max_bytes: 50MiB # Optional (default: "50MiB") fetch_max_wait: 5s # Optional (default: "5s") fetch_min_bytes: 1B # Optional (default: "1B") fetch_max_partition_bytes: 1MiB # Optional (default: "1MiB") transaction_isolation_level: read_uncommitted # Optional (default: "read_uncommitted") consumer_group: dad-joke-raters # Optional commit_period: 5s # Optional (default: "5s") partition_buffer_bytes: 1MB # Optional (default: "1MB") topic_lag_refresh_period: 5s # Optional (default: "5s") max_yield_batch_bytes: 32KB # Optional (default: "32KB") auto_replay_nacks: true # Optional (default: true) pipeline: processors: - bloblang: | root = this let rating = random_int(min: 1, max: 11) root.cringe_rating = $rating root.cringe_level = if $rating <= 3 { "Mild - Almost acceptable" } else if $rating <= 6 { "Medium - Classic dad joke territory" } else if $rating <= 8 { "High - Eye-roll inducing" } else { "EXTREME - Peak dad joke achievement" } root.processed_at = now() root.rating_emoji = match { $rating <= 3 => "😐", $rating <= 6 => "😬", $rating <= 8 => "🤦", _ => "💀" } let age_seconds = (timestamp_unix() - this.timestamp.ts_parse("2006-01-02T15:04:05Z07:00").ts_unix()) root.age_seconds = $age_seconds - log: level: INFO message: | 🎭 JOKE RATED! ${! json("rating_emoji") } Joke: "${! json("joke") }" Cringe Rating: ${! json("cringe_rating") }/10 - ${! json("cringe_level") } Age: ${! json("age_seconds") } seconds old Processed at: ${! json("processed_at") } - catch: - log: level: ERROR message: "❌ Failed to process joke: ${! error() }" output: drop: {} ``` 7. Click **Save** to start your pipeline. 8. Your pipeline details display, and after a few seconds, the pipeline starts running. Check the logs to see a rated joke. For example: ```json { "custom_source": "true", "instance_id": "d454dkn4u2is73ava480", "label": "", "level": "INFO", "message": "🎭 JOKE RATED! 💀\nJoke: \"I used to play piano by ear, but now I use my hands.\"\nCringe Rating: 9/10 - EXTREME - Peak dad joke achievement\nAge: 659 seconds old\nProcessed at: 2026-03-27T17:54:13.340229297Z\n", "path": "root.pipeline.processors.1", "pipeline_id": "d454djahlips73dmcll0", "time": "2026-03-27T17:54:13.341137527Z" } ``` ## [](#enable-egress-for-pipelines-on-byoc-and-dedicated-clusters)Enable egress for pipelines on BYOC and Dedicated clusters The consumer pipeline in this quickstart drops messages instead of writing them to a real destination. When you replace the `drop` output with a connector that writes to your own systems, pipelines on BYOC and Dedicated clusters may need additional egress access. On BYOC and Dedicated clusters, pipelines run inside the Redpanda data plane VPC. By default, pipelines can connect to your Redpanda cluster and to publicly routable endpoints. Outbound connections to other destinations, such as a database with a private address in a peered VPC, are blocked. To allow pipelines to reach these destinations, add an egress allowlist to your cluster using the [Redpanda Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) (version `>= 2.1.1`). The `redpanda_connect.allowed_destination_cidr_ports` attribute on the `redpanda_cluster` resource accepts up to 16 rules, each allowing egress to a destination CIDR block on a port or port range: ```hcl resource "redpanda_cluster" "example" { # ... other cluster arguments ... redpanda_connect = { allowed_destination_cidr_ports = [ { cidr = "10.62.0.0/16" # CIDR of the VPC that hosts your database port_start = 5432 # PostgreSQL } ] } } ``` The allowlist permits outbound traffic from pipelines, but it does not create a network path to the destination. For private destinations, you must also establish routing, for example by peering the Redpanda data plane VPC with the VPC that hosts your database. To set the allowlist with the Cloud API, and for the full rule reference and troubleshooting guidance, see [Configure Egress for Redpanda Connect Pipelines](https://docs.redpanda.com/cloud-data-platform/networking/connect-egress-allowlist/). For a pipeline example that writes to a private PostgreSQL database, see [Enable egress to custom destinations](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/#enable-egress-to-custom-destinations). ## [](#clean-up)Clean up When you’ve finished experimenting with your data pipeline, you can delete the pipelines and the topic you created for this quickstart. 1. On the **Connect** page, click the **…​** icon next to the `joke-generator-producer` pipeline and select **Delete**. Repeat for the `joke-generator-consumer` pipeline. 2. Confirm your deletion to remove the pipelines and associated logs. 3. On the **Topics** page, delete the `dad-jokes` topic. ## [](#next-steps)Next steps - Try one of the [Redpanda Connect cookbooks](https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/). - Choose [connectors for your use case](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/about/). - [Add secrets to your pipeline](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/). - [Monitor a data pipeline on a BYOC or Dedicated cluster](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/monitor-connect/). - [Manually scale resources for a pipeline](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/resource-management/). - [Configure, test, and run a data pipeline locally](https://docs.redpanda.com/connect/get-started/quickstarts/rpk/). --- # Page 289: Cookbooks **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks.md --- # Cookbooks > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Cookbooks latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/index.adoc description: Step-by-step Redpanda Connect cookbooks that walk through complete, real-world pipeline builds in Redpanda Cloud. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- - [DynamoDB CDC Patterns](dynamodb_cdc/) Learn how to capture, filter, transform, and route DynamoDB change data capture (CDC) events with Redpanda Connect. - [Enrichment Workflows](enrichments/) How to configure Redpanda Connect to process a workflow of enrichment services. - [Filtering and Sampling](filtering/) Configure Redpanda Connect to conditionally drop messages. - [Ingest data into Snowflake](snowflake_ingestion/) Configure Redpanda Connect to ingest data from a Redpanda topic into Snowflake using Snowpipe Streaming. - [Joining Streams](joining_streams/) How to hydrate documents by joining multiple streams. - [Redpanda Migrator](redpanda_migrator/) Move your workloads from any Kafka system to Redpanda Cloud using a single command. - [Retrieval-Augmented Generation (RAG)](rag/) How to configure Redpanda Connect to create a RAG pipeline, using PostgreSQL and PGVector. - [Work with Jira Issues](jira/) Learn how to query, filter, and create Jira issues using Redpanda Connect pipelines. --- # Page 290: DynamoDB CDC Patterns **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/dynamodb_cdc.md --- # DynamoDB CDC Patterns > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: DynamoDB CDC Patterns latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/dynamodb_cdc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/dynamodb_cdc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/dynamodb_cdc.adoc description: Learn how to capture, filter, transform, and route DynamoDB change data capture (CDC) events with Redpanda Connect. page-topic-type: cookbook personas: streaming_developer, data_engineer learning-objective-1: Apply reusable patterns for capturing DynamoDB CDC events learning-objective-2: Adapt integration patterns to route CDC data to Redpanda and S3 learning-objective-3: Identify patterns for filtering and transforming change events page-git-created-date: "2026-03-04" page-git-modified-date: "2026-08-11" --- The DynamoDB CDC input enables capturing item-level changes from DynamoDB tables with streams enabled. This cookbook provides reusable patterns for filtering, transforming, and routing DynamoDB CDC events to Redpanda, S3, and other destinations. Use this cookbook to: - Apply reusable patterns for capturing DynamoDB CDC events - Adapt integration patterns to route CDC data to Redpanda and S3 - Identify patterns for filtering and transforming change events ## [](#prerequisites)Prerequisites Before using these patterns, configure the following. ### [](#redpanda-cli)Redpanda CLI Install the Redpanda CLI (`rpk`) to run Redpanda Connect. See [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) for installation instructions. ### [](#dynamodb-streams)DynamoDB Streams The source DynamoDB table must have [DynamoDB Streams](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html) enabled with an appropriate view type: - `KEYS_ONLY`: Only the key attributes of the modified item - `NEW_IMAGE`: The entire item as it appears after the modification - `OLD_IMAGE`: The entire item as it appeared before the modification - `NEW_AND_OLD_IMAGES`: Both the new and old item images (recommended for detecting changes) To enable streams on an existing table using the AWS CLI: ```bash aws dynamodb update-table \ --table-name orders \ --stream-specification StreamEnabled=true,StreamViewType=NEW_AND_OLD_IMAGES ``` ### [](#environment-variables)Environment variables The examples in this cookbook use environment variables for AWS configuration. Environment variables keep credentials out of your pipeline configuration files. ```bash export DYNAMODB_TABLE=orders (1) export AWS_REGION=us-east-1 (2) export REDPANDA_BROKERS=localhost:9092 (3) export S3_BUCKET=cdc-archive (4) ``` | 1 | The name of the DynamoDB table with streams enabled. | | --- | --- | | 2 | The AWS region where your DynamoDB table is located. | | 3 | The Redpanda broker addresses (for Redpanda output examples). | | 4 | The S3 bucket name (for S3 output examples). | Redpanda Connect loads AWS credentials from the standard [credential chain](https://docs.aws.amazon.com/cli/latest/userguide/cli-configure-files.html) (environment variables, `~/.aws/credentials`, or IAM roles). ## [](#capture-cdc-events)Capture CDC events The simplest pattern captures all change events from a DynamoDB table and outputs them with metadata: ```yaml input: aws_dynamodb_cdc: tables: ["${DYNAMODB_TABLE}"] region: ${AWS_REGION} checkpoint_table: redpanda_dynamodb_checkpoints start_from: trim_horizon pipeline: processors: # Extract the change event details - mapping: | root.event_type = this.eventName root.table = this.tableName root.event_id = this.eventID root.keys = this.dynamodb.keys root.new_image = this.dynamodb.newImage root.old_image = this.dynamodb.oldImage root.sequence_number = this.dynamodb.sequenceNumber root.timestamp = now() output: stdout: codec: lines ``` For details on the CDC event message structure and available fields for Bloblang mappings, see the [message structure](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_dynamodb_cdc/#_message_structure) section in the connector reference. ## [](#filter-cdc-events)Filter CDC events Filter events to process only specific change types: ```yaml input: aws_dynamodb_cdc: tables: ["${DYNAMODB_TABLE}"] region: ${AWS_REGION} start_from: latest pipeline: processors: # Filter to only process INSERT and MODIFY events (ignore REMOVE) - mapping: | root = if this.eventName == "REMOVE" { deleted() } else { this } # Transform to a simplified format - mapping: | root.event_type = this.eventName root.keys = this.dynamodb.keys root.new_data = this.dynamodb.newImage root.old_data = this.dynamodb.oldImage output: stdout: codec: lines ``` This example: - Filters out `REMOVE` events using `deleted()` - Transforms the event to a simplified format ## [](#route-to-redpanda)Route to Redpanda Stream DynamoDB changes to Redpanda for real-time processing: ```yaml input: aws_dynamodb_cdc: tables: ["${DYNAMODB_TABLE}"] region: ${AWS_REGION} checkpoint_table: redpanda_dynamodb_checkpoints batch_size: 100 poll_interval: 500ms pipeline: processors: # Transform to a Kafka-friendly format with a composite key - mapping: | let keys = this.dynamodb.keys meta kafka_key = [$keys.pk, $keys.sk].filter(v -> v != null).join("#") root.event_type = this.eventName root.table = this.tableName root.timestamp = now() root.keys = this.dynamodb.keys root.new_image = this.dynamodb.newImage root.old_image = this.dynamodb.oldImage output: redpanda: seed_brokers: - ${REDPANDA_BROKERS} topic: dynamodb-cdc-events key: ${! @kafka_key } partitioner: murmur2_hash compression: snappy batching: count: 100 period: 1s ``` This example: - Creates a composite message key from the DynamoDB primary key - Transforms the DynamoDB format to plain JSON - Batches messages for efficient delivery ## [](#route-to-s3)Route to S3 Archive CDC events to S3 for long-term storage and analytics: ```yaml input: aws_dynamodb_cdc: tables: ["${DYNAMODB_TABLE}"] region: ${AWS_REGION} checkpoint_table: redpanda_dynamodb_checkpoints start_from: trim_horizon pipeline: processors: # Add partitioning metadata for S3 organization - mapping: | let event_time = now() meta s3_path = "year=%s/month=%s/day=%s/hour=%s".format( $event_time.ts_format("2006"), $event_time.ts_format("01"), $event_time.ts_format("02"), $event_time.ts_format("15") ) root.event_type = this.eventName root.table = this.tableName root.sequence_number = this.dynamodb.sequenceNumber root.event_time = $event_time root.keys = this.dynamodb.keys root.new_image = this.dynamodb.newImage root.old_image = this.dynamodb.oldImage output: aws_s3: bucket: ${S3_BUCKET} path: dynamodb-cdc/${DYNAMODB_TABLE}/${! @s3_path }/${! uuid_v4() }.json region: ${AWS_REGION} object_canned_acl: private batching: count: 1000 period: 1m processors: - archive: format: lines ``` This example: - Organizes files by time-based partitions (year/month/day/hour) - Batches events and archives them as newline-delimited JSON - Uses UUID file names to prevent collisions ## [](#route-by-event-type)Route by event type Route different event types to different destinations: ```yaml input: aws_dynamodb_cdc: tables: ["${DYNAMODB_TABLE}"] region: ${AWS_REGION} pipeline: processors: # Transform to a common format - mapping: | root.event_type = this.eventName root.table = this.tableName root.timestamp = now() root.keys = this.dynamodb.keys root.data = if this.dynamodb.exists("newImage") { this.dynamodb.newImage } else { this.dynamodb.oldImage } output: switch: cases: # Route INSERT events to a topic for new records - check: this.event_type == "INSERT" output: redpanda: seed_brokers: - ${REDPANDA_BROKERS} topic: dynamodb-inserts # Route MODIFY events to a topic for updates - check: this.event_type == "MODIFY" output: redpanda: seed_brokers: - ${REDPANDA_BROKERS} topic: dynamodb-updates # Route REMOVE events to a topic for deletes - check: this.event_type == "REMOVE" output: redpanda: seed_brokers: - ${REDPANDA_BROKERS} topic: dynamodb-deletes # Fallback for any unexpected event types - output: drop: {} ``` This pattern: - Separates processing pipelines for inserts, updates, and deletes - Applies different retention policies per event type - Supports specialized downstream consumers ## [](#detect-changed-fields)Detect changed fields Compare old and new images to identify which fields changed: ```yaml input: aws_dynamodb_cdc: tables: ["${DYNAMODB_TABLE}"] region: ${AWS_REGION} pipeline: processors: # Only process MODIFY events - mapping: | root = if this.eventName != "MODIFY" { deleted() } else { this } # Compare old and new images to find changed fields - mapping: | let old_data = this.dynamodb.oldImage let new_data = this.dynamodb.newImage root.table = this.tableName root.keys = this.dynamodb.keys root.timestamp = now() # Find fields that changed by comparing key-value pairs root.changes = $new_data.key_values().filter(kv -> !$old_data.exists(kv.key) || $old_data.get(kv.key) != kv.value).map_each(kv -> {"field": kv.key, "old_value": if $old_data.exists(kv.key) { $old_data.get(kv.key) } else { null }, "new_value": kv.value}) # Find fields that were removed root.removed_fields = $old_data.keys().filter(k -> !$new_data.exists(k)) output: stdout: codec: lines ``` This pattern: - Filters to only MODIFY events - Compares old and new images to find differences - Outputs a list of changed fields with their old and new values > 📝 **NOTE** > > This pattern requires the `NEW_AND_OLD_IMAGES` stream view type. The `.key_values()` method converts an object to an array of key-value pairs that can be filtered and mapped. ## [](#checkpointing)Checkpointing The DynamoDB CDC input automatically manages checkpoints in a separate DynamoDB table: ```yaml input: aws_dynamodb_cdc: tables: - orders checkpoint_table: cdc-checkpoints (1) checkpoint_limit: 500 (2) start_from: trim_horizon (3) ``` | 1 | Custom checkpoint table name (default: redpanda_dynamodb_checkpoints). | | --- | --- | | 2 | Checkpoint after every 500 messages (lower = better recovery, higher = fewer writes). | | 3 | Start from the oldest available record when no checkpoint exists. | If a checkpoint table doesn’t exist, it’s created automatically with the required schema. ## [](#performance-tuning)Performance tuning Optimize throughput and latency with these settings: ```yaml input: aws_dynamodb_cdc: tables: - orders batch_size: 1000 (1) poll_interval: 100ms (2) max_tracked_shards: 10000 (3) throttle_backoff: 50ms (4) ``` | 1 | Maximum records per shard per request (1-1000). | | --- | --- | | 2 | Time between polls when no records are available. | | 3 | Maximum shards to track (for very large tables). | | 4 | Backpressure delay when too many messages are in-flight. | ### [](#throughput-considerations)Throughput considerations - DynamoDB Streams allows 5 `GetRecords` calls per second per shard - Higher `batch_size` improves throughput but increases memory usage - Shorter `poll_interval` reduces latency but increases API calls ## [](#troubleshoot-common-issues)Troubleshoot common issues ### [](#no-events-received)No events received If you’re not receiving events: 1. Verify streams are enabled on the table: ```bash aws dynamodb describe-table --table-name orders \ --query 'Table.StreamSpecification' ``` 2. Check that changes are being made to the table 3. Verify `start_from` is set to `trim_horizon` to capture existing stream data ### [](#duplicate-events)Duplicate events Each stream record appears exactly once in DynamoDB Streams. However, if your pipeline fails before checkpointing, records may be re-read on restart, resulting in at-least-once processing semantics. To handle potential duplicates: - Use idempotent processing in downstream systems - Deduplicate using the `dynamodb_sequence_number` metadata - Lower `checkpoint_limit` to reduce the window of possible duplicates ### [](#stream-retention)Stream retention DynamoDB Streams retains data for 24 hours. If your pipeline is offline longer than that: - Consider using [Kinesis Data Streams for DynamoDB](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/kds.html) with the [`aws_kinesis` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_kinesis/) instead (up to 1 year retention) - Implement a full-table scan fallback for disaster recovery ## [](#next-steps)Next steps - [DynamoDB CDC Input Reference](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_dynamodb_cdc/) - [AWS Configuration Guide](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws/) - [Kinesis Input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_kinesis/) (for Kinesis Data Streams for DynamoDB) - [DynamoDB Streams Documentation](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/Streams.html) --- # Page 291: Enrichment Workflows **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/enrichments.md --- # Enrichment Workflows > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Enrichment Workflows latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/enrichments page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/enrichments.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/enrichments.adoc description: How to configure Redpanda Connect to process a workflow of enrichment services. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- This cookbook demonstrates how to enrich a stream of JSON documents with HTTP services. This method also works with [AWS Lambda functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/aws_lambda/). We will start off by configuring a single enrichment, then we will move onto a workflow of enrichments with a network of dependencies using the [`workflow` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/workflow/). Each enrichment will be performed in parallel across a [pre-batched](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/batching/) stream of documents. Workflow enrichments that do not depend on each other will also be performed in parallel, making this orchestration method very efficient. The imaginary problem we are going to solve is applying a set of NLP based enrichments to a feed of articles in order to detect fake news. We will be consuming and writing to Kafka, but the example works with any [input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/about/) and [output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about/) combination. Articles are received over the topic `articles` and look like this: ```json { "type": "article", "article": { "id": "123foo", "title": "Dogs Stop Barking", "content": "The world was shocked this morning to find that all dogs have stopped barking." } } ``` ## [](#meet-the-enrichments)Meet the enrichments ### [](#claims-detector)Claims detector To start us off we will configure a single enrichment, which is an imaginary 'claims detector' service. This is an HTTP service that wraps a trained machine learning model to extract claims that are made within a body of text. The service expects a `POST` request with JSON payload of the form: ```json { "text": "The world was shocked this morning to find that all dogs have stopped barking." } ``` And returns a JSON payload of the form: ```json { "claims": [ { "entity": "world", "claim": "shocked" }, { "entity": "dogs", "claim": "NOT barking" } ] } ``` Since each request only applies to a single document we will make this enrichment scale by deploying multiple HTTP services and hitting those instances in parallel across our document batches. In order to send a mapped request and map the response back into the original document we will use the [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/), with a child `http` processor. ```yaml input: kafka: addresses: [ TODO ] topics: [ articles ] consumer_group: benthos_articles_group batching: count: 20 # Tune this to set the size of our document batches. period: 1s pipeline: processors: - branch: request_map: 'root.text = this.article.content' processors: - http: url: http://localhost:4197/claims verb: POST result_map: 'root.tmp.claims = this.claims' output: kafka: addresses: [ TODO ] topic: comments_hydrated ``` With this pipeline our documents will come out looking something like this: ```json { "type": "article", "article": { "id": "123foo", "title": "Dogs Stop Barking", "content": "The world was shocked this morning to find that all dogs have stopped barking." }, "tmp": { "claims": [ { "entity": "world", "claim": "shocked" }, { "entity": "dogs", "claim": "NOT barking" } ] } } ``` ### [](#hyperbole-detector)Hyperbole detector Next up is a 'hyperbole detector' that takes a `POST` request containing the article contents and returns a hyperbole score between 0 and 1. This time the format is array-based and therefore supports calculating multiple documents in a single request, making better use of the host machines GPU. A request should take the following form: ```json [ { "text": "The world was shocked this morning to find that all dogs have stopped barking." } ] ``` And the response looks like this: ```json [ { "hyperbole_rank": 0.73 } ] ``` In order to create a single request from a batch of documents, and subsequently map the result back into our batch, we will use the [`archive`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/archive/) and [`unarchive`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/unarchive/) processors in our [`branch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) flow, like this: ```yaml pipeline: processors: - branch: request_map: 'root.text = this.article.content' processors: - archive: format: json_array - http: url: http://localhost:4198/hyperbole verb: POST - unarchive: format: json_array result_map: 'root.tmp.hyperbole_rank = this.hyperbole_rank' ``` The purpose of the `json_array` format `archive` processor is to take a batch of JSON documents and place them into a single document as an array. Subsequently, we then send one single request for each batch. After the request is made we do the opposite with the `unarchive` processor in order to convert it back into a batch of the original size. ### [](#fake-news-detector)Fake news detector Finally, we are going to use a 'fake news detector' that takes the article contents as well as the output of the previous two enrichments and calculates a fake news rank between 0 and 1. This service behaves similarly to the claims detector service and takes a document of the form: ```json { "text": "The world was shocked this morning to find that all dogs have stopped barking.", "hyperbole_rank": 0.73, "claims": [ { "entity": "world", "claim": "shocked" }, { "entity": "dogs", "claim": "NOT barking" } ] } ``` And returns an object of the form: ```json { "fake_news_rank": 0.893 } ``` We then wish to map the field `fake_news_rank` from that result into the original document at the path `article.fake_news_score`. Our [`branch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) block for this enrichment would look like this: ```yaml pipeline: processors: - branch: request_map: | root.text = this.article.content root.claims = this.tmp.claims root.hyperbole_rank = this.tmp.hyperbole_rank processors: - http: url: http://localhost:4199/fakenews verb: POST result_map: 'root.article.fake_news_score = this.fake_news_rank' ``` Note that in our `request_map` we are targeting fields that are populated from the previous two enrichments. If we were to execute all three enrichments in a sequence we’ll end up with a document looking like this: ```json { "type": "article", "article": { "id": "123foo", "title": "Dogs Stop Barking", "content": "The world was shocked this morning to find that all dogs have stopped barking.", "fake_news_score": 0.76 }, "tmp": { "hyperbole_rank": 0.34, "claims": [ { "entity": "world", "claim": "shocked" }, { "entity": "dogs", "claim": "NOT barking" } ] } } ``` Great! However, as a streaming pipeline this set up isn’t ideal as our first two enrichments are independent and could potentially be executed in parallel in order to reduce processing latency. ## [](#combining-into-a-workflow)Combining into a workflow If we configure our enrichments within a [`workflow` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/workflow/) we can use Redpanda Connect to automatically detect our dependency graph, giving us two key benefits: 1. Enrichments at the same level of a dependency graph (claims and hyperbole) will be executed in parallel. 2. When introducing more enrichments to our pipeline the added complexity of resolving the dependency graph is handled automatically by Redpanda Connect. Placing our branches within a [`workflow` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/workflow/) makes our final pipeline configuration look like this: ```yaml input: kafka: addresses: [ TODO ] topics: [ articles ] consumer_group: benthos_articles_group batching: count: 20 # Tune this to set the size of our document batches. period: 1s pipeline: processors: - workflow: meta_path: '' # Don't bother storing branch metadata. branches: claims: request_map: 'root.text = this.article.content' processors: - http: url: http://localhost:4197/claims verb: POST result_map: 'root.tmp.claims = this.claims' hyperbole: request_map: 'root.text = this.article.content' processors: - archive: format: json_array - http: url: http://localhost:4198/hyperbole verb: POST - unarchive: format: json_array result_map: 'root.tmp.hyperbole_rank = this.hyperbole_rank' fake_news: request_map: | root.text = this.article.content root.claims = this.tmp.claims root.hyperbole_rank = this.tmp.hyperbole_rank processors: - http: url: http://localhost:4199/fakenews verb: POST result_map: 'root.article.fake_news_score = this.fake_news_rank' - catch: - log: fields_mapping: 'root.content = content().string()' message: "Enrichments failed due to: ${!error()}" - mapping: | root = this root.tmp = deleted() output: kafka: addresses: [ TODO ] topic: comments_hydrated ``` Since the contents of `tmp` won’t be required downstream we remove it after our enrichments using a [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/). A [`catch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/catch/) processor was added at the end of the pipeline which catches documents that failed enrichment. You can replace the log event with a wide range of recovery actions such as sending to a dead-letter/retry queue, dropping the message entirely, etc. You can read more about error handling [in this article](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/). --- # Page 292: Filtering and Sampling **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/filtering.md --- # Filtering and Sampling > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Filtering and Sampling latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/filtering page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/filtering.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/filtering.adoc description: Configure Redpanda Connect to conditionally drop messages. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Filtering events in Redpanda Connect is both easy and flexible, this cookbook demonstrates a few different types of filtering you can do. All of these examples make use of the [`mapping` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/) but shouldn’t require any prior knowledge. ## [](#the-basic-filter)The basic filter Dropping events with [Bloblang](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/) is done by mapping the function `deleted()` to the `root` of the mapped document. To remove all events indiscriminately you can simply do: ```yaml pipeline: processors: - mapping: root = deleted() ``` But that’s most likely not what you want. We can instead only delete an event under certain conditions with a [`match`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#pattern-matching) or [`if`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#conditional-mapping) expression: ```yaml pipeline: processors: - mapping: | root = if @topic.or("") == "foo" || this.doc.type == "bar" || this.doc.urls.contains("https://www.benthos.dev/").catch(false) { deleted() } ``` The above config removes any events where: - The metadata field `topic` is equal to `foo` - The event field `doc.type` (a string) is equal to `bar` - The event field `doc.urls` (an array) contains the string `https://www.benthos.dev/` Events that do not match any of these conditions will remain unchanged. ## [](#sample-events)Sample events Another type of filter we might want is a sampling filter, we can do that with a random number generator: ```yaml pipeline: processors: - mapping: | # Drop 50% of documents randomly root = if random_int() % 2 == 0 { deleted() } ``` We can also do this in a deterministic way by hashing events and filtering by that hash value: ```yaml pipeline: processors: - mapping: | # Drop ~10% of documents deterministically (same docs filtered each run) root = if content().hash("xxhash64").slice(-8).number() % 10 == 0 { deleted() } ``` --- # Page 293: Work with Jira Issues **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/jira.md --- # Work with Jira Issues > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Work with Jira Issues latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/jira page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/jira.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/jira.adoc description: Learn how to query, filter, and create Jira issues using Redpanda Connect pipelines. page-topic-type: cookbook personas: streaming_developer, data_engineer learning-objective-1: Query Jira issues using JQL patterns with the Jira processor learning-objective-2: Combine generate input with Jira processor for scheduled queries learning-objective-3: Create Jira issues using the HTTP processor and REST API page-git-created-date: "2026-02-18" page-git-modified-date: "2026-02-18" --- The Jira processor enables querying Jira issues using JQL (Jira Query Language) and returning structured data. It’s a processor, so you can use it in pipelines for input-style flows (pair with `generate`) or output-style flows (pair with `drop`). Use this cookbook to: - Query Jira issues on a schedule or on-demand - Filter issues using JQL patterns - Create Jira issues using the HTTP processor ## [](#prerequisites)Prerequisites The examples in this cookbook use the Secrets Store for Jira credentials. This keeps sensitive credentials secure and separate from your pipeline configuration. 1. [Generate a Jira API token](https://id.atlassian.com/manage-profile/security/api-tokens). 2. Add your Jira credentials to the [Secrets Store](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/): - `JIRA_BASE_URL`: Your Jira instance URL (for example, `https://your-domain.atlassian.net`) - `JIRA_USERNAME`: Your Jira account email address - `JIRA_API_TOKEN`: The API token generated from your Atlassian account - `JIRA_AUTH_TOKEN` (optional, for creating issues): Base64-encoded `username:api_token` string ## [](#use-jira-as-an-input)Use Jira as an input To use Jira as an input, combine the `generate` input with the Jira processor. This pattern triggers Jira queries at regular intervals or on-demand. > 💡 **TIP** > > Replace `MYPROJECT` in the examples with your actual Jira project key. ### [](#query-jira-periodically)Query Jira periodically This example queries Jira every 30 seconds for recent issues: ```yaml input: generate: interval: 30s mapping: | root.jql = "project = MYPROJECT AND updated >= -1h ORDER BY updated DESC" root.maxResults = 50 root.fields = ["key", "summary", "status", "assignee", "priority"] pipeline: processors: - jira: base_url: "${secrets.JIRA_BASE_URL}" username: "${secrets.JIRA_USERNAME}" api_token: "${secrets.JIRA_API_TOKEN}" output: stdout: {} ``` ### [](#one-time-query)One-time query For a single query, use `count` instead of `interval`: ```yaml input: generate: count: 1 mapping: | root.jql = "project = MYPROJECT AND status = Open" root.maxResults = 100 pipeline: processors: - jira: base_url: "${secrets.JIRA_BASE_URL}" username: "${secrets.JIRA_USERNAME}" api_token: "${secrets.JIRA_API_TOKEN}" output: stdout: {} ``` ## [](#input-message-format)Input message format The Jira processor expects input messages containing valid Jira queries in JSON format: ```json { "jql": "project = MYPROJECT AND status = Open", "maxResults": 50, "fields": ["key", "summary", "status", "assignee"] } ``` ### [](#required-fields)Required fields - `jql`: The JQL (Jira Query Language) query string ### [](#optional-fields)Optional fields - `maxResults`: Maximum number of results to return (default: 50) - `fields`: Array of field names to include in the response ## [](#jql-query-patterns)JQL query patterns Here are common JQL patterns for filtering issues: ### [](#recent-issues-by-project)Recent issues by project ```jql project = AND created >= -7d ORDER BY created DESC ``` ### [](#issues-assigned-to-current-user)Issues assigned to current user ```jql assignee = currentUser() AND status != Done ``` ### [](#issues-by-status)Issues by status ```jql project = AND status IN (Open, 'In Progress', 'To Do') ``` ### [](#issues-by-priority)Issues by priority ```jql project = AND priority = High ORDER BY created DESC ``` ## [](#output-message-format)Output message format The Jira processor returns individual issue messages, rather than a response object with an `issues` array. Each message output by the Jira processor represents a single issue: ```json { "id": "12345", "key": "DOC-123", "fields": { "summary": "Example issue", "status": { "name": "In Progress" }, "assignee": { "displayName": "John Doe" } } } ``` The Jira processor automatically handles pagination internally. The processor: 1. Makes the initial request with `startAt=0`. 2. Checks if more results are available. 3. Automatically fetches subsequent pages until all results are retrieved. 4. Outputs each issue as an individual message. You don’t need to handle pagination manually. ## [](#create-and-update-jira-issues)Create and update Jira issues The Jira processor is read-only and only supports querying. To create or update Jira issues, use the [`http` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/http/) with the Jira REST API. ### [](#create-a-jira-issue)Create a Jira issue ```yaml input: generate: count: 1 mapping: | root.fields = { "project": {"key": "MYPROJECT"}, "summary": "Issue created from Redpanda Connect", "description": { "type": "doc", "version": 1, "content": [{"type": "paragraph", "content": [{"type": "text", "text": "Created via API"}]}] }, "issuetype": {"name": "Task"} } pipeline: processors: - http: url: "${secrets.JIRA_BASE_URL}/rest/api/3/issue" verb: POST headers: Content-Type: application/json Authorization: "Basic ${secrets.JIRA_AUTH_TOKEN}" output: stdout: {} ``` ## [](#see-also)See also - [Jira processor reference](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/jira/) - [Jira REST API documentation](https://developer.atlassian.com/cloud/jira/platform/rest/v3/intro/) - [JQL query guide](https://www.atlassian.com/software/jira/guides/jql) --- # Page 294: Joining Streams **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/joining_streams.md --- # Joining Streams > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Joining Streams latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/joining_streams page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/joining_streams.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/joining_streams.adoc description: How to hydrate documents by joining multiple streams. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- This cookbook demonstrates how to merge JSON events from parallel streams using content based rules and a [cache](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/) of your choice. The imaginary problem we are going to solve is hydrating a feed of article comments with information from their parent articles. We will be consuming and writing to Kafka, but the example works with any [input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/about/) and [output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/about/) combination. Articles are received over the topic `articles` and look like this: ```json { "type": "article", "article": { "id": "123foo", "title": "Good article", "content": "this is a totally good article" }, "user": { "id": "user1" } } ``` Comments can either be posted on an article or a parent comment, are received over the topic `comments`, and look like this: ```json { "type": "comment", "comment": { "id": "456bar", "parent_id": "123foo", "content": "this article is bad" }, "user": { "id": "user2" } } ``` Our goal is to end up with a single stream of comments, where information about the root article of the comment is attached to the event. The above comment should exit our pipeline looking like this: ```json { "type": "comment", "comment": { "id": "456bar", "parent_id": "123foo", "content": "this article is bad" }, "article": { "title": "Good article", "content": "this is a totally good article" }, "user": { "id": "user2" } } ``` In order to achieve this we will need to cache articles as they pass through our pipelines and then retrieve them for each comment passing through. Since the parent of a comment might be another comment we will also need to cache and retrieve comments in the same way. ## [](#caching-articles)Caching articles Our first pipeline is very simple, we just consume articles, reduce them to only the fields we wish to cache, and then cache them. If we receive the same article multiple times we’re going to assume it’s okay to overwrite the old article in the cache. In this example I’m targeting Redis, but you can choose any of the supported [cache targets](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/about/). The TTL of cached articles is set to one week. ```yaml input: kafka: addresses: [ TODO ] topics: [ articles ] consumer_group: benthos_articles_group pipeline: processors: # Reduce document into only fields we wish to cache. - mapping: 'article = article' # Store reduced articles into our cache. - cache: operator: set resource: hydration_cache key: '${!json("article.id")}' value: '${!content()}' # Drop all articles after they are cached. output: drop: {} cache_resources: - label: hydration_cache redis: url: TODO default_ttl: 168h ``` ## [](#hydrating-comments)Hydrating comments Our second pipeline consumes comments, caches them in case a subsequent comment references them, obtains its parent (article or comment), and attaches the root article to the event before sending it to our output topic `comments_hydrated`. In this config we make use of the [`branch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) processor as it allows us to reduce documents into smaller maps for caching and gives us greater control over how results are mapped back into the document. ```yaml input: kafka: addresses: [ TODO ] topics: [ comments ] consumer_group: benthos_comments_group pipeline: processors: # Perform both hydration and caching within a for_each block as this ensures # that a given message of a batch is cached before the next message is # hydrated, ensuring that when a message of the batch has a parent within # the same batch hydration can still work. - for_each: # Attempt to obtain parent event from cache (if the ID exists). - branch: request_map: 'root = this.comment.parent_id | deleted()' processors: - cache: operator: get resource: hydration_cache key: '${!content()}' # And if successful copy it into the field `article`. result_map: 'root.article = this.article' # Reduce comment into only fields we wish to cache. - branch: request_map: | root.comment.id = this.comment.id root.article = this.article processors: # Store reduced comment into our cache. - cache: operator: set resource: hydration_cache key: '${!json("comment.id")}' value: '${!content()}' # No `result_map` since we don't need to map into the original message. # Send resulting documents to our hydrated topic. output: kafka: addresses: [ TODO ] topic: comments_hydrated cache_resources: - label: hydration_cache redis: url: TODO default_ttl: 168h ``` This pipeline satisfies our basic needs but errors aren’t handled at all, meaning intermittent cache connectivity problems that span beyond our cache retries will result in failed documents entering our `comments_hydrated` topic. This is also the case if a comment arrives in our pipeline before its parent. There are [many patterns for error handling](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/) to choose from in Redpanda Connect. In this example we’re going to introduce a delayed retry queue as it enables us to reprocess failed documents after a grace period, which is isolated from our main pipeline. ## [](#adding-a-retry-queue)Adding a retry queue Our retry queue is going to be another topic called `comments_retried`. Since most errors are related to time we will delay retry attempts by storing the current timestamp after a failed request as a metadata field. We will use an input [`broker`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/broker/) so that we can consume both the `comments` and `comments_retry` topics in the same pipeline. Our config (omitting the caching sections for brevity) now looks like this: ```yaml input: broker: inputs: - kafka: addresses: [ TODO ] topics: [ comments ] consumer_group: benthos_comments_group - kafka: addresses: [ TODO ] topics: [ comments_retry ] consumer_group: benthos_comments_group processors: - for_each: # Calculate time until next retry attempt and sleep for that duration. # This sleep blocks the topic 'comments_retry' but NOT 'comments', # because both topics are consumed independently and these processors # only apply to the 'comments_retry' input. - sleep: duration: '${! 3600 - ( timestamp_unix() - meta("last_attempted").number() ) }s' pipeline: processors: - try: - for_each: # Attempt to obtain parent event from cache. - branch: {} # Omitted # Reduce document into only fields we wish to cache. - branch: {} # Omitted # If we've reached this point then both processors succeeded. - mapping: 'meta output_topic = "comments_hydrated"' - catch: # If we reach here then a processing stage failed. - mapping: | meta output_topic = "comments_retry" meta last_attempted = timestamp_unix() # Send resulting documents either to our hydrated topic or the retry topic. output: kafka: addresses: [ TODO ] topic: '${!meta("output_topic")}' cache_resources: - label: hydration_cache redis: url: TODO default_ttl: 168h ``` You can find a full example [in the project repo](https://github.com/redpanda-data/connect/blob/master/config/examples/joining_streams.yaml), and with this config we can deploy as many instances of Redpanda Connect as we need as the partitions will be balanced across the consumers. --- # Page 295: Retrieval-Augmented Generation (RAG) **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/rag.md --- # Retrieval-Augmented Generation (RAG) > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Retrieval-Augmented Generation (RAG) latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/rag page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/rag.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/rag.adoc description: How to configure Redpanda Connect to create a RAG pipeline, using PostgreSQL and PGVector. page-git-created-date: "2024-09-12" page-git-modified-date: "2026-05-26" --- This cookbook shows you how to create a vector embeddings indexing pipeline for Retrieval-Augmented Generation (RAG), using PostgreSQL and [PGVector](https://github.com/pgvector/pgvector). Follow the cookbook to: - Take textual data from a Redpanda topic and compute vector embeddings for it using [Ollama](https://ollama.ai) - Write the pipeline output into a PostgreSQL table with a [PGVector](https://github.com/pgvector/pgvector) index on the embeddings column. ## [](#compute-the-embeddings)Compute the embeddings Start by creating a Redpanda topic, which you can use as an input for an indexing data pipeline. ```bash rpk topic create articles echo '{ "type": "article", "article": { "id": "123foo", "title": "Dogs Stop Barking", "content": "The world was shocked this morning to find that all dogs have stopped barking." } }' | rpk topic produce articles -f '%v' ``` Your indexing pipeline can read from the Redpanda topic, using the [`kafka`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/kafka/) input: ```yaml input: kafka: addresses: [ "TODO" ] topics: [ articles ] consumer_group: rp_connect_articles_group tls: enabled: true sasl: mechanism: SCRAM-SHA-256 user: "TODO" password: "TODO" ``` Use [Nomic Embed](https://ollama.com/library/nomic-embed-text) to compute embeddings. Since each request only applies to a single document, you can scale this by making requests in parallel across document batches. To send a mapped request and map the response back into the original document, use the [`branch` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/branch/) with a child [`ollama_embeddings`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/ollama_embeddings/) processor. ```yaml pipeline: threads: -1 processors: - branch: request_map: 'root = "search_document: %s\n%s".format(this.article.title, this.article.content)' processors: - ollama_embeddings: model: nomic-embed-text result_map: 'root.article.embeddings = this' ``` With this pipeline, your processed documents should look something like this: ```yaml { "type": "article", "article": { "id": "123foo", "title": "Dogs Stop Barking", "content": "The world was shocked this morning to find that all dogs have stopped barking.", "embeddings": [0.754, 0.19283, 0.231, 0.834], # This vector will actually have 768 dimensions } } ``` Now, try sending this transformed data to PostgreSQL using the [`sql_insert`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/sql_insert/) output. You can take advantage of the `init_statement` functionality to set up `pgvector` and a table to write the data to. ```yaml output: sql_insert: driver: postgres dsn: "TODO" init_statement: | CREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE IF NOT EXISTS searchable_text ( id varchar(128) PRIMARY KEY, title text NOT NULL, body text NOT NULL, embeddings vector(768) NOT NULL ); CREATE INDEX IF NOT EXISTS text_hnsw_index ON searchable_text USING hnsw (embeddings vector_l2_ops); table: searchable_text columns: ["id", "title", "body", "embeddings"] args_mapping: "[this.article.id, this.article.title, this.article.content, this.article.embeddings.vector()]" ``` After deploying this pipeline using the Redpanda Console, you can verify data is being written into PostgreSQL using `psql` to execute `SELECT count(*) FROM searchable_text;`. --- # Page 296: Redpanda Migrator **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/redpanda_migrator.md --- # Redpanda Migrator > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Migrator latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/redpanda_migrator page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/redpanda_migrator.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/redpanda_migrator.adoc description: Move your workloads from any Kafka system to Redpanda Cloud using a single command. page-git-created-date: "2024-10-02" page-git-modified-date: "2026-05-26" --- With Redpanda Migrator, you can move your workloads from any Apache Kafka system to Redpanda using a single command. It lets you migrate Kafka messages, schemas, and ACLs quickly and efficiently. Redpanda Migrator is Redpanda’s alternative to Kafka MirrorMaker 2 for continuous replication into Redpanda, including timestamp-based consumer group offset translation. Redpanda Connect’s Redpanda Migrator uses the unified migrator components (available in Redpanda Connect 4.67.5+): - [`redpanda_migrator` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda_migrator/) connects to the source Kafka cluster and Schema Registry. - [`redpanda_migrator` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_migrator/) handles all migration logic including topic creation, schema synchronization, and consumer group offset translation. > 📝 **NOTE** > > If you’re currently using the legacy `redpanda_migrator_bundle` components, see [Migrate to the Unified Redpanda Migrator](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/migrate-unified-redpanda-migrator/) for migration instructions. ## [](#create-a-kafka-cluster-and-a-redpanda-cloud-cluster)Create a Kafka cluster and a Redpanda Cloud cluster First, you need to provision two clusters, a Kafka one called `source` and a Redpanda Cloud one called `destination`. This cookbook uses the following sample connection details throughout the rest of this cookbook: Source broker: source.cloud.kafka.com:9092 schema registry: https://schema-registry-source.cloud.kafka.com:30081 username: kafka password: testpass Destination broker: destination.cloud.redpanda.com:9092 schema registry: https://schema-registry-destination.cloud.redpanda.com:30081 username: redpanda password: testpass Then you create two topics in the `source` Kafka cluster, `foo` and `bar`, and an ACL for each topic: ```bash cat > ./config.properties < 📝 **NOTE** > > The Brave browser does not fully support code snippets. `generate_data.yaml` ```yaml http: enabled: false input: sequence: inputs: - generate: mapping: | let msg = counter() root.data = $msg meta kafka_topic = match $msg % 2 { 0 => "foo" 1 => "bar" } interval: 1s count: 0 batch_size: 1 processors: - schema_registry_encode: url: "https://schema-registry-source.cloud.kafka.com:30081" subject: ${! metadata("kafka_topic") } avro_raw_json: true basic_auth: enabled: true username: kafka password: testpass output: kafka_franz: seed_brokers: [ "source.cloud.kafka.com:9092" ] topic: ${! @kafka_topic } partitioner: manual partition: ${! random_int(min:0, max:1) } tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: kafka password: testpass ``` > 📝 **NOTE** > > The Brave browser does not fully support code snippets. 5. Click **Create**. Your pipeline details are displayed and the pipeline state changes from **Starting** to **Running**, which may take a few minutes. If you don’t see this state change, refresh your page. Next, add a Redpanda Connect consumer, which reads messages from the `source` cluster topics, and leave it running. This consumer uses the `foobar` consumer group, which is reused in a later step when consuming from the `destination` cluster. 1. Go to the **Connect** page on your cluster and click **Create pipeline**. 2. In **Pipeline name**, enter a name and add a short description. 3. For **Compute units**, leave the default value of **1**. 4. For **Configuration**, paste the following configuration. `read_data_source.yaml` ```yaml http: enabled: false input: kafka_franz: seed_brokers: [ "source.cloud.kafka.com:9092" ] topics: - '^[^_]' # Skip topics which start with `_` regexp_topics: true consumer_group: foobar tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: kafka password: testpass processors: - schema_registry_decode: url: "https://schema-registry-source.cloud.kafka.com:30081" avro_raw_json: true basic_auth: enabled: true username: kafka password: testpass output: stdout: {} processors: - mapping: | root = this.merge({"count": counter(), "topic": @kafka_topic, "partition": @kafka_partition}) ``` > 📝 **NOTE** > > The Brave browser does not fully support code snippets. 5. Click **Create**. Your pipeline details are displayed and the pipeline state changes from **Starting** to **Running**, which may take a few minutes. If you don’t see this state change, refresh your page. At this point, the `source` cluster has some data in both `foo` and `bar` topics, and the consumer prints the messages it reads from these topics to `stdout`. ## [](#required-permissions)Required permissions This cookbook authenticates as a superuser for simplicity. In production, when the source and destination clusters enforce authorization (ACLs), a basic consumer or producer ACL is not sufficient for Redpanda Migrator. The migrator authenticates to the source cluster with the credentials in the `redpanda_migrator` input, and to the destination cluster with the credentials in the `redpanda_migrator` output, so grant ACLs to each principal as follows. > ❗ **IMPORTANT** > > To recreate each topic on the destination with matching settings, Redpanda Migrator reads the source topic’s configuration with a `DescribeConfigs` request, which requires the `DESCRIBE_CONFIGS` operation on the topic. > > A consumer ACL (`READ`) implicitly grants `DESCRIBE`, but it does **not** grant `DESCRIBE_CONFIGS`. If the source principal has only `READ`, the migrator consumes messages successfully but fails to create topics, logging an error such as: > > ```text > level=error msg="Failed to send message to redpanda_migrator: creating records: sync topics: > create topic : get topic details : TOPIC_AUTHORIZATION_FAILED: > Not authorized to access topics: [Topic authorization failed.]" > ``` > > Despite the `create topic` wording, this failure is the `DescribeConfigs` read against the **source** cluster, not the topic creation on the destination. ### [](#source-cluster)Source cluster Grant the source principal (the `redpanda_migrator` input credentials) these ACLs: | Resource | Operations | Purpose | | --- | --- | --- | | Topic (migrated topics) | READ, DESCRIBE_CONFIGS | Consume records (READ, which also grants DESCRIBE for metadata and offsets) and read topic configurations to replicate them (DESCRIBE_CONFIGS). | | Group (the input’s consumer_group) | READ | Join the migrator’s own consumer group and track progress. | | Group (migrated groups) | DESCRIBE | Read source consumer group offsets. Required when consumer group migration is enabled (consumer_groups.enabled, the default). | | Cluster | DESCRIBE | List source consumer groups. Also required to read source ACLs when sync_topic_acls is enabled on the output. | For example, to grant the least-privilege source ACLs to `User:migrator` with `rpk`: ```bash # Consume records and read topic configs (READ also grants DESCRIBE; DESCRIBE_CONFIGS does not come with READ) rpk security acl create --allow-principal User:migrator \ --operation read,describe_configs \ --topic # READ on the migrator's own consumer group rpk security acl create --allow-principal User:migrator \ --operation read \ --group # DESCRIBE on the groups being migrated (omit if consumer_groups.enabled is false) rpk security acl create --allow-principal User:migrator \ --operation describe \ --group # List consumer groups (and describe ACLs if sync_topic_acls is enabled) rpk security acl create --allow-principal User:migrator \ --operation describe \ --cluster ``` ### [](#destination-cluster)Destination cluster Grant the destination principal (the `redpanda_migrator` output credentials) these ACLs: | Resource | Operations | Purpose | | --- | --- | --- | | Topic (migrated topics) | CREATE, WRITE, ALTER, DESCRIBE_CONFIGS | Create topics (CREATE and DESCRIBE_CONFIGS), produce migrated records (WRITE), and add partitions to match the source (ALTER). These operations also grant DESCRIBE. | | Cluster | CREATE | Allow creating destination topics whose names are not known in advance. Use instead of per-topic CREATE. | | Group (migrated groups) | READ | Commit translated consumer group offsets. Required when consumer group migration is enabled (consumer_groups.enabled, the default). | | Cluster | ALTER | Create migrated ACLs on the destination. Required only when sync_topic_acls is enabled. | > 💡 **TIP** > > Run `rpk security acl --help-operations` to see which ACL operation each Kafka request requires. ## [](#configure-and-start-redpanda-migrator)Configure and start Redpanda Migrator The unified Redpanda Migrator does the following: - The `redpanda_migrator` input connects to the source Kafka cluster and Schema Registry to consume messages and schema information. - The `redpanda_migrator` output handles all migration logic: - Schema migration: reads schemas from the source Schema Registry and synchronizes them to the destination. - Topic creation: automatically creates destination topics that don’t exist with proper configurations. - ACL migration: migrates access control lists according to the migration rules. - Message streaming: processes and routes messages from source to destination topics. - Consumer group offset translation: maps source consumer group offsets to equivalent destination positions. - If new topics are created in the source cluster while the migrator is running, they are migrated when messages are written to them. ACL migration for topics adheres to the following principles: - `ALLOW WRITE` ACLs for topics are not migrated - `ALLOW ALL` ACLs for topics are downgraded to `ALLOW READ` - Group ACLs are not migrated > 📝 **NOTE** > > Changing topic configurations, such as partition count, isn’t currently supported. Now, use the following unified Redpanda Migrator configuration. See the [`redpanda_migrator` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda_migrator/) and [`redpanda_migrator` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_migrator/) docs for details. 1. Go to the **Connect** page on your cluster and click **Create pipeline**. 2. In **Pipeline name**, enter a name and add a short description. 3. For **Compute units**, leave the default value of **1**. 4. For **Configuration**, paste the following configuration. `redpanda_migrator.yaml` ```yaml input: label: "migration_pipeline" (1) redpanda_migrator: # Source Kafka settings seed_brokers: [ "source.cloud.kafka.com:9092" ] topics: - '^[^_]' # Skip internal topics which start with `_` regexp_topics: true consumer_group: migrator tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: kafka password: testpass # Source Schema Registry settings schema_registry: url: "https://schema-registry-source.cloud.kafka.com:30081" basic_auth: enabled: true username: kafka password: testpass output: label: "migration_pipeline" (2) redpanda_migrator: # Destination Redpanda settings seed_brokers: [ "destination.cloud.redpanda.com:9092" ] tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: redpanda password: testpass # Destination Schema Registry and migration settings schema_registry: url: https://schema-registry-destination.cloud.redpanda.com:30081 include_deleted: true translate_ids: true basic_auth: enabled: true username: redpanda password: testpass # Consumer group migration settings consumer_groups: enabled: true interval: 30s serverless: false (3) ``` > 💡 **TIP** > > Label names must be between 3 and 128 characters and can only contain alphanumeric characters, hyphens, and underscores (`A-Za-z0-9-_`). ## [](#check-the-status-of-migrated-topics)Check the status of migrated topics You can use the Redpanda [`rpk` CLI tool](https://docs.redpanda.com/streaming/current/get-started/rpk/) to check which topics and ACLs have been migrated to the `destination` cluster. You can quickly [install `rpk`](https://docs.redpanda.com/streaming/current/get-started/rpk-install/) if you don’t already have it. > 📝 **NOTE** > > For now, users require manual migration. However, this step is not required for the current demo. Similarly, roles are specific to Redpanda and, for now, also require manual migration if the `source` cluster is based on Redpanda. ```bash rpk -X brokers=destination.cloud.redpanda.com:9092 -X tls.enabled=true -X sasl.mechanism=SCRAM-SHA-256 -X user=redpanda -X pass=testpass topic list NAME PARTITIONS REPLICAS _schemas 1 1 bar 2 1 foo 2 1 rpk -X brokers=destination.cloud.redpanda.com:9092 -X tls.enabled=true -X sasl.mechanism=SCRAM-SHA-256 -X user=redpanda -X pass=testpass security acl list PRINCIPAL HOST RESOURCE-TYPE RESOURCE-NAME RESOURCE-PATTERN-TYPE OPERATION PERMISSION ERROR User:redpanda * TOPIC bar LITERAL READ DENY User:redpanda * TOPIC foo LITERAL READ ALLOW ``` ## [](#check-metrics-to-monitor-progress)Check metrics to monitor progress Redpanda Connect provides a comprehensive suite of metrics in various formats, such as Prometheus, which you can use to monitor its performance in your observability stack. Besides the [standard Redpanda Connect metrics](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/about/#metric-names), the `redpanda_migrator` input also emits an `input_redpanda_migrator_lag` metric for monitoring the migration progress of each topic and partition. To monitor the migration progress, use the Redpanda Cloud OpenMetrics endpoint, which exposes all Redpanda and connector metrics for your cluster. You can integrate this endpoint with Prometheus, Datadog, or other observability platforms. For step-by-step instructions on configuring monitoring and connecting your observability tool, see [Monitor Redpanda Cloud](https://docs.redpanda.com/cloud-data-platform/manage/monitor-cloud/). After ingesting the metrics, search for the `input_redpanda_migrator_lag` metric in your monitoring tool and filter by `topic` and `partition` as needed to track migration lag for each topic and partition. ## [](#read-from-the-migrated-topics)Read from the migrated topics Stop the `read_data_source.yaml` consumer you started earlier and, afterwards, start a similar consumer for the `destination` cluster. Before starting the consumer up on the `destination` cluster, make sure you give the migrator bundle some time to replicate the translated offset. 1. On the **Connect** page, stop the `read_data_source` pipeline you created earlier. 2. Go to the **Connect** page on your cluster and click **Create pipeline**. 3. In **Pipeline name**, enter a name and add a short description. 4. For **Compute units**, leave the default value of **1**. 5. For **Configuration**, paste the following configuration. `read_data_destination.yaml` ```yaml http: enabled: false input: kafka_franz: seed_brokers: [ "destination.cloud.redpanda.com:9092" ] topics: - '^[^_]' # Skip topics which start with `_` regexp_topics: true consumer_group: foobar sasl: - mechanism: SCRAM-SHA-256 username: redpanda password: testpass processors: - schema_registry_decode: url: "https://schema-registry-destination.cloud.redpanda.com:30081" avro_raw_json: true basic_auth: enabled: true username: redpanda password: testpass output: stdout: {} processors: - mapping: | root = this.merge({"count": counter(), "topic": @kafka_topic, "partition": @kafka_partition}) ``` > 📝 **NOTE** > > The Brave browser does not fully support code snippets. 6. Click **Create**. Your pipeline details are displayed and the pipeline state changes from **Starting** to **Running**, which may take a few minutes. If you don’t see this state change, refresh your page. The `source` cluster consumer uses the same `foobar` consumer group. This consumer resumes reading messages from where the `source` consumer left off. Redpanda Migrator performs offset remapping when migrating consumer group offsets to the `destination` cluster. While more sophisticated approaches are possible, Redpanda chose to use a simple timestamp-based approach. So, for each migrated offset, the `destination` cluster is queried to find the latest offset before the received offset timestamp. Redpanda Migrator then writes this offset as the `destination` consumer group offset for the corresponding topic and partition pair. Although the timestamp-based approach doesn’t guarantee exactly-once delivery, it minimizes the likelihood of message duplication and avoids the need for complex and error-prone offset remapping logic. --- # Page 297: Ingest data into Snowflake **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/snowflake_ingestion.md --- # Ingest data into Snowflake > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Ingest data into Snowflake latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/cookbooks/snowflake_ingestion page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/cookbooks/snowflake_ingestion.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/cookbooks/snowflake_ingestion.adoc description: Configure Redpanda Connect to ingest data from a Redpanda topic into Snowflake using Snowpipe Streaming. page-git-created-date: "2025-01-28" page-git-modified-date: "2026-05-26" --- Configure a Redpanda Connect pipeline to generate and write data into a Redpanda Serverless topic, and then ingest that data into [Snowflake](https://www.snowflake.com/en/) using [Snowpipe Streaming](https://docs.snowflake.com/en/user-guide/data-load-snowpipe-streaming-overview). ## [](#prerequisites)Prerequisites - A [Redpanda Cloud account](https://cloud.redpanda.com/sign-up) - [`rpk` installed](https://docs.redpanda.com/streaming/current/get-started/rpk-install/) and [signed into your Cloud account](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-login/) - A [Snowflake account](https://trial.snowflake.com/) - `openssl` command-line tool ## [](#set-up-your-redpanda-cluster)Set up your Redpanda cluster In [Redpanda Cloud](https://cloud.redpanda.com/), create a new Serverless Standard cluster. When the cluster is ready, run `rpk cloud cluster select` to select the cluster and set it to be your current [rpk profile](https://docs.redpanda.com/streaming/current/get-started/config-rpk-profile/). Next, create a `demo_topic` to use as the data source for ingesting data into Snowflake: ```bash rpk topic create demo_topic ``` Create a user with minimal [ACLs](https://docs.redpanda.com/streaming/current/manage/security/authorization/acl/) to run the ingestion pipeline into Snowflake: ```bash rpk security user create ingestion_user --password Testing1234 ``` Now that the user exists, give them read permissions to `demo_topic`, as well as full control over any consumer group with the prefix `redpanda_connect`: ```bash rpk security acl create --allow-principal ingestion_user --operation read --topic demo_topic rpk security acl create --allow-principal ingestion_user --resource-pattern-type prefixed --operation all --group redpanda_connect ``` ## [](#set-up-your-snowflake-account)Set up your Snowflake account Log in to your Snowflake account with a user who has the ACCOUNTADMIN role. Then, run the following SQL commands in a worksheet. They set up another user with minimal permissions to write data into a specified database and schema, ready for streaming data to Snowflake. ```sql -- Set default values for multiple variables SET PWD = 'Test1234567'; SET USER = 'STREAMING_USER'; SET DB = 'STREAMING_DB'; SET ROLE = 'REDPANDA_CONNECT'; SET WH = 'STREAMING_WH'; USE ROLE ACCOUNTADMIN; -- Create users CREATE USER IF NOT EXISTS IDENTIFIER($USER) PASSWORD=$PWD COMMENT='STREAMING USER FOR REDPANDA CONNECT'; -- Create roles CREATE OR REPLACE ROLE IDENTIFIER($ROLE); -- Create the destination database and virtual warehouse CREATE DATABASE IF NOT EXISTS IDENTIFIER($DB); USE IDENTIFIER($DB); CREATE OR REPLACE WAREHOUSE IDENTIFIER($WH) WITH WAREHOUSE_SIZE = 'SMALL'; -- Grant privileges GRANT CREATE WAREHOUSE ON ACCOUNT TO ROLE IDENTIFIER($ROLE); GRANT ROLE IDENTIFIER($ROLE) TO USER IDENTIFIER($USER); GRANT OWNERSHIP ON DATABASE IDENTIFIER($DB) TO ROLE IDENTIFIER($ROLE); GRANT USAGE ON WAREHOUSE IDENTIFIER($WH) TO ROLE IDENTIFIER($ROLE); -- Set defaults ALTER USER IDENTIFIER($USER) SET DEFAULT_ROLE=$ROLE; ALTER USER IDENTIFIER($USER) SET DEFAULT_WAREHOUSE=$WH; -- Run the following commands to find your account identifier. Copy it down for later use. -- It will be something like `organization_name-account_name` -- e.g. ykmxgak-wyb52636 WITH HOSTLIST AS (SELECT * FROM TABLE(FLATTEN(INPUT => PARSE_JSON(SYSTEM$allowlist())))) SELECT REPLACE(VALUE:host,'.snowflakecomputing.com','') AS ACCOUNT_IDENTIFIER FROM HOSTLIST WHERE VALUE:type = 'SNOWFLAKE_DEPLOYMENT_REGIONLESS'; ``` ### [](#create-an-rsa-key-pair)Create an RSA key pair Create an [RSA key pair](https://docs.snowflake.com/en/user-guide/key-pair-auth) using `openssl` to authenticate Redpanda Connect to Snowflake. When you’re prompted to give an encryption password, record it for later. ```bash openssl genrsa 2048 | openssl pkcs8 -topk8 -inform PEM -passout pass:Testing123 -out rsa_key.p8 ``` Create a public key. You’re prompted to enter your encryption password. ```bash openssl rsa -in rsa_key.p8 -pubout -passout pass:Testing123 -out rsa_key.pub ``` To register the public key in Snowflake, remove the public key delimiters and output only the base64-encoded portion of the PEM file. Run the following bash command to print it: ```bash cat rsa_key.pub | sed -e '1d' -e '$d' | tr -d '\n' ``` In the Snowflake worksheet, add the output of the bash command you just ran to the following SQL command and execute it: ```sql use role accountadmin; alter user streaming_user set rsa_public_key='< PubKeyWithoutDelimiters >'; ``` ### [](#create-a-schema-using-streaming_user)Create a schema using `streaming_user` Log out of Snowflake and sign back in as the default user (`streaming_user`) with the associated password (default: `Test1234567`). You created these credentials in [Set up your Snowflake account](#set-up-your-snowflake-account). Run the following SQL commands in a worksheet to create a schema (e.g. `STREAMING_SCHEMA`) in the default database (e.g. `STREAMING_DB`): ```sql SET DB = 'STREAMING_DB'; SET SCHEMA = 'STREAMING_SCHEMA'; USE IDENTIFIER($DB); CREATE OR REPLACE SCHEMA IDENTIFIER($SCHEMA); ``` ## [](#create-a-pipeline-from-your-redpanda-cluster-to-snowflake)Create a pipeline from your Redpanda cluster to Snowflake You can now create the pipeline. First create [secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) for the passwords and keys you created during setup. On your Serverless cluster, go to the **Connect** page, select the **Secrets** tab and then create three secrets: - `REDPANDA_PASS` with the value `Testing1234` - `SNOWFLAKE_KEY` with the output value of `awk '{printf "%s\\n", $0}' rsa_key.p8` - `SNOWFLAKE_KEY_PASS` with the value `Testing123` Select the **Pipelines** tab and create a pipeline called **RedpandaToSnowflake**. Use the following YAML configuration: ```yaml input: # Reads data from our `demo_topic` kafka_franz: seed_brokers: ["${REDPANDA_BROKERS}"] topics: ["demo_topic"] consumer_group: "redpanda_connect_to_snowflake" tls: {enabled: true} checkpoint_limit: 4096 sasl: - mechanism: SCRAM-SHA-256 username: ingestion_user password: ${secrets.REDPANDA_PASS} # Define the batching policy. This cookbook creates small batches, # but in a production environment use the largest file size you can. batching: count: 100 # Collect 10 messages before flushing period: 10s # or after 10 seconds, whichever comes first output: snowflake_streaming: # Replace this placeholder with your account identifier account: "< OrgName-AccountName >" user: STREAMING_USER role: REDPANDA_CONNECT database: STREAMING_DB schema: STREAMING_SCHEMA table: STREAMING_DATA # Inject your private key and password private_key_file: "${secrets.SNOWFLAKE_KEY}" private_key_pass: "${secrets.SNOWFLAKE_KEY_PASS}" schema_evolution: enabled: true max_in_flight: 1 ``` You now can produce some data using `rpk` to test that everything works: ```bash echo '{"animal":"redpanda","attributes":"cute","age":6}' | rpk topic produce demo_topic -f '%v\n' echo '{"animal":"polar bear","attributes":"cool","age":13}' | rpk topic produce demo_topic -f '%v\n' echo '{"animal":"unicorn","attributes":"rare","age":999}' | rpk topic produce demo_topic -f '%v\n' ``` The data produced into the `demo_topic` is consumed and streamed into Snowflake in seconds. Go back to the Snowflake worksheet and run the following query to see data arrive in Snowflake with the schema from the JSON data you produced. ```sql SELECT * FROM STREAMING_DB.STREAMING_SCHEMA.STREAMING_DATA LIMIT 50; ``` See also: - The [`kafka_franz` input](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/kafka_franz/) - The [`snowflake_streaming`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/snowflake_streaming/) output --- # Page 298: Guides **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides.md --- # Guides > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Guides latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/index.adoc description: Guides for building and operating Redpanda Connect pipelines in Redpanda Cloud, from Bloblang mappings to deployment patterns. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- - [Bloblang](bloblang/about/) Learn what Bloblang is and how to use the native mapping language. - Cloud Credentials - [Amazon Web Services](cloud/aws/) Find out about AWS components in Redpanda Connect. - [Google Cloud Platform](cloud/gcp/) Find out about GCP components in Redpanda Connect. - [Ingest Real-Time Sensor Telemetry with the HTTP Gateway](cloud/gateway/) Learn how to stream sensor telemetry data into Redpanda Cloud using the gateway input in Redpanda Connect. - [Synchronous Responses](sync_responses/) Understand synchronous response handling in Redpanda Connect, ensuring reliable and efficient data processing. - [Migrate to the Unified Redpanda Migrator](migrate-unified-redpanda-migrator/) Learn how to migrate from legacy migrator components to the unified \`redpanda\_migrator\` input/output pair in Redpanda Connect 4.67.5+. --- # Page 299: Bloblang **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about.md --- # Bloblang > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Bloblang latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/bloblang/about page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/bloblang/about.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/bloblang/about.adoc description: Learn what Bloblang is and how to use the native mapping language. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Bloblang, or blobl for short, is a language designed for mapping data of a wide variety of forms. It’s a safe, fast, and powerful way to perform document mapping within Redpanda Connect. It also has a [Go API for writing your own functions and methods](https://pkg.go.dev/github.com/redpanda-data/connect/v4/public/bloblang) as plugins. Bloblang is available as a [processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/mapping/) and it’s also possible to use blobl queries in [function interpolations](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/#bloblang-queries). This document outlines the core features of the Bloblang language, but if you’re totally new to Bloblang then it’s worth following [the walkthrough first](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/walkthrough/). ## [](#learn-bloblang)Learn Bloblang [learnbloblang.com](https://www.learnbloblang.com) is an interactive resource for learning Bloblang with hands-on exercises. ## [](#assignment)Assignment A Bloblang mapping expresses how to create a new document by extracting data from an existing input document. Assignments consist of a dot separated path segments on the left-hand side describing a field to be created within the new document, and a right-hand side query describing what the content of the new field should be. The keyword `root` on the left-hand side refers to the root of the new document, the keyword `this` on the right-hand side refers to the current context of the query, which is the read-only input document when querying from the root of a mapping: ```bloblang root.id = this.thing.id root.type = "yo" # Both `root` and `this` are optional, and will be inferred in their absence. content = thing.doc.message # In: {"thing":{"id":"wat1","doc":{"title":"wut","message":"hello world"}}} ``` Since the document being created starts off empty it is sometimes useful to begin a mapping by copying the entire contents of the input document, which can be expressed by assigning `this` to `root`. ```bloblang root = this root.foo = "added value" # In: {"id":"wat1","message":"hello world"} ``` If the new document `root` is never assigned to or otherwise mutated then the original document remains unchanged. ### [](#special-characters-in-paths)Special characters in paths Quotes can be used to describe sections of a field path that contain whitespace, dots or other special characters: ```bloblang # Use quotes around a path segment in order to include whitespace or dots within # the path root."foo.bar".baz = this."buz bev".fub # In: {"buz bev":{"fub":"hello world"}} ``` ### [](#non-structured-data)Non-structured data Bloblang is able to map data that is unstructured, whether it’s a log line or a binary blob, by referencing it with the [`content` function](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#content), which returns the raw bytes of the input document: ```bloblang # Parse a base64 encoded JSON document root = content().decode("base64").parse_json() # In: eyJmb28iOiJiYXIifQ== ``` And your newly mapped document can also be unstructured, simply assign a value type to the `root` of your document: ```bloblang root = this.foo # In: {"foo":"hello world"} ``` And the resulting message payload will be the raw value you’ve assigned. ### [](#deleting)Deleting It’s possible to selectively delete fields from an object by assigning the function `deleted()` to the field path: ```bloblang root = this root.bar = deleted() # In: {"id":"wat1","message":"hello world","bar":"remove me"} ``` ### [](#variables)Variables Another type of assignment is a `let` statement, which creates a variable that can be referenced elsewhere within a mapping. Variables are discarded at the end of the mapping and are mostly useful for query reuse. Variables are referenced within queries with `$`: ```bloblang # Set a temporary variable let foo = "yo" root.new_doc.type = $foo ``` ### [](#metadata)Metadata Redpanda Connect messages contain metadata that is separate from the main payload, in Bloblang you can modify the metadata of the resulting message with the `meta` assignment keyword. Metadata values of the resulting message are referenced within queries with the `@` operator or the [`metadata()` function](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#metadata): ```bloblang # Reference a metadata value root.new_doc.bar = @kafka_topic # Or `@.kafka_topic` or `metadata("kafka_topic")` # Delete all metadata meta = deleted() # Set metadata values meta bar = "hello world" meta baz = { "something": "structured" } # Get an object of key/values for all metadata root.meta_obj = @ # Or `metadata()` ``` ## [](#coalesce)Coalesce The pipe operator (`|`) used within brackets allows you to coalesce multiple candidates for a path segment. The first field that exists and has a non-null value will be selected: ```bloblang root.new_doc.type = this.thing.(article | comment | this).type # In: {"thing":{"article":{"type":"foo"}}} # In: {"thing":{"comment":{"type":"bar"}}} # In: {"thing":{"type":"baz"}} ``` Opening brackets on a field begins a query where the context of `this` changes to value of the path it is opened upon, therefore in the above example `this` within the brackets refers to the contents of `this.thing`. ## [](#literals)Literals Bloblang supports number, boolean, string, null, array and object literals: ```bloblang root = [ 7, false, "string", null, { "first": 11, "second": {"foo":"bar"}, "third": """multiple lines on this string""" } ] # In: {} ``` The values within literal arrays and objects can be dynamic query expressions, as well as the keys of object literals. ## [](#comments)Comments You might’ve already spotted, comments are started with a hash (`#`) and end with a line break: ```bloblang root = this.some.value # And now this is a comment ``` ## [](#boolean-logic-and-arithmetic)Boolean logic and arithmetic Bloblang supports a range of boolean operators `!`, `>`, `>=`, `==`, `<`, `<=`, `&&`, `||` and mathematical operators `+`, `-`, `*`, `/`, `%`: ```bloblang root.is_big = this.number > 100 root.multiplied = this.number * 7 # In: {"number":50} # In: {"number":150} ``` For more information about these operators and how they work check out [the arithmetic page](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/arithmetic/). ## [](#conditional-mapping)Conditional mapping Use `if` as either a statement or an expression in order to perform maps conditionally: ```bloblang root = this root.sorted_foo = if this.foo.type() == "array" { this.foo.sort() } if this.foo.type() == "string" { root.upper_foo = this.foo.uppercase() root.lower_foo = this.foo.lowercase() } # In: {"foo":"FooBar"} # In: {"foo":["foo","bar"]} ``` And add as many `else if` queries as you like, followed by an optional final fallback `else`: ```bloblang root.sound = if this.type == "cat" { this.cat.meow } else if this.type == "dog" { this.dog.woof.uppercase() } else { "sweet sweet silence" } # In: {"type":"cat","cat":{"meow":"meeeeooooow!"}} # In: {"type":"dog","dog":{"woof":"guurrrr woof woof!"}} # In: {"type":"caterpillar","caterpillar":{"name":"oleg"}} ``` ## [](#pattern-matching)Pattern matching A `match` expression allows you to perform conditional mappings on a value, each case should be either a boolean expression, a literal value to compare against the target value, or an underscore (`_`) which captures values that have not matched a prior case: ```bloblang root.new_doc = match this.doc { this.type == "article" => this.article this.type == "comment" => this.comment _ => this } # In: {"doc":{"type":"article","article":{"id":"foo","content":"qux"}}} # In: {"doc":{"type":"comment","comment":{"id":"bar","content":"quz"}}} # In: {"doc":{"type":"neither","content":"some other stuff unchanged"}} ``` Within a match block the context of `this` changes to the pattern matched expression, therefore `this` within the match expression above refers to `this.doc`. Match cases can specify a literal value for simple comparison: ```bloblang root = this root.type = match this.type { "doc" => "document", "art" => "article", _ => this } # In: {"type":"doc","foo":"bar"} ``` The match expression can also be left unset which means the context remains unchanged, and the catch-all case can also be omitted: ```bloblang root.new_doc = match { this.doc.type == "article" => this.doc.article this.doc.type == "comment" => this.doc.comment } # In: {"doc":{"type":"neither","content":"some other stuff unchanged"}} ``` If no case matches then the mapping is skipped entirely, hence we would end up with the original document in this case. ## [](#functions)Functions Functions can be placed anywhere and allow you to extract information from your environment, generate values, or access data from the underlying message being mapped: ```bloblang root.doc.id = uuid_v4() root.doc.received_at = now() root.doc.host = hostname() ``` Functions support both named and nameless style arguments: ```bloblang root.values_one = range(start: 0, stop: this.max, step: 2) root.values_two = range(0, this.max, 2) # In: {"max":10} ``` You can find a full list of functions and their parameters in [the functions page](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/). ## [](#methods)Methods Methods are similar to functions but enact upon a target value, these provide most of the power in Bloblang as they allow you to augment query values and can be added to any expression (including other methods): ```bloblang root.doc.id = this.thing.id.string().catch(uuid_v4()) root.doc.reduced_nums = this.thing.nums.map_each(num -> if num < 10 { deleted() } else { num - 10 }) root.has_good_taste = ["pikachu","mewtwo","magmar"].contains(this.user.fav_pokemon) # In: {"thing":{"id":123,"nums":[5,12,8,15,20]},"user":{"fav_pokemon":"pikachu"}} ``` Methods also support both named and nameless style arguments: ```bloblang root.foo_one = this.(bar | baz).trim().replace_all(old: "dog", new: "cat") root.foo_two = this.(bar | baz).trim().replace_all("dog", "cat") # In: {"bar":" I love my dog "} ``` You can find a full list of methods and their parameters in [the methods page](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/). ## [](#maps)Maps Defining named maps allows you to reuse common mappings on values with the [`apply` method](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#apply): ```bloblang map things { root.first = this.thing_one root.second = this.thing_two } root.foo = this.value_one.apply("things") root.bar = this.value_two.apply("things") # In: {"value_one":{"thing_one":"hey","thing_two":"yo"},"value_two":{"thing_one":"sup","thing_two":"waddup"}} ``` Within a map the keyword `root` refers to a newly created document that will replace the target of the map, and `this` refers to the original value of the target. The argument of `apply` is a string, which allows you to dynamically resolve the mapping to apply. ## [](#import-maps)Import maps It’s possible to import maps defined in a file with an `import` statement: ```bloblang import "./common_maps.blobl" root.foo = this.value_one.apply("things") root.bar = this.value_two.apply("things") # In: {"value_one":{"thing_one":"hey","thing_two":"yo"},"value_two":{"thing_one":"sup","thing_two":"waddup"}} ``` Imports from a Bloblang mapping within a Redpanda Connect config are relative to the process running the config. Imports from an imported file are relative to the file that is importing it. ## [](#filtering)Filtering By assigning the root of a mapped document to the `deleted()` function you can delete a message entirely: ```bloblang # Filter all messages that have fewer than 10 URLs. root = if this.doc.urls.length() < 10 { deleted() } # In: {"doc":{"urls":["a","b","c"]}} # In: {"doc":{"urls":["a","b","c","d","e","f","g","h","i","j"]}} ``` ## [](#error-handling)Error handling Functions and methods can fail under certain circumstances, such as when they receive types they aren’t able to act upon. These failures, when not caught, will cause the entire mapping to fail. However, the [method `catch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#catch) can be used in order to return a value when a failure occurs instead: ```bloblang # Map an empty array to `foo` if the field `bar` is not a string. root.foo = this.bar.split(",").catch([]) # In: {"bar":"a,b,c"} # In: {"bar":123} ``` Since `catch` is a method it can also be attached to bracketed map expressions: ```bloblang # Map `false` if any of the operations in this boolean query fail. root.thing = ( this.foo > this.bar && this.baz.contains("wut") ).catch(false) # In: {"foo":10,"bar":5,"baz":"wut wut"} # In: {"foo":"not a number","bar":5,"baz":"wut wut"} ``` And one of the more powerful features of Bloblang is that a single `catch` method at the end of a chain of methods can recover errors from any method in the chain: ```bloblang # Catch errors caused by: # - foo not existing # - foo not being a string # - an element from split foo not being a valid JSON string root.things = this.foo.split(",").map_each( ele -> ele.parse_json() ).catch([]) # Specifically catch a JSON parse error root.things = this.foo.split(",").map_each( ele -> ele.parse_json().catch({}) ) # In: {"foo":"{\"a\":1},{\"b\":2}"} # In: {"foo":"not valid json"} ``` However, the `catch` method only acts on errors, sometimes it’s also useful to set a fall back value when a query returns `null` in which case the [method `or`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#or) can be used the same way: ```bloblang # Map "default" if either the element index 5 does not exist, or the underlying # element is `null`. root.foo = this.bar.index(5).or("default") # In: {"bar":["a","b","c"]} # In: {"bar":["a","b","c","d","e","f","g"]} ``` ## [](#unit-testing)Unit testing It’s possible to execute unit tests for your Bloblang mappings using the standard Redpanda Connect unit test capabilities outlined [in this document](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/unit_testing/). ## [](#troubleshooting)Troubleshooting 1. I’m seeing `unable to reference message as structured (with 'this')` when I try to run mappings with `rpk connect blobl`. That particular error message means the mapping is failing to parse what’s being fed in as a JSON document. Make sure that the data you are feeding in is valid JSON, and also that the documents _do not_ contain line breaks as `rpk connect blobl` will parse each line individually. Why? That’s a good question. Bloblang supports non-JSON formats too, so it can’t delimit documents with a streaming JSON parser like tools such as `jq`, so instead it uses line breaks to determine the boundaries of each message. --- # Page 300: Bloblang Arithmetic **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/arithmetic.md --- # Bloblang Arithmetic > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Bloblang Arithmetic latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/bloblang/arithmetic page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/bloblang/arithmetic.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/bloblang/arithmetic.adoc description: How arithmetic works within Bloblang page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Bloblang supports a range of comparison operators `!`, `>`, `>=`, `==`, `<`, `<=`, `&&`, `||` and mathematical operators `+`, `-`, `*`, `/`, `%`. How these operators behave is dependent on the type of the values they’re used with, and therefore it’s worth fully understanding these behaviors if you intend to use them heavily in your mappings. ## [](#mathematical)Mathematical All mathematical operators (`+`, `-`, `*`, `/`, `%`) are valid against number values, and addition (`+`) is also supported when both the left and right hand side arguments are strings. If a mathematical operator is used with an argument that is non-numeric (with the aforementioned string exception) then a [recoverable mapping error will be thrown](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#error-handling). ### [](#number-degradation)Number degradation In Bloblang any number resulting from a method, function or arithmetic is either a 64-bit signed integer or a 64-bit floating point value. Numbers from input documents can be any combination of size and be signed or unsigned. When a mathematical operation is performed with two or more integer values Bloblang will create an integer result, with the exception of division. However, if any number within a mathematical operation is a floating point then the result will be a floating point value. In order to explicitly coerce numbers into integer types you can use the [`.ceil()`, `.floor()`, or `.round()` methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#number-manipulation). ## [](#comparison)Comparison The not (`!`) operator reverses the boolean value of the expression immediately following it, and is valid to place before any query that yields a boolean value. If the following expression yields a non-boolean value then a [recoverable mapping error will be thrown](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#error-handling). If you wish to reverse the boolean result of a complex query then simply place the query within brackets (`!(this.foo > this.bar)`). ### [](#equality)Equality The equality operators (`==` and `!=`) are valid to use against any value type. In order for arguments to be considered equal they must match in both their basic type (`string`, `number`, `null`, `bool`, etc) as well as their value. If you wish to compare mismatched value types then use [coercion methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#type-coercion). Number arguments are considered equal if their value is the same when represented the same way, which means their underlying representations (integer, float, etc) do not need to match in order for them to be considered equal. ### [](#numerical)Numerical Numerical comparisons (`>`, `>=`, `<`, `<=`) are valid to use against number values only. If a non-number value is used as an argument then a [recoverable mapping error will be thrown](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#error-handling). ### [](#boolean)Boolean Boolean comparison operators (`||`, `&&`) are valid to use against boolean values only (`true` or `false`). If a non-boolean value is used as an argument then a [recoverable mapping error will be thrown](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#error-handling). --- # Page 301: Bloblang Functions **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions.md --- # Bloblang Functions > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Bloblang Functions latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/bloblang/functions page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/bloblang/functions.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/bloblang/functions.adoc description: A list of Bloblang functions page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Functions can be placed anywhere and allow you to extract information from your environment, generate values, or access data from the underlying message being mapped: ```bloblang root.doc.id = uuid_v4() root.doc.received_at = now() root.doc.host = hostname() ``` Functions support both named and nameless style arguments: ```bloblang root.values_one = range(start: 0, stop: this.max, step: 2) root.values_two = range(0, this.max, 2) # In: {"max":10} ``` ## [](#batch_index)batch_index Returns the zero-based index of the current message within its batch. Use this to conditionally process messages based on their position, or to create sequential identifiers within a batch. ### [](#examples)Examples ```bloblang root = if batch_index() > 0 { deleted() } ``` Create a unique identifier combining batch position with timestamp: ```bloblang root.id = "%v-%v".format(timestamp_unix(), batch_index()) ``` ## [](#batch_size)batch_size Returns the total number of messages in the current batch. Use this to determine batch boundaries or compute relative positions. ### [](#examples-2)Examples ```bloblang root.total = batch_size() ``` Check if processing the last message in a batch: ```bloblang root.is_last = batch_index() == batch_size() - 1 ``` ## [](#bytes)bytes Creates a zero-initialized byte array of specified length. Use this to allocate fixed-size byte buffers for binary data manipulation or to generate padding. ### [](#parameters)Parameters | Name | Type | Description | | --- | --- | --- | | length | integer | The size of the resulting byte array. | ### [](#examples-3)Examples ```bloblang root.data = bytes(5) ``` Create a buffer for binary operations: ```bloblang root.header = bytes(16) root.payload = content() ``` ## [](#content)content Returns the raw message payload as bytes, regardless of the current mapping context. Use this to access the original message when working within nested contexts, or to store the entire message as a field. ### [](#examples-4)Examples ```bloblang root.doc = content().string() # In: {"foo":"bar"} # Out: {"doc":"{\"foo\":\"bar\"}"} ``` Preserve original message while adding metadata: ```bloblang root.original = content().string() root.processed_by = "ai" # In: {"foo":"bar"} # Out: {"original":"{\"foo\":\"bar\"}","processed_by":"ai"} ``` ## [](#count)count > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. The `count` function is a counter starting at 1 which increments after each time it is called. Count takes an argument which is an identifier for the counter, allowing you to specify multiple unique counters in your configuration. ### [](#parameters-2)Parameters | Name | Type | Description | | --- | --- | --- | | name | string | An identifier for the counter. | ### [](#examples-5)Examples ```bloblang root = this root.id = count("bloblang_function_example") # In: {"message":"foo"} # Out: {"id":1,"message":"foo"} # In: {"message":"bar"} # Out: {"id":2,"message":"bar"} ``` ## [](#counter)counter Generates an incrementing sequence of integers starting from a minimum value (default 1). Each counter instance maintains its own independent state across message processing. When the maximum value is reached, the counter automatically resets to the minimum. ### [](#parameters-3)Parameters | Name | Type | Description | | --- | --- | --- | | min | query expression | The starting value of the counter. This is the first value yielded. Evaluated once when the mapping is initialized. | | max | query expression | The maximum value before the counter resets to min. Evaluated once when the mapping is initialized. | | set (optional) | query expression | An optional query that controls counter behavior: when it resolves to a non-negative integer, the counter is set to that value; when it resolves to null, the counter is read without incrementing; when it resolves to a deletion, the counter resets to min; otherwise the counter increments normally. | ### [](#examples-6)Examples Generate sequential IDs for each message: ```bloblang root.id = counter() # In: {} # Out: {"id":1} # In: {} # Out: {"id":2} ``` Use a custom range for the counter: ```bloblang root.batch_num = counter(min: 100, max: 200) # In: {} # Out: {"batch_num":100} # In: {} # Out: {"batch_num":101} ``` Increment a counter multiple times within a single mapping using a named map: ```bloblang map increment { root = counter() } root.first_id = null.apply("increment") root.second_id = null.apply("increment") # In: {} # Out: {"first_id":1,"second_id":2} # In: {} # Out: {"first_id":3,"second_id":4} ``` Conditionally reset a counter based on input data: ```bloblang root.streak = counter(set: if this.status != "success" { 0 }) # In: {"status":"success"} # Out: {"streak":1} # In: {"status":"success"} # Out: {"streak":2} # In: {"status":"failure"} # Out: {"streak":0} # In: {"status":"success"} # Out: {"streak":1} ``` Peek at the current counter value without incrementing by using null in the set parameter: ```bloblang root.count = counter(set: if this.peek { null }) # In: {"peek":false} # Out: {"count":1} # In: {"peek":false} # Out: {"count":2} # In: {"peek":true} # Out: {"count":2} # In: {"peek":false} # Out: {"count":3} ``` ## [](#deleted)deleted Returns a deletion marker that removes the target field or message. When applied to root, the entire message is dropped while still being acknowledged as successfully processed. Use this to filter data or conditionally remove fields. ### [](#examples-7)Examples ```bloblang root = this root.bar = deleted() # In: {"bar":"bar_value","baz":"baz_value","foo":"foo value"} # Out: {"baz":"baz_value","foo":"foo value"} ``` Filter array elements by returning deleted for unwanted items: ```bloblang root.new_nums = this.nums.map_each(num -> if num < 10 { deleted() } else { num - 10 }) # In: {"nums":[3,11,4,17]} # Out: {"new_nums":[1,7]} ``` ## [](#env)env Reads an environment variable and returns its value as a string. Returns `null` if the variable is not set. By default, values are cached for performance. ### [](#parameters-4)Parameters | Name | Type | Description | | --- | --- | --- | | name | string | The name of the environment variable to read. | | no_cache | bool | Disable caching to read the latest value on each invocation. | ### [](#examples-8)Examples ```bloblang root.api_key = env("API_KEY") ``` ```bloblang root.database_url = env("DB_URL").or("localhost:5432") ``` Use `no_cache` to read updated environment variables during runtime, useful for dynamic configuration changes: ```bloblang root.config = env(name: "DYNAMIC_CONFIG", no_cache: true) ``` ## [](#error)error Returns the error message string if the message has failed processing, otherwise `null`. Use this in error handling pipelines to log or route failed messages based on their error details. ### [](#examples-9)Examples ```bloblang root.doc.error = error() ``` Route messages to different outputs based on error presence: ```bloblang root = this root.error_msg = error() root.has_error = error() != null ``` ## [](#error_source_label)error_source_label Returns the user-defined label of the component that caused the error, empty string if no label is set, or `null` if the message has no error. Use this for more human-readable error tracking when components have custom labels. ### [](#examples-10)Examples ```bloblang root.doc.error_source_label = error_source_label() ``` Route errors based on component labels: ```bloblang root.error_category = error_source_label().or("unknown") ``` ## [](#error_source_name)error_source_name Returns the component name that caused the error, or `null` if the message has no error or the error has no associated component. Use this to identify which processor or component in your pipeline caused a failure. ### [](#examples-11)Examples ```bloblang root.doc.error_source_name = error_source_name() ``` Create detailed error logs with component information: ```bloblang root.error_details = if errored() { { "message": error(), "component": error_source_name(), "timestamp": now() } } ``` ## [](#error_source_path)error_source_path Returns the dot-separated path to the component that caused the error, or `null` if the message has no error. Use this to identify the exact location of a failed component in nested pipeline configurations. ### [](#examples-12)Examples ```bloblang root.doc.error_source_path = error_source_path() ``` Build comprehensive error context for debugging: ```bloblang root.error_info = { "path": error_source_path(), "component": error_source_name(), "message": error() } ``` ## [](#errored)errored Returns true if the message has failed processing, false otherwise. Use this for conditional logic in error handling workflows or to route failed messages to dead letter queues. ### [](#examples-13)Examples ```bloblang root.doc.status = if errored() { 400 } else { 200 } ``` Send only failed messages to a separate stream: ```bloblang root = if errored() { this } else { deleted() } ``` ## [](#fake)fake Generates realistic fake data for testing and development purposes. Supports a wide variety of data types including personal information, network addresses, dates/times, financial data, and UUIDs. Useful for creating mock data, populating test databases, or anonymizing sensitive information. Supported functions: `latitude`, `longitude`, `unix_time`, `date`, `time_string`, `month_name`, `year_string`, `day_of_week`, `day_of_month`, `timestamp`, `century`, `timezone`, `time_period`, `email`, `mac_address`, `domain_name`, `url`, `username`, `ipv4`, `ipv6`, `password`, `jwt`, `word`, `sentence`, `paragraph`, `cc_type`, `cc_number`, `currency`, `amount_with_currency`, `title_male`, `title_female`, `first_name`, `first_name_male`, `first_name_female`, `last_name`, `name`, `gender`, `chinese_first_name`, `chinese_last_name`, `chinese_name`, `phone_number`, `toll_free_phone_number`, `e164_phone_number`, `uuid_hyphenated`, `uuid_digit`. ### [](#parameters-5)Parameters | Name | Type | Description | | --- | --- | --- | | function | string | The name of the faker function to use. See description for full list of supported functions. | ### [](#examples-14)Examples Generate fake user profile data for testing: ```bloblang root.user = { "id": fake("uuid_hyphenated"), "name": fake("name"), "email": fake("email"), "created_at": fake("timestamp") } ``` Create realistic test data for network monitoring: ```bloblang root.event = { "source_ip": fake("ipv4"), "mac_address": fake("mac_address"), "url": fake("url") } ``` ## [](#file)file Reads a file and returns its contents as bytes. Paths are resolved from the process working directory. For paths relative to the mapping file, use `file_rel`. By default, files are cached after first read. ### [](#parameters-6)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | The absolute or relative path to the file. | | no_cache | bool | Disable caching to read the latest file contents on each invocation. | ### [](#examples-15)Examples ```bloblang root.config = file("/etc/config.json").parse_json() ``` ```bloblang root.template = file("./templates/email.html").string() ``` Use `no_cache` to read updated file contents during runtime, useful for hot-reloading configuration: ```bloblang root.rules = file(path: "/etc/rules.yaml", no_cache: true).parse_yaml() ``` ## [](#file_rel)file_rel Reads a file and returns its contents as bytes. Paths are resolved relative to the mapping file’s directory, making it portable across different environments. By default, files are cached after first read. ### [](#parameters-7)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | The path to the file, relative to the mapping file’s directory. | | no_cache | bool | Disable caching to read the latest file contents on each invocation. | ### [](#examples-16)Examples ```bloblang root.schema = file_rel("./schemas/user.json").parse_json() ``` ```bloblang root.lookup = file_rel("../data/lookup.csv").parse_csv() ``` Use `no_cache` to read updated file contents during runtime, useful for reloading data files without restarting: ```bloblang root.translations = file_rel(path: "./i18n/en.yaml", no_cache: true).parse_yaml() ``` ## [](#hostname)hostname Returns the hostname of the machine running Benthos. Useful for identifying which instance processed a message in distributed deployments. ### [](#examples-17)Examples ```bloblang root.processed_by = hostname() ``` ## [](#json)json Returns a field from the original JSON message by dot path, always accessing the root document regardless of mapping context. Use this to reference the source message when working in nested contexts or to extract specific fields. ### [](#parameters-8)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | An optional [dot path][field_paths] identifying a field to obtain. | ### [](#examples-18)Examples ```bloblang root.mapped = json("foo.bar") # In: {"foo":{"bar":"hello world"}} # Out: {"mapped":"hello world"} ``` Access the original message from within nested mapping contexts: ```bloblang root.doc = json() # In: {"foo":{"bar":"hello world"}} # Out: {"doc":{"foo":{"bar":"hello world"}}} ``` ## [](#ksuid)ksuid Generates a K-Sortable Unique Identifier with built-in timestamp ordering. Use this for distributed unique IDs that sort chronologically and remain collision-resistant without coordination between generators. ### [](#examples-19)Examples ```bloblang root.id = ksuid() ``` Create sortable event IDs for logging: ```bloblang root.event = { "id": ksuid(), "type": this.event_type, "data": this.payload } ``` ## [](#meta)meta > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Returns the value of a metadata key from the input message as a string, or `null` if the key does not exist. Since values are extracted from the read-only input message they do NOT reflect changes made from within the map. In order to query metadata mutations made within a mapping use the [`root_meta` function](#root_meta). This function supports extracting metadata from other messages of a batch with the `from` method. ### [](#parameters-9)Parameters | Name | Type | Description | | --- | --- | --- | | key | string | An optional key of a metadata value to obtain. | ### [](#examples-20)Examples ```bloblang root.topic = meta("kafka_topic") ``` The key parameter is optional and if omitted the entire metadata contents are returned as an object: ```bloblang root.all_metadata = meta() ``` ## [](#metadata)metadata Returns metadata from the input message by key, or `null` if the key doesn’t exist. This reads the original metadata; to access modified metadata during mapping, use the `@` operator instead. Use this to extract message properties like topics, headers, or timestamps. ### [](#parameters-10)Parameters | Name | Type | Description | | --- | --- | --- | | key | string | An optional key of a metadata value to obtain. | ### [](#examples-21)Examples ```bloblang root.topic = metadata("kafka_topic") ``` Retrieve all metadata as an object by omitting the key parameter: ```bloblang root.all_metadata = metadata() ``` Copy specific metadata fields to the message body: ```bloblang root.source = { "topic": metadata("kafka_topic"), "partition": metadata("kafka_partition"), "timestamp": metadata("kafka_timestamp_unix") } ``` ## [](#nanoid)nanoid Generates a URL-safe unique identifier using Nano ID. Use this for compact, URL-friendly IDs with good collision resistance. Customize the length (default 21) or provide a custom alphabet for specific character requirements. ### [](#parameters-11)Parameters | Name | Type | Description | | --- | --- | --- | | length (optional) | integer | An optional length. | | alphabet (optional) | string | An optional custom alphabet to use for generating IDs. When specified the field length must also be present. | ### [](#examples-22)Examples ```bloblang root.id = nanoid() ``` Generate a longer ID for additional uniqueness: ```bloblang root.id = nanoid(54) ``` Use a custom alphabet for domain-specific IDs: ```bloblang root.id = nanoid(54, "abcde") ``` ## [](#nothing)nothing ## [](#now)now Returns the current timestamp as an RFC 3339 formatted string with nanosecond precision. Use this to add processing timestamps to messages or measure time between events. Chain with `ts_format` to customize the format or timezone. ### [](#examples-23)Examples ```bloblang root.received_at = now() ``` Format the timestamp in a custom format and timezone: ```bloblang root.received_at = now().ts_format("Mon Jan 2 15:04:05 -0700 MST 2006", "UTC") ``` ## [](#pi)pi Returns the value of the mathematical constant Pi. ### [](#examples-24)Examples ```bloblang root.radians = this.degrees * (pi() / 180) # In: {"degrees":45} # Out: {"radians":0.7853981633974483} ``` ```bloblang root.degrees = this.radians * (180 / pi()) # In: {"radians":0.78540} # Out: {"degrees":45.00010522957486} ``` ## [](#random_int)random_int Generates a pseudo-random non-negative 64-bit integer. Use this for creating random IDs, sampling data, or generating test values. Provide a seed for reproducible randomness, or use a dynamic seed like `timestamp_unix_nano()` for unique values per mapping instance. Optional `min` and `max` parameters constrain the output range (both inclusive). For dynamic ranges based on message data, use the modulo operator instead: `random_int() % dynamic_max + dynamic_min`. ### [](#parameters-12)Parameters | Name | Type | Description | | --- | --- | --- | | seed | query expression | A seed to use, if a query is provided it will only be resolved once during the lifetime of the mapping. | | min | integer | The minimum value the random generated number will have. The default value is 0. | | max | integer | The maximum value the random generated number will have. The default value is 9223372036854775806 (math.MaxInt64 - 1). | ### [](#examples-25)Examples ```bloblang root.first = random_int() root.second = random_int(1) root.third = random_int(max:20) root.fourth = random_int(min:10, max:20) root.fifth = random_int(timestamp_unix_nano(), 5, 20) root.sixth = random_int(seed:timestamp_unix_nano(), max:20) ``` Use a dynamic seed for unique random values per mapping instance: ```bloblang root.random_id = random_int(timestamp_unix_nano()) root.sample_percent = random_int(seed: timestamp_unix_nano(), min: 0, max: 100) ``` ## [](#range)range Creates an array of integers from start (inclusive) to stop (exclusive) with an optional step. Use this to generate sequences for iteration, indexing, or creating numbered lists. ### [](#parameters-13)Parameters | Name | Type | Description | | --- | --- | --- | | start | integer | The start value. | | stop | integer | The stop value. | | step | integer | The step value. | ### [](#examples-26)Examples ```bloblang root.a = range(0, 10) root.b = range(start: 0, stop: this.max, step: 2) # Using named params root.c = range(0, -this.max, -2) # In: {"max":10} # Out: {"a":[0,1,2,3,4,5,6,7,8,9],"b":[0,2,4,6,8],"c":[0,-2,-4,-6,-8]} ``` Generate a sequence for batch processing: ```bloblang root.pages = range(0, this.total_items, 100).map_each(offset -> { "offset": offset, "limit": 100 }) # In: {"total_items":250} # Out: {"pages":[{"limit":100,"offset":0},{"limit":100,"offset":100}]} ``` ## [](#root_meta)root_meta > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Returns the value of a metadata key from the new message being created as a string, or `null` if the key does not exist. Changes made to metadata during a mapping will be reflected by this function. ### [](#parameters-14)Parameters | Name | Type | Description | | --- | --- | --- | | key | string | An optional key of a metadata value to obtain. | ### [](#examples-27)Examples ```bloblang root.topic = root_meta("kafka_topic") ``` The key parameter is optional and if omitted the entire metadata contents are returned as an object: ```bloblang root.all_metadata = root_meta() ``` ## [](#snowflake_id)snowflake_id Generates a unique, time-ordered Snowflake ID. Snowflake IDs are 64-bit integers that encode timestamp, node ID, and sequence information, making them ideal for distributed systems where sortable unique identifiers are needed. Returns a string representation of the ID. ### [](#parameters-15)Parameters | Name | Type | Description | | --- | --- | --- | | node_id | integer | Optional node identifier (0-1023) to distinguish IDs generated by different machines in a distributed system. Defaults to 1. | ### [](#examples-28)Examples Generate a unique Snowflake ID for each message: ```bloblang root.id = snowflake_id() root.payload = this ``` Generate Snowflake IDs with different node IDs for multi-datacenter deployments: ```bloblang root.id = snowflake_id(42) root.data = this ``` ## [](#throw)throw Immediately fails the mapping with a custom error message. Use this to halt processing when data validation fails or required fields are missing, causing the message to be routed to error handlers. ### [](#parameters-16)Parameters | Name | Type | Description | | --- | --- | --- | | why | string | A string explanation for why an error was thrown, this will be added to the resulting error message. | ### [](#examples-29)Examples ```bloblang root.doc.type = match { this.exists("header.id") => "foo" this.exists("body.data") => "bar" _ => throw("unknown type") } root.doc.contents = (this.body.content | this.thing.body) # In: {"header":{"id":"first"},"thing":{"body":"hello world"}} # Out: {"doc":{"contents":"hello world","type":"foo"}} # In: {"nothing":"matches"} # Out: Error("failed assignment (line 1): unknown type") ``` Validate required fields before processing: ```bloblang root = if this.exists("user_id") { this } else { throw("missing required field: user_id") } # In: {"user_id":123,"name":"alice"} # Out: {"name":"alice","user_id":123} # In: {"name":"bob"} # Out: Error("failed assignment (line 1): missing required field: user_id") ``` ## [](#timestamp_unix)timestamp_unix Returns the current Unix timestamp in seconds since epoch. Use this for numeric timestamps compatible with most systems, or as a seed for random number generation. ### [](#examples-30)Examples ```bloblang root.received_at = timestamp_unix() ``` Create a sortable ID combining timestamp with a counter: ```bloblang root.id = "%v-%v".format(timestamp_unix(), batch_index()) ``` ## [](#timestamp_unix_micro)timestamp_unix_micro Returns the current Unix timestamp in microseconds since epoch. Use this for high-precision timing measurements or when microsecond resolution is required. ### [](#examples-31)Examples ```bloblang root.received_at = timestamp_unix_micro() ``` Measure elapsed time between events: ```bloblang root.processing_duration_us = timestamp_unix_micro() - this.start_time_us ``` ## [](#timestamp_unix_milli)timestamp_unix_milli Returns the current Unix timestamp in milliseconds since epoch. Use this for millisecond-precision timestamps common in web APIs and JavaScript systems. ### [](#examples-32)Examples ```bloblang root.received_at = timestamp_unix_milli() ``` Add processing time metadata: ```bloblang meta processing_time_ms = timestamp_unix_milli() ``` ## [](#timestamp_unix_nano)timestamp_unix_nano Returns the current Unix timestamp in nanoseconds since epoch. Use this for the highest precision timing or as a unique seed value that changes on every invocation. ### [](#examples-33)Examples ```bloblang root.received_at = timestamp_unix_nano() ``` Generate unique random values on each mapping: ```bloblang root.random_value = random_int(timestamp_unix_nano()) ``` ## [](#tracing_id)tracing_id Returns the OpenTelemetry trace ID for the message, or an empty string if no tracing span exists. Use this to correlate logs and events with distributed traces. ### [](#examples-34)Examples ```bloblang meta trace_id = tracing_id() ``` Add trace ID to structured logs: ```bloblang root.log_entry = this root.log_entry.trace_id = tracing_id() ``` ## [](#tracing_span)tracing_span Returns the OpenTelemetry tracing span attached to the message as a text map object, or `null` if no span exists. Use this to propagate trace context to downstream systems via headers or metadata. ### [](#examples-35)Examples ```bloblang root.headers.traceparent = tracing_span().traceparent # In: {"some_stuff":"just can't be explained by science"} # Out: {"headers":{"traceparent":"00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01"}} ``` Forward all tracing fields to output metadata: ```bloblang meta = tracing_span() ``` ## [](#ulid)ulid Generates a Universally Unique Lexicographically Sortable Identifier (ULID). ULIDs are 128-bit identifiers that are sortable by creation time, URL-safe, and case-insensitive. They consist of a 48-bit timestamp (millisecond precision) and 80 bits of randomness, making them ideal for distributed systems that need time-ordered unique IDs without coordination. ### [](#parameters-17)Parameters | Name | Type | Description | | --- | --- | --- | | encoding | string | Encoding format for the ULID. "crockford" produces 26-character Base32 strings (recommended). "hex" produces 32-character hexadecimal strings. | | random_source | string | Randomness source: "secure_random" uses cryptographically secure random (recommended for production), "fast_random" uses faster but non-secure random (only for non-sensitive testing). | ### [](#examples-36)Examples Generate time-sortable IDs for distributed message ordering: ```bloblang root.message_id = ulid() root.timestamp = now() root.data = this ``` Generate hex-encoded ULIDs for systems that prefer hexadecimal format: ```bloblang root.id = ulid("hex") ``` ## [](#uuid_v4)uuid_v4 Generates a random RFC-4122 version 4 UUID. Use this for creating unique identifiers that don’t reveal timing information or require ordering. Each invocation produces a new globally unique ID. ### [](#examples-37)Examples ```bloblang root.id = uuid_v4() ``` Add unique request IDs for tracing: ```bloblang root = this root.request_id = uuid_v4() ``` ## [](#uuid_v7)uuid_v7 Generates a time-ordered UUID version 7 with millisecond timestamp precision. Use this for sortable unique identifiers that maintain chronological ordering, ideal for database keys or event IDs. Optionally specify a custom timestamp. ### [](#parameters-18)Parameters | Name | Type | Description | | --- | --- | --- | | time (optional) | timestamp | An optional timestamp to use for the time ordered portion of the UUID. | ### [](#examples-38)Examples ```bloblang root.id = uuid_v7() ``` Generate a UUID with a specific timestamp for backdating events: ```bloblang root.id = uuid_v7(now().ts_sub_iso8601("PT1M")) ``` ## [](#var)var ### [](#parameters-19)Parameters | Name | Type | Description | | --- | --- | --- | | name | string | The name of the target variable. | ## [](#with_schema_registry_header)with_schema_registry_header Prepends a Confluent Schema Registry wire format header to message bytes. The header is 5 bytes: a magic byte (0x00) followed by a 4-byte big-endian schema ID. This format is required when producing messages to Kafka topics that use Confluent Schema Registry for schema validation and evolution. ### [](#parameters-20)Parameters | Name | Type | Description | | --- | --- | --- | | schema_id | unknown | The schema ID from your Schema Registry (0 to 4294967295). This ID references the schema version used to encode the message. | | message | unknown | The serialized message bytes (e.g., Avro, Protobuf, or JSON Schema encoded data) to prepend the header to. | ### [](#examples-39)Examples Add Schema Registry header to Avro-encoded message: ```bloblang root = with_schema_registry_header(123, content()) ``` Use schema ID from metadata to add header dynamically: ```bloblang root = with_schema_registry_header(meta("schema_id").number(), content()) ``` --- # Page 302: Bloblang Methods **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods.md --- # Bloblang Methods > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Bloblang Methods latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/bloblang/methods page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/bloblang/methods.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/bloblang/methods.adoc description: A list of Bloblang methods page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Methods provide most of the power in Bloblang as they allow you to augment values and can be added to any expression (including other methods): ```bloblang root.doc.id = this.thing.id.string().catch(uuid_v4()) root.doc.reduced_nums = this.thing.nums.map_each(num -> if num < 10 { deleted() } else { num - 10 }) root.has_good_taste = ["pikachu","mewtwo","magmar"].contains(this.user.fav_pokemon) # In: {"thing":{"id":123,"nums":[5,12,18,7,25]},"user":{"fav_pokemon":"pikachu"}} ``` Methods support both named and nameless style arguments: ```bloblang root.foo_one = this.(bar | baz).trim().replace_all(old: "dog", new: "cat") root.foo_two = this.(bar | baz).trim().replace_all("dog", "cat") # In: {"bar":" I love my dog "} ``` ## [](#general)General ### [](#apply)apply Apply a declared mapping to a target value. #### [](#parameters)Parameters | Name | Type | Description | | --- | --- | --- | | mapping | string | The mapping to apply. | #### [](#examples)Examples ```bloblang map thing { root.inner = this.first } root.foo = this.doc.apply("thing") # In: {"doc":{"first":"hello world"}} # Out: {"foo":{"inner":"hello world"}} ``` ```bloblang map create_foo { root.name = "a foo" root.purpose = "to be a foo" } root = this root.foo = null.apply("create_foo") # In: {"id":"1234"} # Out: {"foo":{"name":"a foo","purpose":"to be a foo"},"id":"1234"} ``` ### [](#catch)catch If the result of a target query fails (due to incorrect types, failed parsing, etc) the argument is returned instead. #### [](#parameters-2)Parameters | Name | Type | Description | | --- | --- | --- | | fallback | query expression | A value to yield, or query to execute, if the target query fails. | #### [](#examples-2)Examples ```bloblang root.doc.id = this.thing.id.string().catch(uuid_v4()) ``` The fallback argument can be a mapping, allowing you to capture the error string and yield structured data back: ```bloblang root.url = this.url.parse_url().catch(err -> {"error":err,"input":this.url}) # In: {"url":"invalid %&# url"} # Out: {"url":{"error":"field `this.url`: parse \"invalid %&\": invalid URL escape \"%&\"","input":"invalid %&# url"}} ``` When the input document is not structured attempting to reference structured fields with `this` will result in an error. Therefore, a convenient way to delete non-structured data is with a catch: ```bloblang root = this.catch(deleted()) # In: {"doc":{"foo":"bar"}} # Out: {"doc":{"foo":"bar"}} # In: not structured data # Out: ``` ### [](#from)from Modifies a target query such that certain functions are executed from the perspective of another message in the batch. This allows you to mutate events based on the contents of other messages. Functions that support this behavior are `content`, `json` and `meta`. #### [](#parameters-3)Parameters | Name | Type | Description | | --- | --- | --- | | index | integer | The message index to use as a perspective. | #### [](#examples-3)Examples For example, the following map extracts the contents of the JSON field `foo` specifically from message index `1` of a batch, effectively overriding the field `foo` for all messages of a batch to that of message 1: ```bloblang root = this root.foo = json("foo").from(1) ``` ### [](#from_all)from_all Modifies a target query such that certain functions are executed from the perspective of each message in the batch, and returns the set of results as an array. Functions that support this behavior are `content`, `json` and `meta`. #### [](#examples-4)Examples ```bloblang root = this root.foo_summed = json("foo").from_all().sum() ``` ### [](#map)map Executes a query on the target value, allowing you to transform or extract data from the current context. #### [](#parameters-4)Parameters | Name | Type | Description | | --- | --- | --- | | query | query expression | A query to execute on the target. | ### [](#not)not Returns the logical NOT (negation) of a boolean value. Converts true to false and false to true. ### [](#or)or If the result of the target query fails or resolves to `null`, returns the argument instead. This is an explicit method alternative to the coalesce pipe operator `|`. #### [](#parameters-5)Parameters | Name | Type | Description | | --- | --- | --- | | fallback | query expression | A value to yield, or query to execute, if the target query fails or resolves to null. | #### [](#examples-5)Examples ```bloblang root.doc.id = this.thing.id.or(uuid_v4()) ``` ## [](#encoding-and-encryption)Encoding and encryption ### [](#compress)compress Compresses a string or byte array using the specified compression algorithm. Returns compressed data as bytes. Useful for reducing payload size before transmission or storage. #### [](#parameters-6)Parameters | Name | Type | Description | | --- | --- | --- | | algorithm | string | The compression algorithm: flate, gzip, pgzip (parallel gzip), lz4, snappy, zlib, or zstd. | | level | integer | Compression level (default: -1 for default compression). Higher values increase compression ratio but use more CPU. Range and effect varies by algorithm. | #### [](#examples-6)Examples Compress and encode for safe transmission: ```bloblang root.compressed = content().bytes().compress("gzip").encode("base64") # In: {"message":"hello world I love space"} # Out: {"compressed":"H4sIAAAJbogA/wAmANn/eyJtZXNzYWdlIjoiaGVsbG8gd29ybGQgSSBsb3ZlIHNwYWNlIn0DAHEvdwomAAAA"} ``` Compare compression ratios across algorithms: ```bloblang root.original_size = content().length() root.gzip_size = content().compress("gzip").length() root.lz4_size = content().compress("lz4").length() # In: The quick brown fox jumps over the lazy dog. The quick brown fox jumps over the lazy dog. # Out: {"gzip_size":114,"lz4_size":85,"original_size":89} ``` ### [](#decode)decode Decodes an encoded string according to a chosen scheme. #### [](#parameters-7)Parameters | Name | Type | Description | | --- | --- | --- | | scheme | string | The decoding scheme to use. | #### [](#examples-7)Examples ```bloblang root.decoded = this.value.decode("hex").string() # In: {"value":"68656c6c6f20776f726c64"} # Out: {"decoded":"hello world"} ``` ```bloblang root = this.encoded.decode("ascii85") # In: {"encoded":"FD,B0+DGm>FDl80Ci\"A>F`)8BEckl6F`M&(+Cno&@/"} # Out: this is totally unstructured data ``` ### [](#decompress)decompress Decompresses a byte array using the specified decompression algorithm. Returns decompressed data as bytes. Use with data that was previously compressed using the corresponding algorithm. #### [](#parameters-8)Parameters | Name | Type | Description | | --- | --- | --- | | algorithm | string | The decompression algorithm: gzip, pgzip (parallel gzip), zlib, bzip2, flate, snappy, lz4, or zstd. | #### [](#examples-8)Examples Decompress base64-encoded compressed data: ```bloblang root = this.compressed.decode("base64").decompress("gzip") # In: {"compressed":"H4sIAN12MWkAA8tIzcnJVyjPL8pJUfBUyMkvS1UoLkhMTgUAQpDxbxgAAAA="} # Out: hello world I love space ``` Convert decompressed bytes to string for JSON output: ```bloblang root.message = this.compressed.decode("base64").decompress("gzip").string() # In: {"compressed":"H4sIAN12MWkAA8tIzcnJVyjPL8pJUfBUyMkvS1UoLkhMTgUAQpDxbxgAAAA="} # Out: {"message":"hello world I love space"} ``` ### [](#decrypt_aes)decrypt_aes Decrypts an AES-encrypted string or byte array. #### [](#parameters-9)Parameters | Name | Type | Description | | --- | --- | --- | | scheme | string | The scheme to use for decryption, one of ctr, gcm, ofb, cbc. | | key | string | A key to decrypt with. | | iv | string | An initialization vector / nonce. | #### [](#examples-9)Examples ```bloblang let key = "2b7e151628aed2a6abf7158809cf4f3c".decode("hex") let vector = "f0f1f2f3f4f5f6f7f8f9fafbfcfdfeff".decode("hex") root.decrypted = this.value.decode("hex").decrypt_aes("ctr", $key, $vector).string() # In: {"value":"84e9b31ff7400bdf80be7254"} # Out: {"decrypted":"hello world!"} ``` ### [](#encode)encode Encodes a string or byte array according to a chosen scheme. #### [](#parameters-10)Parameters | Name | Type | Description | | --- | --- | --- | | scheme | string | The encoding scheme to use. | #### [](#examples-10)Examples ```bloblang root.encoded = this.value.encode("hex") # In: {"value":"hello world"} # Out: {"encoded":"68656c6c6f20776f726c64"} ``` ```bloblang root.encoded = content().encode("ascii85") # In: this is totally unstructured data # Out: {"encoded":"FD,B0+DGm>FDl80Ci\"A>F`)8BEckl6F`M&(+Cno&@/"} ``` ### [](#encrypt_aes)encrypt_aes Encrypts a string or byte array using AES encryption. #### [](#parameters-11)Parameters | Name | Type | Description | | --- | --- | --- | | scheme | string | The scheme to use for encryption, one of ctr, gcm, ofb, cbc. | | key | string | A key to encrypt with. | | iv | string | An initialization vector / nonce. | #### [](#examples-11)Examples ```bloblang let key = "2b7e151628aed2a6abf7158809cf4f3c".decode("hex") let vector = "f0f1f2f3f4f5f6f7f8f9fafbfcfdfeff".decode("hex") root.encrypted = this.value.encrypt_aes("ctr", $key, $vector).encode("hex") # In: {"value":"hello world!"} # Out: {"encrypted":"84e9b31ff7400bdf80be7254"} ``` ### [](#hash)hash Hashes a string or byte array using a specified algorithm. #### [](#parameters-12)Parameters | Name | Type | Description | | --- | --- | --- | | algorithm | string | The hashing algorithm to use. | | key (optional) | string | An optional key to use. | | polynomial | string | An optional polynomial key to use when selecting the crc32 algorithm, otherwise ignored. Options are IEEE (default), Castagnoli and Koopman | #### [](#examples-12)Examples ```bloblang root.h1 = this.value.hash("sha1").encode("hex") root.h2 = this.value.hash("hmac_sha1","static-key").encode("hex") # In: {"value":"hello world"} # Out: {"h1":"2aae6c35c94fcfb415dbe95f408b9ce91ee846ed","h2":"d87e5f068fa08fe90bb95bc7c8344cb809179d76"} ``` The `crc32` algorithm supports options for the polynomial: ```bloblang root.h1 = this.value.hash(algorithm: "crc32", polynomial: "Castagnoli").encode("hex") root.h2 = this.value.hash(algorithm: "crc32", polynomial: "Koopman").encode("hex") # In: {"value":"hello world"} # Out: {"h1":"c99465aa","h2":"df373d3c"} ``` ### [](#uuid_v5)uuid_v5 Generates a version 5 UUID from a namespace and name. #### [](#parameters-13)Parameters | Name | Type | Description | | --- | --- | --- | | ns (optional) | string | An optional namespace name or UUID. It supports the dns, url, oid and x500 predefined namespaces and any valid RFC-9562 UUID. If empty, the nil UUID will be used. | #### [](#examples-13)Examples ```bloblang root.id = "example".uuid_v5() ``` ```bloblang root.id = "example".uuid_v5("x500") ``` ```bloblang root.id = "example".uuid_v5("77f836b7-9f61-46c0-851e-9b6ca3535e69") ``` ## [](#geoip)GeoIP ### [](#geoip_anonymous_ip)geoip_anonymous_ip Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the anonymous IP associated with it. #### [](#parameters-14)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ### [](#geoip_asn)geoip_asn Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the ASN associated with it. #### [](#parameters-15)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ### [](#geoip_city)geoip_city Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the city associated with it. #### [](#parameters-16)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ### [](#geoip_connection_type)geoip_connection_type Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the connection type associated with it. #### [](#parameters-17)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ### [](#geoip_country)geoip_country Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the country associated with it. #### [](#parameters-18)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ### [](#geoip_domain)geoip_domain Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the domain associated with it. #### [](#parameters-19)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ### [](#geoip_enterprise)geoip_enterprise Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the enterprise associated with it. #### [](#parameters-20)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ### [](#geoip_isp)geoip_isp Looks up an IP address against a [MaxMind database file](https://www.maxmind.com/en/home) and, if found, returns an object describing the ISP associated with it. #### [](#parameters-21)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A path to an mmdb (maxmind) file. | ## [](#json-web-tokens)JSON web tokens ### [](#parse_jwt_es256)parse_jwt_es256 Parses a claims object from a JWT string encoded with ES256. This method does not validate JWT claims. #### [](#parameters-22)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The ES256 secret that was used for signing the token. | #### [](#examples-14)Examples ```bloblang root.claims = this.signed.parse_jwt_es256("""-----BEGIN PUBLIC KEY----- MFkwEwYHKoZIzj0CAQYIKoZIzj0DAQcDQgAEGtLqIBePHmIhQcf0JLgc+F/4W/oI dp0Gta53G35VerNDgUUXmp78J2kfh4qLdh0XtmOMI587tCaqjvDAXfs//w== -----END PUBLIC KEY-----""") # In: {"signed":"eyJhbGciOiJFUzI1NiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.GIRajP9JJbpTlqSCdNEz4qpQkRvzX4Q51YnTwVyxLDM9tKjR_a8ggHWn9CWj7KG0x8J56OWtmUxn112SRTZVhQ"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_es384)parse_jwt_es384 Parses a claims object from a JWT string encoded with ES384. This method does not validate JWT claims. #### [](#parameters-23)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The ES384 secret that was used for signing the token. | #### [](#examples-15)Examples ```bloblang root.claims = this.signed.parse_jwt_es384("""-----BEGIN PUBLIC KEY----- MHYwEAYHKoZIzj0CAQYFK4EEACIDYgAERoz74/B6SwmLhs8X7CWhnrWyRrB13AuU 8OYeqy0qHRu9JWNw8NIavqpTmu6XPT4xcFanYjq8FbeuM11eq06C52mNmS4LLwzA 2imlFEgn85bvJoC3bnkuq4mQjwt9VxdH -----END PUBLIC KEY-----""") # In: {"signed":"eyJhbGciOiJFUzM4NCIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.H2HBSlrvQBaov2tdreGonbBexxtQB-xzaPL4-tNQZ6TVh7VH8VBcSwcWHYa1lBAHmdsKOFcB2Wk0SB7QWeGT3ptSgr-_EhDMaZ8bA5spgdpq5DsKfaKHrd7DbbQlmxNq"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_es512)parse_jwt_es512 Parses a claims object from a JWT string encoded with ES512. This method does not validate JWT claims. #### [](#parameters-24)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The ES512 secret that was used for signing the token. | #### [](#examples-16)Examples ```bloblang root.claims = this.signed.parse_jwt_es512("""-----BEGIN PUBLIC KEY----- MIGbMBAGByqGSM49AgEGBSuBBAAjA4GGAAQAkHLdts9P56fFkyhpYQ31M/Stwt3w vpaxhlfudxnXgTO1IP4RQRgryRxZ19EUzhvWDcG3GQIckoNMY5PelsnCGnIBT2Xh 9NQkjWF5K6xS4upFsbGSAwQ+GIyyk5IPJ2LHgOyMSCVh5gRZXV3CZLzXujx/umC9 UeYyTt05zRRWuD+p5bY= -----END PUBLIC KEY-----""") # In: {"signed":"eyJhbGciOiJFUzUxMiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.ACrpLuU7TKpAnncDCpN9m85nkL55MJ45NFOBl6-nEXmNT1eIxWjiP4pwWVbFH9et_BgN14119jbL_KqEJInPYc9nAXC6dDLq0aBU-dalvNl4-O5YWpP43-Y-TBGAsWnbMTrchILJ4-AEiICe73Ck5yWPleKg9c3LtkEFWfGs7BoPRguZ"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_hs256)parse_jwt_hs256 Parses a claims object from a JWT string encoded with HS256. This method does not validate JWT claims. #### [](#parameters-25)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The HS256 secret that was used for signing the token. | #### [](#examples-17)Examples ```bloblang root.claims = this.signed.parse_jwt_hs256("""dont-tell-anyone""") # In: {"signed":"eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.YwXOM8v3gHVWcQRRRQc_zDlhmLnM62fwhFYGpiA0J1A"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_hs384)parse_jwt_hs384 Parses a claims object from a JWT string encoded with HS384. This method does not validate JWT claims. #### [](#parameters-26)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The HS384 secret that was used for signing the token. | #### [](#examples-18)Examples ```bloblang root.claims = this.signed.parse_jwt_hs384("""dont-tell-anyone""") # In: {"signed":"eyJhbGciOiJIUzM4NCIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.2Y8rf_ijwN4t8hOGGViON_GrirLkCQVbCOuax6EoZ3nluX0tCGezcJxbctlIfsQ2"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_hs512)parse_jwt_hs512 Parses a claims object from a JWT string encoded with HS512. This method does not validate JWT claims. #### [](#parameters-27)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The HS512 secret that was used for signing the token. | #### [](#examples-19)Examples ```bloblang root.claims = this.signed.parse_jwt_hs512("""dont-tell-anyone""") # In: {"signed":"eyJhbGciOiJIUzUxMiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.utRb0urG6LGGyranZJVo5Dk0Fns1QNcSUYPN0TObQ-YzsGGB8jrxHwM5NAJccjJZzKectEUqmmKCaETZvuX4Fg"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_rs256)parse_jwt_rs256 Parses a claims object from a JWT string encoded with RS256. This method does not validate JWT claims. #### [](#parameters-28)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The RS256 secret that was used for signing the token. | #### [](#examples-20)Examples ```bloblang root.claims = this.signed.parse_jwt_rs256("""-----BEGIN PUBLIC KEY----- MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAs/ibN8r68pLMR6gRzg4S 8v8l6Q7yi8qURjkEbcNeM1rkokC7xh0I4JVTwxYSVv/JIW8qJdyspl5NIfuAVi32 WfKvSAs+NIs+DMsNPYw3yuQals4AX8hith1YDvYpr8SD44jxhz/DR9lYKZFGhXGB +7NqQ7vpTWp3BceLYocazWJgusZt7CgecIq57ycM5hjM93BvlrUJ8nQ1a46wfL/8 Cy4P0et70hzZrsjjN41KFhKY0iUwlyU41yEiDHvHDDsTMBxAZosWjSREGfJL6Mfp XOInTHs/Gg6DZMkbxjQu6L06EdJ+Q/NwglJdAXM7Zo9rNELqRig6DdvG5JesdMsO +QIDAQAB -----END PUBLIC KEY-----""") # In: {"signed":"eyJhbGciOiJSUzI1NiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.b0lH3jEupZZ4zoaly4Y_GCvu94HH6UKdKY96zfGNsIkPZpQLHIkZ7jMWlLlNOAd8qXlsBGP_i8H2qCKI4zlWJBGyPZgxXDzNRPVrTDfFpn4t4nBcA1WK2-ntXP3ehQxsaHcQU8Z_nsogId7Pme5iJRnoHWEnWtbwz5DLSXL3ZZNnRdrHM9MdI7QSDz9mojKDCaMpGN9sG7Xl-tGdBp1XzXuUOzG8S03mtZ1IgVR1uiBL2N6oohHIAunk8DIAmNWI-zgycTgzUGU7mvPkKH43qO8Ua1-13tCUBKKa8VxcotZ67Mxm1QAvBGoDnTKwWMwghLzs6d6WViXQg6eWlJcpBA"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_rs384)parse_jwt_rs384 Parses a claims object from a JWT string encoded with RS384. This method does not validate JWT claims. #### [](#parameters-29)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The RS384 secret that was used for signing the token. | #### [](#examples-21)Examples ```bloblang root.claims = this.signed.parse_jwt_rs384("""-----BEGIN PUBLIC KEY----- MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAs/ibN8r68pLMR6gRzg4S 8v8l6Q7yi8qURjkEbcNeM1rkokC7xh0I4JVTwxYSVv/JIW8qJdyspl5NIfuAVi32 WfKvSAs+NIs+DMsNPYw3yuQals4AX8hith1YDvYpr8SD44jxhz/DR9lYKZFGhXGB +7NqQ7vpTWp3BceLYocazWJgusZt7CgecIq57ycM5hjM93BvlrUJ8nQ1a46wfL/8 Cy4P0et70hzZrsjjN41KFhKY0iUwlyU41yEiDHvHDDsTMBxAZosWjSREGfJL6Mfp XOInTHs/Gg6DZMkbxjQu6L06EdJ+Q/NwglJdAXM7Zo9rNELqRig6DdvG5JesdMsO +QIDAQAB -----END PUBLIC KEY-----""") # In: {"signed":"eyJhbGciOiJSUzM4NCIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.orcXYBcjVE5DU7mvq4KKWFfNdXR4nEY_xupzWoETRpYmQZIozlZnM_nHxEk2dySvpXlAzVm7kgOPK2RFtGlOVaNRIa3x-pMMr-bhZTno4L8Hl4sYxOks3bWtjK7wql4uqUbqThSJB12psAXw2-S-I_FMngOPGIn4jDT9b802ottJSvTpXcy0-eKTjrV2PSkRRu-EYJh0CJZW55MNhqlt6kCGhAXfbhNazN3ASX-dmpd_JixyBKphrngr_zRA-FCn_Xf3QQDA-5INopb4Yp5QiJ7UxVqQEKI80X_JvJqz9WE1qiAw8pq5-xTen1t7zTP-HT1NbbD3kltcNa3G8acmNg"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#parse_jwt_rs512)parse_jwt_rs512 Parses a claims object from a JWT string encoded with RS512. This method does not validate JWT claims. #### [](#parameters-30)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The RS512 secret that was used for signing the token. | #### [](#examples-22)Examples ```bloblang root.claims = this.signed.parse_jwt_rs512("""-----BEGIN PUBLIC KEY----- MIIBIjANBgkqhkiG9w0BAQEFAAOCAQ8AMIIBCgKCAQEAs/ibN8r68pLMR6gRzg4S 8v8l6Q7yi8qURjkEbcNeM1rkokC7xh0I4JVTwxYSVv/JIW8qJdyspl5NIfuAVi32 WfKvSAs+NIs+DMsNPYw3yuQals4AX8hith1YDvYpr8SD44jxhz/DR9lYKZFGhXGB +7NqQ7vpTWp3BceLYocazWJgusZt7CgecIq57ycM5hjM93BvlrUJ8nQ1a46wfL/8 Cy4P0et70hzZrsjjN41KFhKY0iUwlyU41yEiDHvHDDsTMBxAZosWjSREGfJL6Mfp XOInTHs/Gg6DZMkbxjQu6L06EdJ+Q/NwglJdAXM7Zo9rNELqRig6DdvG5JesdMsO +QIDAQAB -----END PUBLIC KEY-----""") # In: {"signed":"eyJhbGciOiJSUzUxMiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.rsMp_X5HMrUqKnZJIxo27aAoscovRA6SSQYR9rq7pifIj0YHXxMyNyOBDGnvVALHKTi25VUGHpfNUW0VVMmae0A4t_ObNU6hVZHguWvetKZZq4FZpW1lgWHCMqgPGwT5_uOqwYCH6r8tJuZT3pqXeL0CY4putb1AN2w6CVp620nh3l8d3XWb4jaifycd_4CEVCqHuWDmohfug4VhmoVKlIXZkYoAQowgHlozATDssBSWdYtv107Wd2AzEoiXPu6e3pflsuXULlyqQnS4ELEKPYThFLafh1NqvZDPddqozcPZ-iODBW-xf3A4DYDdivnMYLrh73AZOGHexxu8ay6nDA"} # Out: {"claims":{"iat":1516239022,"mood":"Disdainful","sub":"1234567890"}} ``` ### [](#sign_jwt_es256)sign_jwt_es256 Hash and sign an object representing JSON Web Token (JWT) claims using ES256. #### [](#parameters-31)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-23)Examples ```bloblang root.signed = this.claims.sign_jwt_es256("""-----BEGIN EC PRIVATE KEY----- ... signature data ... -----END EC PRIVATE KEY-----""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJFUzI1NiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.-8LrOdkEiv_44ADWW08lpbq41ZmHCel58NMORPq1q4Dyw0zFhqDVLrRoSvCvuyyvgXAFb9IHfR-9MlJ_2ShA9A"} ``` ```bloblang root.signed = this.claims.sign_jwt_es256(signing_secret: """-----BEGIN EC PRIVATE KEY----- ... signature data ... -----END EC PRIVATE KEY-----""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_es384)sign_jwt_es384 Hash and sign an object representing JSON Web Token (JWT) claims using ES384. #### [](#parameters-32)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-24)Examples ```bloblang root.signed = this.claims.sign_jwt_es384("""-----BEGIN EC PRIVATE KEY----- ... signature data ... -----END EC PRIVATE KEY-----""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJFUzM4NCIsInR5cCI6IkpXVCJ9.eyJzdWIiOiJ1c2VyMTIzIn0.8FmTKH08dl7dyxrNu0rmvhegiIBCy-O9cddGco2e9lpZtgv5mS5qHgPkgBC5eRw1d7SRJsHwHZeehzdqT5Ba7aZJIhz9ds0sn37YQ60L7jT0j2gxCzccrt4kECHnUnLw"} ``` ```bloblang root.signed = this.claims.sign_jwt_es384(signing_secret: """-----BEGIN EC PRIVATE KEY----- ... signature data ... -----END EC PRIVATE KEY-----""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_es512)sign_jwt_es512 Hash and sign an object representing JSON Web Token (JWT) claims using ES512. #### [](#parameters-33)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-25)Examples ```bloblang root.signed = this.claims.sign_jwt_es512("""-----BEGIN EC PRIVATE KEY----- ... signature data ... -----END EC PRIVATE KEY-----""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJFUzUxMiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiJ1c2VyMTIzIn0.AQbEWymoRZxDJEJtKSFFG2k2VbDCTYSuBwAZyMqexCspr3If8aERTVGif8HXG3S7TzMBCCzxkcKr3eIU441l3DlpAMNfQbkcOlBqMvNBn-CX481WyKf3K5rFHQ-6wRonz05aIsWAxCDvAozI_9J0OWllxdQ2MBAuTPbPJ38OqXsYkCQs"} ``` ```bloblang root.signed = this.claims.sign_jwt_es512(signing_secret: """-----BEGIN EC PRIVATE KEY----- ... signature data ... -----END EC PRIVATE KEY-----""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_hs256)sign_jwt_hs256 Hash and sign an object representing JSON Web Token (JWT) claims using HS256. #### [](#parameters-34)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-26)Examples ```bloblang root.signed = this.claims.sign_jwt_hs256("""dont-tell-anyone""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiJ1c2VyMTIzIn0.hUl-nngPMY_3h9vveWJUPsCcO5PeL6k9hWLnMYeFbFQ"} ``` ```bloblang root.signed = this.claims.sign_jwt_hs256(signing_secret: """dont-tell-anyone""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_hs384)sign_jwt_hs384 Hash and sign an object representing JSON Web Token (JWT) claims using HS384. #### [](#parameters-35)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-27)Examples ```bloblang root.signed = this.claims.sign_jwt_hs384("""dont-tell-anyone""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJIUzM4NCIsInR5cCI6IkpXVCJ9.eyJzdWIiOiJ1c2VyMTIzIn0.zGYLr83aToon1efUNq-hw7XgT20lPvZb8sYei8x6S6mpHwb433SJdXJXx0Oio8AZ"} ``` ```bloblang root.signed = this.claims.sign_jwt_hs384(signing_secret: """dont-tell-anyone""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_hs512)sign_jwt_hs512 Hash and sign an object representing JSON Web Token (JWT) claims using HS512. #### [](#parameters-36)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-28)Examples ```bloblang root.signed = this.claims.sign_jwt_hs512("""dont-tell-anyone""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJIUzUxMiIsInR5cCI6IkpXVCJ9.eyJzdWIiOiJ1c2VyMTIzIn0.zBNR9o_6EDwXXKkpKLNJhG26j8Dc-mV-YahBwmEdCrmiWt5les8I9rgmNlWIowpq6Yxs4kLNAdFhqoRz3NXT3w"} ``` ```bloblang root.signed = this.claims.sign_jwt_hs512(signing_secret: """dont-tell-anyone""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_rs256)sign_jwt_rs256 Hash and sign an object representing JSON Web Token (JWT) claims using RS256. #### [](#parameters-37)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-29)Examples ```bloblang root.signed = this.claims.sign_jwt_rs256("""-----BEGIN RSA PRIVATE KEY----- ... signature data ... -----END RSA PRIVATE KEY-----""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJSUzI1NiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.b0lH3jEupZZ4zoaly4Y_GCvu94HH6UKdKY96zfGNsIkPZpQLHIkZ7jMWlLlNOAd8qXlsBGP_i8H2qCKI4zlWJBGyPZgxXDzNRPVrTDfFpn4t4nBcA1WK2-ntXP3ehQxsaHcQU8Z_nsogId7Pme5iJRnoHWEnWtbwz5DLSXL3ZZNnRdrHM9MdI7QSDz9mojKDCaMpGN9sG7Xl-tGdBp1XzXuUOzG8S03mtZ1IgVR1uiBL2N6oohHIAunk8DIAmNWI-zgycTgzUGU7mvPkKH43qO8Ua1-13tCUBKKa8VxcotZ67Mxm1QAvBGoDnTKwWMwghLzs6d6WViXQg6eWlJcpBA"} ``` ```bloblang root.signed = this.claims.sign_jwt_rs256(signing_secret: """-----BEGIN RSA PRIVATE KEY----- ... signature data ... -----END RSA PRIVATE KEY-----""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_rs384)sign_jwt_rs384 Hash and sign an object representing JSON Web Token (JWT) claims using RS384. #### [](#parameters-38)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-30)Examples ```bloblang root.signed = this.claims.sign_jwt_rs384("""-----BEGIN RSA PRIVATE KEY----- ... signature data ... -----END RSA PRIVATE KEY-----""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJSUzM4NCIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.orcXYBcjVE5DU7mvq4KKWFfNdXR4nEY_xupzWoETRpYmQZIozlZnM_nHxEk2dySvpXlAzVm7kgOPK2RFtGlOVaNRIa3x-pMMr-bhZTno4L8Hl4sYxOks3bWtjK7wql4uqUbqThSJB12psAXw2-S-I_FMngOPGIn4jDT9b802ottJSvTpXcy0-eKTjrV2PSkRRu-EYJh0CJZW55MNhqlt6kCGhAXfbhNazN3ASX-dmpd_JixyBKphrngr_zRA-FCn_Xf3QQDA-5INopb4Yp5QiJ7UxVqQEKI80X_JvJqz9WE1qiAw8pq5-xTen1t7zTP-HT1NbbD3kltcNa3G8acmNg"} ``` ```bloblang root.signed = this.claims.sign_jwt_rs384(signing_secret: """-----BEGIN RSA PRIVATE KEY----- ... signature data ... -----END RSA PRIVATE KEY-----""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ### [](#sign_jwt_rs512)sign_jwt_rs512 Hash and sign an object representing JSON Web Token (JWT) claims using RS512. #### [](#parameters-39)Parameters | Name | Type | Description | | --- | --- | --- | | signing_secret | string | The secret to use for signing the token. | | headers (optional) | unknown | Optional object of JWT header fields to include in the token. Keys "alg", "typ", "jku", "jwk", "x5u", "x5c", "x5t","x5t#S256" and "crit" will be ignored if provided. | #### [](#examples-31)Examples ```bloblang root.signed = this.claims.sign_jwt_rs512("""-----BEGIN RSA PRIVATE KEY----- ... signature data ... -----END RSA PRIVATE KEY-----""") # In: {"claims":{"sub":"user123"}} # Out: {"signed":"eyJhbGciOiJSUzUxMiIsInR5cCI6IkpXVCJ9.eyJpYXQiOjE1MTYyMzkwMjIsIm1vb2QiOiJEaXNkYWluZnVsIiwic3ViIjoiMTIzNDU2Nzg5MCJ9.rsMp_X5HMrUqKnZJIxo27aAoscovRA6SSQYR9rq7pifIj0YHXxMyNyOBDGnvVALHKTi25VUGHpfNUW0VVMmae0A4t_ObNU6hVZHguWvetKZZq4FZpW1lgWHCMqgPGwT5_uOqwYCH6r8tJuZT3pqXeL0CY4putb1AN2w6CVp620nh3l8d3XWb4jaifycd_4CEVCqHuWDmohfug4VhmoVKlIXZkYoAQowgHlozATDssBSWdYtv107Wd2AzEoiXPu6e3pflsuXULlyqQnS4ELEKPYThFLafh1NqvZDPddqozcPZ-iODBW-xf3A4DYDdivnMYLrh73AZOGHexxu8ay6nDA"} ``` ```bloblang root.signed = this.claims.sign_jwt_rs512(signing_secret: """-----BEGIN RSA PRIVATE KEY----- ... signature data ... -----END RSA PRIVATE KEY-----""", headers: {"kid": "my-key", "x": "y"}) # In: {"claims":{"sub":"user123"}} # Out: {"signed":""} ``` ## [](#number-manipulation)Number manipulation ### [](#abs)abs Returns the absolute value of an int64 or float64 number. As a special case, when an integer is provided that is the minimum value it is converted to the maximum value. #### [](#examples-32)Examples ```bloblang root.outs = this.ins.map_each(ele -> ele.abs()) # In: {"ins":[9,-18,1.23,-4.56]} # Out: {"outs":[9,18,1.23,4.56]} ``` ### [](#bitwise_and)bitwise_and Performs a bitwise AND operation between the integer and the specified value. #### [](#parameters-40)Parameters | Name | Type | Description | | --- | --- | --- | | value | integer | The value to AND with | #### [](#examples-33)Examples ```bloblang root.new_value = this.value.bitwise_and(6) # In: {"value":12} # Out: {"new_value":4} ``` ```bloblang root.masked = this.flags.bitwise_and(15) # In: {"flags":127} # Out: {"masked":15} ``` ### [](#bitwise_or)bitwise_or Performs a bitwise OR operation between the integer and the specified value. #### [](#parameters-41)Parameters | Name | Type | Description | | --- | --- | --- | | value | integer | The value to OR with | #### [](#examples-34)Examples ```bloblang root.new_value = this.value.bitwise_or(6) # In: {"value":12} # Out: {"new_value":14} ``` ```bloblang root.combined = this.flags.bitwise_or(8) # In: {"flags":4} # Out: {"combined":12} ``` ### [](#bitwise_xor)bitwise_xor Performs a bitwise XOR (exclusive OR) operation between the integer and the specified value. #### [](#parameters-42)Parameters | Name | Type | Description | | --- | --- | --- | | value | integer | The value to XOR with | #### [](#examples-35)Examples ```bloblang root.new_value = this.value.bitwise_xor(6) # In: {"value":12} # Out: {"new_value":10} ``` ```bloblang root.toggled = this.flags.bitwise_xor(5) # In: {"flags":3} # Out: {"toggled":6} ``` ### [](#ceil)ceil Rounds a number up to the nearest integer. Returns an integer if the result fits in 64-bit, otherwise returns a float. #### [](#examples-36)Examples ```bloblang root.new_value = this.value.ceil() # In: {"value":5.3} # Out: {"new_value":6} # In: {"value":-5.9} # Out: {"new_value":-5} ``` ```bloblang root.result = this.price.ceil() # In: {"price":19.99} # Out: {"result":20} ``` ### [](#cos)cos Calculates the cosine of a given angle specified in radians. #### [](#examples-37)Examples ```bloblang root.new_value = (this.value * (pi() / 180)).cos() # In: {"value":45} # Out: {"new_value":0.7071067811865476} # In: {"value":0} # Out: {"new_value":1} # In: {"value":180} # Out: {"new_value":-1} ``` ### [](#float32)float32 Converts a numerical type into a 32-bit floating point number, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 32-bit floating point number. Please refer to the [`strconv.ParseFloat` documentation](https://pkg.go.dev/strconv#ParseFloat) for details regarding the supported formats. #### [](#examples-38)Examples ```bloblang root.out = this.in.float32() # In: {"in":"6.674282313423543523453425345e-11"} # Out: {"out":6.674283e-11} ``` ### [](#float64)float64 Converts a numerical type into a 64-bit floating point number, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 64-bit floating point number. Please refer to the [`strconv.ParseFloat` documentation](https://pkg.go.dev/strconv#ParseFloat) for details regarding the supported formats. #### [](#examples-39)Examples ```bloblang root.out = this.in.float64() # In: {"in":"6.674282313423543523453425345e-11"} # Out: {"out":6.674282313423544e-11} ``` ### [](#floor)floor Rounds a number down to the nearest integer. Returns an integer if the result fits in 64-bit, otherwise returns a float. #### [](#examples-40)Examples ```bloblang root.new_value = this.value.floor() # In: {"value":5.7} # Out: {"new_value":5} # In: {"value":-3.2} # Out: {"new_value":-4} ``` ```bloblang root.whole_seconds = this.duration_seconds.floor() # In: {"duration_seconds":12.345} # Out: {"whole_seconds":12} ``` ### [](#int16)int16 Converts a numerical type into a 16-bit signed integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 16-bit signed integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-41)Examples ```bloblang root.a = this.a.int16() root.b = this.b.round().int16() root.c = this.c.int16() root.d = this.d.int16().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":-12} ``` ```bloblang root = this.int16() # In: "0xDE" # Out: 222 ``` ### [](#int32)int32 Converts a numerical type into a 32-bit signed integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 32-bit signed integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-42)Examples ```bloblang root.a = this.a.int32() root.b = this.b.round().int32() root.c = this.c.int32() root.d = this.d.int32().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":-12} ``` ```bloblang root = this.int32() # In: "0xDEAD" # Out: 57005 ``` ### [](#int64)int64 Converts a numerical type into a 64-bit signed integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 64-bit signed integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-43)Examples ```bloblang root.a = this.a.int64() root.b = this.b.round().int64() root.c = this.c.int64() root.d = this.d.int64().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":-12} ``` ```bloblang root = this.int64() # In: "0xDEADBEEF" # Out: 3735928559 ``` ### [](#int8)int8 Converts a numerical type into a 8-bit signed integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 8-bit signed integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-44)Examples ```bloblang root.a = this.a.int8() root.b = this.b.round().int8() root.c = this.c.int8() root.d = this.d.int8().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":-12} ``` ```bloblang root = this.int8() # In: "0xD" # Out: 13 ``` ### [](#log)log Calculates the natural logarithm (base e) of a number. #### [](#examples-45)Examples ```bloblang root.new_value = this.value.log().round() # In: {"value":1} # Out: {"new_value":0} # In: {"value":2.7183} # Out: {"new_value":1} ``` ```bloblang root.ln_result = this.number.log() # In: {"number":10} # Out: {"ln_result":2.302585092994046} ``` ### [](#log10)log10 Calculates the base-10 logarithm of a number. #### [](#examples-46)Examples ```bloblang root.new_value = this.value.log10() # In: {"value":100} # Out: {"new_value":2} # In: {"value":1000} # Out: {"new_value":3} ``` ```bloblang root.log_value = this.magnitude.log10() # In: {"magnitude":10000} # Out: {"log_value":4} ``` ### [](#max)max Returns the largest number from an array. All elements must be numbers and the array cannot be empty. #### [](#examples-47)Examples ```bloblang root.biggest = this.values.max() # In: {"values":[0,3,2.5,7,5]} # Out: {"biggest":7} ``` ```bloblang root.highest_temp = this.temperatures.max() # In: {"temperatures":[20.5,22.1,19.8,23.4]} # Out: {"highest_temp":23.4} ``` ### [](#min)min Returns the smallest number from an array. All elements must be numbers and the array cannot be empty. #### [](#examples-48)Examples ```bloblang root.smallest = this.values.min() # In: {"values":[0,3,-2.5,7,5]} # Out: {"smallest":-2.5} ``` ```bloblang root.lowest_temp = this.temperatures.min() # In: {"temperatures":[20.5,22.1,19.8,23.4]} # Out: {"lowest_temp":19.8} ``` ### [](#pow)pow Returns the number raised to the specified exponent. #### [](#parameters-43)Parameters | Name | Type | Description | | --- | --- | --- | | exponent | float | The exponent you want to raise to the power of. | #### [](#examples-49)Examples ```bloblang root.new_value = this.value * 10.pow(-2) # In: {"value":2} # Out: {"new_value":0.02} ``` ```bloblang root.new_value = this.value.pow(-2) # In: {"value":2} # Out: {"new_value":0.25} ``` ### [](#round)round Rounds a number to the nearest integer. Values at .5 round away from zero. Returns an integer if the result fits in 64-bit, otherwise returns a float. #### [](#examples-50)Examples ```bloblang root.new_value = this.value.round() # In: {"value":5.3} # Out: {"new_value":5} # In: {"value":5.9} # Out: {"new_value":6} ``` ```bloblang root.rounded = this.score.round() # In: {"score":87.5} # Out: {"rounded":88} ``` ### [](#sin)sin Calculates the sine of a given angle specified in radians. #### [](#examples-51)Examples ```bloblang root.new_value = (this.value * (pi() / 180)).sin() # In: {"value":45} # Out: {"new_value":0.7071067811865475} # In: {"value":0} # Out: {"new_value":0} # In: {"value":90} # Out: {"new_value":1} ``` ### [](#tan)tan Calculates the tangent of a given angle specified in radians. #### [](#examples-52)Examples ```bloblang root.new_value = "%f".format((this.value * (pi() / 180)).tan()) # In: {"value":0} # Out: {"new_value":"0.000000"} # In: {"value":45} # Out: {"new_value":"1.000000"} # In: {"value":180} # Out: {"new_value":"-0.000000"} ``` ### [](#uint16)uint16 Converts a numerical type into a 16-bit unsigned integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 16-bit unsigned integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-53)Examples ```bloblang root.a = this.a.uint16() root.b = this.b.round().uint16() root.c = this.c.uint16() root.d = this.d.uint16().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":0} ``` ```bloblang root = this.uint16() # In: "0xDE" # Out: 222 ``` ### [](#uint32)uint32 Converts a numerical type into a 32-bit unsigned integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 32-bit unsigned integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-54)Examples ```bloblang root.a = this.a.uint32() root.b = this.b.round().uint32() root.c = this.c.uint32() root.d = this.d.uint32().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":0} ``` ```bloblang root = this.uint32() # In: "0xDEAD" # Out: 57005 ``` ### [](#uint64)uint64 Converts a numerical type into a 64-bit unsigned integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 64-bit unsigned integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-55)Examples ```bloblang root.a = this.a.uint64() root.b = this.b.round().uint64() root.c = this.c.uint64() root.d = this.d.uint64().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":0} ``` ```bloblang root = this.uint64() # In: "0xDEADBEEF" # Out: 3735928559 ``` ### [](#uint8)uint8 Converts a numerical type into a 8-bit unsigned integer, this is for advanced use cases where a specific data type is needed for a given component (such as the ClickHouse SQL driver). If the value is a string then an attempt will be made to parse it as a 8-bit unsigned integer. If the target value exceeds the capacity of an integer or contains decimal values then this method will throw an error. In order to convert a floating point number containing decimals first use [`.round()`](#round) on the value. Please refer to the [`strconv.ParseInt` documentation](https://pkg.go.dev/strconv#ParseInt) for details regarding the supported formats. #### [](#examples-56)Examples ```bloblang root.a = this.a.uint8() root.b = this.b.round().uint8() root.c = this.c.uint8() root.d = this.d.uint8().catch(0) # In: {"a":12,"b":12.34,"c":"12","d":-12} # Out: {"a":12,"b":12,"c":12,"d":0} ``` ```bloblang root = this.uint8() # In: "0xD" # Out: 13 ``` ## [](#object-array-manipulation)Object & array manipulation ### [](#all)all Tests whether all elements in an array satisfy a condition. Returns true only if the query evaluates to true for every element. Returns false for empty arrays. #### [](#parameters-44)Parameters | Name | Type | Description | | --- | --- | --- | | test | query expression | A test query to apply to each element. | #### [](#examples-57)Examples ```bloblang root.all_over_21 = this.patrons.all(patron -> patron.age >= 21) # In: {"patrons":[{"id":"1","age":18},{"id":"2","age":23}]} # Out: {"all_over_21":false} # In: {"patrons":[{"id":"1","age":45},{"id":"2","age":23}]} # Out: {"all_over_21":true} ``` ```bloblang root.all_positive = this.values.all(v -> v > 0) # In: {"values":[1,2,3,4,5]} # Out: {"all_positive":true} # In: {"values":[1,-2,3,4,5]} # Out: {"all_positive":false} ``` ### [](#any)any Tests whether at least one element in an array satisfies a condition. Returns true if the query evaluates to true for any element. Returns false for empty arrays. #### [](#parameters-45)Parameters | Name | Type | Description | | --- | --- | --- | | test | query expression | A test query to apply to each element. | #### [](#examples-58)Examples ```bloblang root.any_over_21 = this.patrons.any(patron -> patron.age >= 21) # In: {"patrons":[{"id":"1","age":18},{"id":"2","age":23}]} # Out: {"any_over_21":true} # In: {"patrons":[{"id":"1","age":10},{"id":"2","age":12}]} # Out: {"any_over_21":false} ``` ```bloblang root.has_errors = this.results.any(r -> r.status == "error") # In: {"results":[{"status":"ok"},{"status":"error"},{"status":"ok"}]} # Out: {"has_errors":true} # In: {"results":[{"status":"ok"},{"status":"ok"}]} # Out: {"has_errors":false} ``` ### [](#append)append Adds one or more elements to the end of an array and returns the new array. The original array is not modified. #### [](#examples-59)Examples ```bloblang root.foo = this.foo.append("and", "this") # In: {"foo":["bar","baz"]} # Out: {"foo":["bar","baz","and","this"]} ``` ```bloblang root.combined = this.items.append(this.new_item) # In: {"items":["apple","banana"],"new_item":"orange"} # Out: {"combined":["apple","banana","orange"]} ``` ### [](#assign)assign Merges two objects or arrays with override behavior. For objects, source values replace destination values on key conflicts. Arrays are concatenated. To preserve both values on conflict, use the merge method instead. #### [](#parameters-46)Parameters | Name | Type | Description | | --- | --- | --- | | with | unknown | A value to merge the target value with. | #### [](#examples-60)Examples ```bloblang root = this.foo.assign(this.bar) # In: {"foo":{"first_name":"fooer","likes":"bars"},"bar":{"second_name":"barer","likes":"foos"}} # Out: {"first_name":"fooer","likes":"foos","second_name":"barer"} ``` Override defaults with user settings: ```bloblang root.config = this.defaults.assign(this.user_settings) # In: {"defaults":{"timeout":30,"retries":3},"user_settings":{"timeout":60}} # Out: {"config":{"retries":3,"timeout":60}} ``` ### [](#collapse)collapse Flattens a nested structure into a flat object with dot-notation keys. #### [](#parameters-47)Parameters | Name | Type | Description | | --- | --- | --- | | include_empty | bool | Whether to include empty objects and arrays in the resulting object. | #### [](#examples-61)Examples ```bloblang root.result = this.collapse() # In: {"foo":[{"bar":"1"},{"bar":{}},{"bar":"2"},{"bar":[]}]} # Out: {"result":{"foo.0.bar":"1","foo.2.bar":"2"}} ``` Set include\_empty to true to preserve empty objects and arrays in the output: ```bloblang root.result = this.collapse(include_empty: true) # In: {"foo":[{"bar":"1"},{"bar":{}},{"bar":"2"},{"bar":[]}]} # Out: {"result":{"foo.0.bar":"1","foo.1.bar":{},"foo.2.bar":"2","foo.3.bar":[]}} ``` ### [](#concat)concat Concatenates an array value with one or more argument arrays. #### [](#examples-62)Examples ```bloblang root.foo = this.foo.concat(this.bar, this.baz) # In: {"foo":["a","b"],"bar":["c"],"baz":["d","e","f"]} # Out: {"foo":["a","b","c","d","e","f"]} ``` ### [](#contains)contains Tests if an array or object contains a value. #### [](#parameters-48)Parameters | Name | Type | Description | | --- | --- | --- | | value | unknown | A value to test against elements of the target. | #### [](#examples-63)Examples ```bloblang root.has_foo = this.thing.contains("foo") # In: {"thing":["this","foo","that"]} # Out: {"has_foo":true} # In: {"thing":["this","bar","that"]} # Out: {"has_foo":false} ``` ```bloblang root.has_bar = this.thing.contains(20) # In: {"thing":[10.3,20.0,"huh",3]} # Out: {"has_bar":true} # In: {"thing":[2,3,40,67]} # Out: {"has_bar":false} ``` ```bloblang root.has_foo = this.thing.contains("foo") # In: {"thing":"this foo that"} # Out: {"has_foo":true} # In: {"thing":"this bar that"} # Out: {"has_foo":false} ``` ### [](#diff)diff Compares the current value with another value and returns a detailed changelog describing all differences. The changelog contains operations (create, update, delete) with their paths and values, enabling you to track changes between data versions, implement audit logs, or synchronize data between systems. #### [](#parameters-49)Parameters | Name | Type | Description | | --- | --- | --- | | other | unknown | The value to compare against the current value. Can be any structured data (object or array). | #### [](#examples-64)Examples Compare two objects to track field changes: ```bloblang root.changes = this.before.diff(this.after) # In: {"before":{"name":"Alice","age":30},"after":{"name":"Alice","age":31,"city":"NYC"}} # Out: {"changes":[{"From":30,"Path":["age"],"To":31,"Type":"update"},{"From":null,"Path":["city"],"To":"NYC","Type":"create"}]} ``` Detect deletions in configuration changes: ```bloblang root.changelog = this.old_config.diff(this.new_config) # In: {"old_config":{"debug":true,"timeout":30},"new_config":{"timeout":60}} # Out: {"changelog":[{"From":true,"Path":["debug"],"To":null,"Type":"delete"},{"From":30,"Path":["timeout"],"To":60,"Type":"update"}]} ``` ### [](#enumerated)enumerated Transforms an array into an array of objects with index and value fields, making it easy to access both the position and content of each element. #### [](#examples-65)Examples ```bloblang root.foo = this.foo.enumerated() # In: {"foo":["bar","baz"]} # Out: {"foo":[{"index":0,"value":"bar"},{"index":1,"value":"baz"}]} ``` Useful for filtering by index position: ```bloblang root.first_two = this.items.enumerated().filter(item -> item.index < 2).map_each(item -> item.value) # In: {"items":["a","b","c","d"]} # Out: {"first_two":["a","b"]} ``` ### [](#exists)exists Checks whether a field exists at the specified dot path within an object. Returns true if the field is present (even if null), false otherwise. #### [](#parameters-50)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A dot path to a field. | #### [](#examples-66)Examples ```bloblang root.result = this.foo.exists("bar.baz") # In: {"foo":{"bar":{"baz":"yep, I exist"}}} # Out: {"result":true} # In: {"foo":{"bar":{}}} # Out: {"result":false} # In: {"foo":{}} # Out: {"result":false} ``` Also returns true for null values if the field exists: ```bloblang root.has_field = this.data.exists("optional_field") # In: {"data":{"optional_field":null}} # Out: {"has_field":true} # In: {"data":{}} # Out: {"has_field":false} ``` ### [](#explode)explode Expands a nested field into multiple documents. #### [](#parameters-51)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A dot path to a field to explode. | #### [](#examples-67)Examples ##### [](#on-arrays)On arrays When exploding an array, each element becomes a separate document with the array element replacing the original field: ```bloblang root = this.explode("value") # In: {"id":1,"value":["foo","bar","baz"]} # Out: [{"id":1,"value":"foo"},{"id":1,"value":"bar"},{"id":1,"value":"baz"}] ``` ##### [](#on-objects)On objects When exploding an object, the output keys match the nested object’s keys, with values being the full document where the target field is replaced by each nested value: ```bloblang root = this.explode("value") # In: {"id":1,"value":{"foo":2,"bar":[3,4],"baz":{"bev":5}}} # Out: {"bar":{"id":1,"value":[3,4]},"baz":{"id":1,"value":{"bev":5}},"foo":{"id":1,"value":2}} ``` ### [](#filter)filter Filters array or object elements based on a condition. #### [](#parameters-52)Parameters | Name | Type | Description | | --- | --- | --- | | test | query expression | A query to apply to each element, if this query resolves to any value other than a boolean true the element will be removed from the result. | #### [](#examples-68)Examples ```bloblang root.new_nums = this.nums.filter(num -> num > 10) # In: {"nums":[3,11,4,17]} # Out: {"new_nums":[11,17]} ``` ##### [](#on-objects-2)On objects When filtering objects, the query receives a context with `key` and `value` fields for each entry: ```bloblang root.new_dict = this.dict.filter(item -> item.value.contains("foo")) # In: {"dict":{"first":"hello foo","second":"world","third":"this foo is great"}} # Out: {"new_dict":{"first":"hello foo","third":"this foo is great"}} ``` ### [](#find)find Searches an array for a matching value and returns the index of the first occurrence. Returns -1 if no match is found. Numeric types are compared by value regardless of representation. #### [](#parameters-53)Parameters | Name | Type | Description | | --- | --- | --- | | value | unknown | A value to find. | #### [](#examples-69)Examples ```bloblang root.index = this.find("bar") # In: ["foo", "bar", "baz"] # Out: {"index":1} ``` ```bloblang root.index = this.things.find(this.goal) # In: {"goal":"bar","things":["foo", "bar", "baz"]} # Out: {"index":1} ``` ### [](#find_all)find_all Searches an array for all occurrences of a value and returns an array of matching indexes. Returns an empty array if no matches are found. Numeric types are compared by value regardless of representation. #### [](#parameters-54)Parameters | Name | Type | Description | | --- | --- | --- | | value | unknown | A value to find. | #### [](#examples-70)Examples ```bloblang root.index = this.find_all("bar") # In: ["foo", "bar", "baz", "bar"] # Out: {"index":[1,3]} ``` ```bloblang root.indexes = this.things.find_all(this.goal) # In: {"goal":"bar","things":["foo", "bar", "baz", "bar", "buz"]} # Out: {"indexes":[1,3]} ``` ### [](#find_all_by)find_all_by Searches an array for all elements that satisfy a condition and returns an array of their indexes. Returns an empty array if no elements match. #### [](#parameters-55)Parameters | Name | Type | Description | | --- | --- | --- | | query | query expression | A query to execute for each element. | #### [](#examples-71)Examples ```bloblang root.index = this.find_all_by(v -> v != "bar") # In: ["foo", "bar", "baz"] # Out: {"index":[0,2]} ``` Find all indexes matching criteria: ```bloblang root.error_indexes = this.logs.find_all_by(log -> log.level == "error") # In: {"logs":[{"level":"info"},{"level":"error"},{"level":"warn"},{"level":"error"}]} # Out: {"error_indexes":[1,3]} ``` ### [](#find_by)find_by Searches an array for the first element that satisfies a condition and returns its index. Returns -1 if no element matches the query. #### [](#parameters-56)Parameters | Name | Type | Description | | --- | --- | --- | | query | query expression | A query to execute for each element. | #### [](#examples-72)Examples ```bloblang root.index = this.find_by(v -> v != "bar") # In: ["foo", "bar", "baz"] # Out: {"index":0} ``` Find first object matching criteria: ```bloblang root.first_adult = this.users.find_by(u -> u.age >= 18) # In: {"users":[{"name":"Alice","age":15},{"name":"Bob","age":22},{"name":"Carol","age":19}]} # Out: {"first_adult":1} ``` ### [](#flatten)flatten Flattens an array by one level, expanding nested arrays into the parent array. Only the first level of nesting is removed. #### [](#examples-73)Examples ```bloblang root.result = this.flatten() # In: ["foo",["bar","baz"],"buz"] # Out: {"result":["foo","bar","baz","buz"]} ``` Deeper nesting requires multiple flatten calls: ```bloblang root.result = this.data.flatten() # In: {"data":["a",["b",["c","d"]],"e"]} # Out: {"result":["a","b",["c","d"],"e"]} ``` ### [](#fold)fold Reduces an array to a single value by iteratively applying a function. Also known as reduce or aggregate. The query receives an accumulator (tally) and current element (value) for each iteration. #### [](#parameters-57)Parameters | Name | Type | Description | | --- | --- | --- | | initial | unknown | The initial value to start the fold with. For example, an empty object {}, a zero count 0, or an empty string "". | | query | query expression | A query to apply for each element. The query is provided an object with two fields; tally containing the current tally, and value containing the value of the current element. The query should result in a new tally to be passed to the next element query. | #### [](#examples-74)Examples Sum numbers in an array: ```bloblang root.sum = this.foo.fold(0, item -> item.tally + item.value) # In: {"foo":[3,8,11]} # Out: {"sum":22} ``` Concatenate strings: ```bloblang root.result = this.foo.fold("", item -> "%v%v".format(item.tally, item.value)) # In: {"foo":["hello ", "world"]} # Out: {"result":"hello world"} ``` Merge an array of objects into a single object: ```bloblang root.smoothie = this.fruits.fold({}, item -> item.tally.merge(item.value)) # In: {"fruits":[{"apple":5},{"banana":3},{"orange":8}]} # Out: {"smoothie":{"apple":5,"banana":3,"orange":8}} ``` ### [](#get)get Extract a field value, identified via a [dot path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/), from an object. #### [](#parameters-58)Parameters | Name | Type | Description | | --- | --- | --- | | path | string | A dot path identifying a field to obtain. | #### [](#examples-75)Examples ```bloblang root.result = this.foo.get(this.target) # In: {"foo":{"bar":"from bar","baz":"from baz"},"target":"bar"} # Out: {"result":"from bar"} # In: {"foo":{"bar":"from bar","baz":"from baz"},"target":"baz"} # Out: {"result":"from baz"} ``` ### [](#index)index Extract an element from an array by an index. The index can be negative, and if so the element will be selected from the end counting backwards starting from -1. E.g. an index of -1 returns the last element, an index of -2 returns the element before the last, and so on. #### [](#parameters-59)Parameters | Name | Type | Description | | --- | --- | --- | | index | integer | The index to obtain from an array. | #### [](#examples-76)Examples ```bloblang root.last_name = this.names.index(-1) # In: {"names":["rachel","stevens"]} # Out: {"last_name":"stevens"} ``` It is also possible to use this method on byte arrays, in which case the selected element will be returned as an integer: ```bloblang root.last_byte = this.name.bytes().index(-1) # In: {"name":"foobar bazson"} # Out: {"last_byte":110} ``` ### [](#join)join Joins an array of strings with an optional delimiter. #### [](#parameters-60)Parameters | Name | Type | Description | | --- | --- | --- | | delimiter (optional) | string | An optional delimiter to add between each string. | #### [](#examples-77)Examples ```bloblang root.joined_words = this.words.join() root.joined_numbers = this.numbers.map_each(this.string()).join(",") # In: {"words":["hello","world"],"numbers":[3,8,11]} # Out: {"joined_numbers":"3,8,11","joined_words":"helloworld"} ``` ### [](#json_path)json_path Executes the given JSONPath expression on an object or array and returns the result. The JSONPath expression syntax can be found at [https://goessner.net/articles/JsonPath/](https://goessner.net/articles/JsonPath/). For more complex logic, you can use Gval expressions ([https://github.com/PaesslerAG/gval](https://github.com/PaesslerAG/gval)). #### [](#parameters-61)Parameters | Name | Type | Description | | --- | --- | --- | | expression | string | The JSONPath expression to execute. | #### [](#examples-78)Examples ```bloblang root.all_names = this.json_path("$..name") # In: {"name":"alice","foo":{"name":"bob"}} # Out: {"all_names":["alice","bob"]} # In: {"thing":["this","bar",{"name":"alice"}]} # Out: {"all_names":["alice"]} ``` ```bloblang root.text_objects = this.json_path("$.body[?(@.type=='text')]") # In: {"body":[{"type":"image","id":"foo"},{"type":"text","id":"bar"}]} # Out: {"text_objects":[{"id":"bar","type":"text"}]} ``` ### [](#json_schema)json_schema Checks a [JSON schema](https://json-schema.org/) against a value and returns the value if it matches or throws and error if it does not. #### [](#parameters-62)Parameters | Name | Type | Description | | --- | --- | --- | | schema | string | The schema to check values against. | #### [](#examples-79)Examples ```bloblang root = this.json_schema("""{ "type":"object", "properties":{ "foo":{ "type":"string" } } }""") # In: {"foo":"bar"} # Out: {"foo":"bar"} # In: {"foo":5} # Out: Error("failed assignment (line 1): field `this`: foo invalid type. expected: string, given: integer") ``` In order to load a schema from a file use the `file` function: ```bloblang root = this.json_schema(file(env("BENTHOS_TEST_BLOBLANG_SCHEMA_FILE"))) ``` ### [](#key_values)key_values Converts an object into an array of key-value pair objects. Each element has a 'key' field and a 'value' field. Order is not guaranteed unless sorted. #### [](#examples-80)Examples ```bloblang root.foo_key_values = this.foo.key_values().sort_by(pair -> pair.key) # In: {"foo":{"bar":1,"baz":2}} # Out: {"foo_key_values":[{"key":"bar","value":1},{"key":"baz","value":2}]} ``` Filter object entries by value: ```bloblang root.large_items = this.items.key_values().filter(pair -> pair.value > 15).map_each(pair -> pair.key) # In: {"items":{"a":5,"b":15,"c":20,"d":3}} # Out: {"large_items":["c"]} ``` ### [](#keys)keys Extracts all keys from an object and returns them as a sorted array. #### [](#examples-81)Examples ```bloblang root.foo_keys = this.foo.keys() # In: {"foo":{"bar":1,"baz":2}} # Out: {"foo_keys":["bar","baz"]} ``` Check if specific keys exist: ```bloblang root.has_id = this.data.keys().contains("id") # In: {"data":{"id":123,"name":"test"}} # Out: {"has_id":true} ``` ### [](#length)length Returns the length of an array, object, or string. #### [](#examples-82)Examples ```bloblang root.foo_len = this.foo.length() # In: {"foo":"hello world"} # Out: {"foo_len":11} ``` ```bloblang root.foo_len = this.foo.length() # In: {"foo":["first","second"]} # Out: {"foo_len":2} # In: {"foo":{"first":"bar","second":"baz"}} # Out: {"foo_len":2} ``` ### [](#map_each)map_each Applies a mapping to each element of an array or object. #### [](#parameters-63)Parameters | Name | Type | Description | | --- | --- | --- | | query | query expression | A query that will be used to map each element. | #### [](#examples-83)Examples ##### [](#on-arrays-2)On arrays Transforms each array element using a query. Return deleted() to remove an element, or the new value to replace it: ```bloblang root.new_nums = this.nums.map_each(num -> if num < 10 { deleted() } else { num - 10 }) # In: {"nums":[3,11,4,17]} # Out: {"new_nums":[1,7]} ``` ##### [](#on-objects-3)On objects Transforms each object value using a query. The query receives an object with 'key' and 'value' fields for each entry: ```bloblang root.new_dict = this.dict.map_each(item -> item.value.uppercase()) # In: {"dict":{"foo":"hello","bar":"world"}} # Out: {"new_dict":{"bar":"WORLD","foo":"HELLO"}} ``` ### [](#map_each_key)map_each_key Transforms object keys using a mapping query. #### [](#parameters-64)Parameters | Name | Type | Description | | --- | --- | --- | | query | query expression | A query that will be used to map each key. | #### [](#examples-84)Examples ```bloblang root.new_dict = this.dict.map_each_key(key -> key.uppercase()) # In: {"dict":{"keya":"hello","keyb":"world"}} # Out: {"new_dict":{"KEYA":"hello","KEYB":"world"}} ``` Conditionally transform keys: ```bloblang root = this.map_each_key(key -> if key.contains("kafka") { "_" + key }) # In: {"amqp_key":"foo","kafka_key":"bar","kafka_topic":"baz"} # Out: {"_kafka_key":"bar","_kafka_topic":"baz","amqp_key":"foo"} ``` ### [](#merge)merge Combines two objects or arrays. When merging objects, conflicting keys create arrays containing both values. Arrays are concatenated. For key override behavior instead, use the assign method. #### [](#parameters-65)Parameters | Name | Type | Description | | --- | --- | --- | | with | unknown | A value to merge the target value with. | #### [](#examples-85)Examples ```bloblang root = this.foo.merge(this.bar) # In: {"foo":{"first_name":"fooer","likes":"bars"},"bar":{"second_name":"barer","likes":"foos"}} # Out: {"first_name":"fooer","likes":["bars","foos"],"second_name":"barer"} ``` Merge arrays: ```bloblang root.combined = this.list1.merge(this.list2) # In: {"list1":["a","b"],"list2":["c","d"]} # Out: {"combined":["a","b","c","d"]} ``` ### [](#patch)patch Applies a changelog (created by the diff method) to the current value, transforming it according to the specified operations. This enables you to synchronize data, replay changes, or implement event sourcing patterns by applying recorded changes to reconstruct state. #### [](#parameters-66)Parameters | Name | Type | Description | | --- | --- | --- | | changelog | unknown | The changelog array to apply. Should be in the format returned by the diff method, containing Type, Path, From, and To fields for each change. | #### [](#examples-86)Examples Apply recorded changes to update an object: ```bloblang root.updated = this.current.patch(this.changelog) # In: {"current":{"name":"Alice","age":30},"changelog":[{"Type":"update","Path":["age"],"From":30,"To":31},{"Type":"create","Path":["city"],"From":null,"To":"NYC"}]} # Out: {"updated":{"age":31,"city":"NYC","name":"Alice"}} ``` Restore previous state by applying inverse changes: ```bloblang root.restored = this.modified.patch(this.reverse_changelog) # In: {"modified":{"timeout":60},"reverse_changelog":[{"Type":"create","Path":["debug"],"From":null,"To":true},{"Type":"update","Path":["timeout"],"From":60,"To":30}]} # Out: {"restored":{"debug":true,"timeout":30}} ``` ### [](#slice)slice Extracts a portion of an array or string. #### [](#parameters-67)Parameters | Name | Type | Description | | --- | --- | --- | | low | integer | The low bound, which is the first element of the selection, or if negative selects from the end. | | high (optional) | integer | An optional high bound. | #### [](#examples-87)Examples ```bloblang root.beginning = this.value.slice(0, 2) root.end = this.value.slice(4) # In: {"value":"foo bar"} # Out: {"beginning":"fo","end":"bar"} ``` A negative low index can be used, indicating an offset from the end of the sequence. If the low index is greater than the length of the sequence then an empty result is returned: ```bloblang root.last_chunk = this.value.slice(-4) root.the_rest = this.value.slice(0, -4) # In: {"value":"foo bar"} # Out: {"last_chunk":" bar","the_rest":"foo"} ``` ```bloblang root.beginning = this.value.slice(0, 2) root.end = this.value.slice(4) # In: {"value":["foo","bar","baz","buz","bev"]} # Out: {"beginning":["foo","bar"],"end":["bev"]} ``` A negative low index can be used, indicating an offset from the end of the sequence. If the low index is greater than the length of the sequence then an empty result is returned: ```bloblang root.last_chunk = this.value.slice(-2) root.the_rest = this.value.slice(0, -2) # In: {"value":["foo","bar","baz","buz","bev"]} # Out: {"last_chunk":["buz","bev"],"the_rest":["foo","bar","baz"]} ``` ### [](#sort)sort Sorts array elements in ascending order. #### [](#parameters-68)Parameters | Name | Type | Description | | --- | --- | --- | | compare (optional) | query expression | An optional query that should explicitly compare elements left and right and provide a boolean result. | #### [](#examples-88)Examples ```bloblang root.sorted = this.foo.sort() # In: {"foo":["bbb","ccc","aaa"]} # Out: {"sorted":["aaa","bbb","ccc"]} ``` Custom comparison for complex objects - return true if left < right: ```bloblang root.sorted = this.foo.sort(item -> item.left.v < item.right.v) # In: {"foo":[{"id":"foo","v":"bbb"},{"id":"bar","v":"ccc"},{"id":"baz","v":"aaa"}]} # Out: {"sorted":[{"id":"baz","v":"aaa"},{"id":"foo","v":"bbb"},{"id":"bar","v":"ccc"}]} ``` ### [](#sort_by)sort_by Sorts array elements by a specified field or expression. #### [](#parameters-69)Parameters | Name | Type | Description | | --- | --- | --- | | query | query expression | A query to apply to each element that yields a value used for sorting. | #### [](#examples-89)Examples ```bloblang root.sorted = this.foo.sort_by(ele -> ele.id) # In: {"foo":[{"id":"bbb","message":"bar"},{"id":"aaa","message":"foo"},{"id":"ccc","message":"baz"}]} # Out: {"sorted":[{"id":"aaa","message":"foo"},{"id":"bbb","message":"bar"},{"id":"ccc","message":"baz"}]} ``` Sort by numeric field: ```bloblang root.sorted = this.items.sort_by(item -> item.priority) # In: {"items":[{"name":"low","priority":3},{"name":"high","priority":1},{"name":"med","priority":2}]} # Out: {"sorted":[{"name":"high","priority":1},{"name":"med","priority":2},{"name":"low","priority":3}]} ``` ### [](#squash)squash Squashes an array of objects into a single object, where key collisions result in the values being merged (following similar rules as the `.merge()` method). #### [](#examples-90)Examples ```bloblang root.locations = this.locations.map_each(loc -> {loc.state: [loc.name]}).squash() # In: {"locations":[{"name":"Seattle","state":"WA"},{"name":"New York","state":"NY"},{"name":"Bellevue","state":"WA"},{"name":"Olympia","state":"WA"}]} # Out: {"locations":{"NY":["New York"],"WA":["Seattle","Bellevue","Olympia"]}} ``` ### [](#sum)sum Returns the sum of numeric values in an array. #### [](#examples-91)Examples ```bloblang root.sum = this.foo.sum() # In: {"foo":[3,8,4]} # Out: {"sum":15} ``` Works with decimals: ```bloblang root.total = this.prices.sum() # In: {"prices":[10.5,20.25,5.00]} # Out: {"total":35.75} ``` ### [](#unique)unique Returns an array with duplicate elements removed. #### [](#parameters-70)Parameters | Name | Type | Description | | --- | --- | --- | | emit (optional) | query expression | An optional query that can be used in order to yield a value for each element to determine uniqueness. | #### [](#examples-92)Examples ```bloblang root.uniques = this.foo.unique() # In: {"foo":["a","b","a","c"]} # Out: {"uniques":["a","b","c"]} ``` Use a query to determine uniqueness by a field: ```bloblang root.unique_users = this.users.unique(u -> u.id) # In: {"users":[{"id":1,"name":"Alice"},{"id":2,"name":"Bob"},{"id":1,"name":"Alice Duplicate"}]} # Out: {"unique_users":[{"id":1,"name":"Alice"},{"id":2,"name":"Bob"}]} ``` ### [](#values)values Returns an array of all values from an object. #### [](#examples-93)Examples ```bloblang root.foo_vals = this.foo.values().sort() # In: {"foo":{"bar":1,"baz":2}} # Out: {"foo_vals":[1,2]} ``` Find max value in object: ```bloblang root.max = this.scores.values().sort().index(-1) # In: {"scores":{"player1":85,"player2":92,"player3":78}} # Out: {"max":92} ``` ### [](#with)with Returns an object where all but one or more [field path](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/field_paths/) arguments are removed. Each path specifies a specific field to be retained from the input object, allowing for nested fields. If a key within a nested path does not exist then it is ignored. #### [](#examples-94)Examples ```bloblang root = this.with("inner.a","inner.c","d") # In: {"inner":{"a":"first","b":"second","c":"third"},"d":"fourth","e":"fifth"} # Out: {"d":"fourth","inner":{"a":"first","c":"third"}} ``` ### [](#without)without Returns an object with specified keys removed. #### [](#examples-95)Examples ```bloblang root = this.without("inner.a","inner.c","d") # In: {"inner":{"a":"first","b":"second","c":"third"},"d":"fourth","e":"fifth"} # Out: {"e":"fifth","inner":{"b":"second"}} ``` Remove sensitive fields: ```bloblang root = this.without("password","ssn","creditCard") # In: {"username":"alice","password":"secret","email":"alice@example.com","ssn":"123-45-6789"} # Out: {"email":"alice@example.com","username":"alice"} ``` ### [](#zip)zip Zip an array value with one or more argument arrays. Each array must match in length. #### [](#examples-96)Examples ```bloblang root.foo = this.foo.zip(this.bar, this.baz) # In: {"foo":["a","b","c"],"bar":[1,2,3],"baz":[4,5,6]} # Out: {"foo":[["a",1,4],["b",2,5],["c",3,6]]} ``` ## [](#parsing)Parsing ### [](#bloblang)bloblang Executes an argument Bloblang mapping on the target. This method can be used in order to execute dynamic mappings. Imports and functions that interact with the environment, such as `file` and `env`, or that access message information directly, such as `content` or `json`, are not enabled for dynamic Bloblang mappings. #### [](#parameters-71)Parameters | Name | Type | Description | | --- | --- | --- | | mapping | string | The mapping to execute. | #### [](#examples-97)Examples ```bloblang root.body = this.body.bloblang(this.mapping) # In: {"body":{"foo":"hello world"},"mapping":"root.foo = this.foo.uppercase()"} # Out: {"body":{"foo":"HELLO WORLD"}} # In: {"body":{"foo":"hello world 2"},"mapping":"root.foo = this.foo.capitalize()"} # Out: {"body":{"foo":"Hello World 2"}} ``` ### [](#format_json)format_json Formats a value as a JSON string. #### [](#parameters-72)Parameters | Name | Type | Description | | --- | --- | --- | | indent | string | Indentation string. Each element in a JSON object or array will begin on a new, indented line followed by one or more copies of indent according to the indentation nesting. | | no_indent | bool | Disable indentation. | | escape_html | bool | Escape problematic HTML characters. | #### [](#examples-98)Examples ```bloblang root = this.doc.format_json() # In: {"doc":{"foo":"bar"}} # Out: { "foo": "bar" } ``` Pass a string to the `indent` parameter in order to customise the indentation: ```bloblang root = this.format_json(" ") # In: {"doc":{"foo":"bar"}} # Out: { "doc": { "foo": "bar" } } ``` Use the `.string()` method in order to coerce the result into a string: ```bloblang root.doc = this.doc.format_json().string() # In: {"doc":{"foo":"bar"}} # Out: {"doc":"{\n \"foo\": \"bar\"\n}"} ``` Set the `no_indent` parameter to true to disable indentation. The result is equivalent to calling `bytes()`: ```bloblang root = this.doc.format_json(no_indent: true) # In: {"doc":{"foo":"bar"}} # Out: {"foo":"bar"} ``` Escapes problematic HTML characters: ```bloblang root = this.doc.format_json() # In: {"doc":{"email":"foo&bar@benthos.dev","name":"foo>bar"}} # Out: { "email": "foo\u0026bar@benthos.dev", "name": "foo\u003ebar" } ``` Set the `escape_html` parameter to false to disable escaping of problematic HTML characters: ```bloblang root = this.doc.format_json(escape_html: false) # In: {"doc":{"email":"foo&bar@benthos.dev","name":"foo>bar"}} # Out: { "email": "foo&bar@benthos.dev", "name": "foo>bar" } ``` ### [](#format_msgpack)format_msgpack Serializes structured data into MessagePack binary format. MessagePack is a compact binary serialization that is faster and more space-efficient than JSON, making it ideal for network transmission and storage of structured data. Returns a byte array that can be further encoded as needed. #### [](#examples-99)Examples Serialize object to MessagePack and encode as hex for transmission: ```bloblang root = this.format_msgpack().encode("hex") # In: {"foo":"bar"} # Out: 81a3666f6fa3626172 ``` Serialize data to MessagePack and base64 encode for embedding in JSON: ```bloblang root.msgpack_payload = this.data.format_msgpack().encode("base64") # In: {"data":{"foo":"bar"}} # Out: {"msgpack_payload":"gaNmb2+jYmFy"} ``` ### [](#format_xml)format_xml Serializes an object into an XML document. Converts structured data to XML format with support for attributes (prefixed with hyphen), custom indentation, and configurable root element. Returns XML as a byte array. #### [](#parameters-73)Parameters | Name | Type | Description | | --- | --- | --- | | indent | string | String to use for each level of indentation (default is 4 spaces). Each nested XML element will be indented by this string. | | no_indent | bool | Disable indentation and newlines to produce compact XML on a single line. | | root_tag (optional) | string | Custom name for the root XML element. By default, the root element name is derived from the first key in the object. | #### [](#examples-100)Examples Serialize object to pretty-printed XML with default indentation: ```bloblang root = this.format_xml() # In: {"foo":{"bar":{"baz":"foo bar baz"}}} # Out: foo bar baz ``` Create compact XML without indentation for smaller message size: ```bloblang root = this.format_xml(no_indent: true) # In: {"foo":{"bar":{"baz":"foo bar baz"}}} # Out: foo bar baz ``` ### [](#format_yaml)format_yaml Formats a value as a YAML string. #### [](#examples-101)Examples ```bloblang root = this.doc.format_yaml() # In: {"doc":{"foo":"bar"}} # Out: foo: bar ``` Use the `.string()` method in order to coerce the result into a string: ```bloblang root.doc = this.doc.format_yaml().string() # In: {"doc":{"foo":"bar"}} # Out: {"doc":"foo: bar\n"} ``` ### [](#infer_schema)infer_schema Attempt to infer the schema of a given value. The resulting schema can then be used as an input to schema conversion and enforcement methods. ### [](#parse_csv)parse_csv Parses CSV data into an array. #### [](#parameters-74)Parameters | Name | Type | Description | | --- | --- | --- | | parse_header_row | bool | Whether to reference the first row as a header row. If set to true the output structure for messages will be an object where field keys are determined by the header row. Otherwise, the output will be an array of row arrays. | | delimiter | string | The delimiter to use for splitting values in each record. It must be a single character. | | lazy_quotes | bool | If set to true, a quote may appear in an unquoted field and a non-doubled quote may appear in a quoted field. | #### [](#examples-102)Examples Parses CSV data with a header row: ```bloblang root.orders = this.orders.parse_csv() # In: {"orders":"foo,bar\nfoo 1,bar 1\nfoo 2,bar 2"} # Out: {"orders":[{"bar":"bar 1","foo":"foo 1"},{"bar":"bar 2","foo":"foo 2"}]} ``` Parses CSV data without a header row: ```bloblang root.orders = this.orders.parse_csv(false) # In: {"orders":"foo 1,bar 1\nfoo 2,bar 2"} # Out: {"orders":[["foo 1","bar 1"],["foo 2","bar 2"]]} ``` Parses CSV data delimited by dots: ```bloblang root.orders = this.orders.parse_csv(delimiter:".") # In: {"orders":"foo.bar\nfoo 1.bar 1\nfoo 2.bar 2"} # Out: {"orders":[{"bar":"bar 1","foo":"foo 1"},{"bar":"bar 2","foo":"foo 2"}]} ``` Parses CSV data containing a quote in an unquoted field: ```bloblang root.orders = this.orders.parse_csv(lazy_quotes:true) # In: {"orders":"foo,bar\nfoo 1,bar 1\nfoo\" \"2,bar\" \"2"} # Out: {"orders":[{"bar":"bar 1","foo":"foo 1"},{"bar":"bar\" \"2","foo":"foo\" \"2"}]} ``` ### [](#parse_form_url_encoded)parse_form_url_encoded Attempts to parse a url-encoded query string (from an x-www-form-urlencoded request body) and returns a structured result. #### [](#examples-103)Examples ```bloblang root.values = this.body.parse_form_url_encoded() # In: {"body":"noise=meow&animal=cat&fur=orange&fur=fluffy"} # Out: {"values":{"animal":"cat","fur":["orange","fluffy"],"noise":"meow"}} ``` ### [](#parse_json)parse_json Parses a JSON string into a structured value. #### [](#parameters-75)Parameters | Name | Type | Description | | --- | --- | --- | | use_number (optional) | bool | An optional flag that when set makes parsing numbers as json.Number instead of the default float64. | #### [](#examples-104)Examples ```bloblang root.doc = this.doc.parse_json() # In: {"doc":"{\"foo\":\"bar\"}"} # Out: {"doc":{"foo":"bar"}} ``` ```bloblang root.doc = this.doc.parse_json(use_number: true) # In: {"doc":"{\"foo\":\"11380878173205700000000000000000000000000000000\"}"} # Out: {"doc":{"foo":"11380878173205700000000000000000000000000000000"}} ``` ### [](#parse_logfmt)parse_logfmt Parses logfmt formatted data into an object. #### [](#examples-105)Examples ```bloblang root = this.msg.parse_logfmt() # In: {"msg":"level=info msg=\"hello world\" dur=1.5s"} # Out: {"dur":"1.5s","level":"info","msg":"hello world"} ``` ### [](#parse_msgpack)parse_msgpack Parses MessagePack binary data into a structured object. MessagePack is an efficient binary serialization format that is more compact than JSON while maintaining similar data structures. Commonly used for high-performance APIs and data interchange between microservices. #### [](#examples-106)Examples Parse MessagePack data from hex-encoded content: ```bloblang root = content().decode("hex").parse_msgpack() # In: 81a3666f6fa3626172 # Out: {"foo":"bar"} ``` Parse MessagePack from base64-encoded field: ```bloblang root.decoded = this.msgpack_data.decode("base64").parse_msgpack() # In: {"msgpack_data":"gaNmb2+jYmFy"} # Out: {"decoded":{"foo":"bar"}} ``` ### [](#parse_parquet)parse_parquet Parses Apache Parquet binary data into an array of objects. Parquet is a columnar storage format optimized for analytics, commonly used with big data systems like Apache Spark, Hive, and cloud data warehouses. Each row in the Parquet file becomes an object in the output array. #### [](#parameters-76)Parameters | Name | Type | Description | | --- | --- | --- | | byte_array_as_string | bool | Deprecated: This parameter is no longer used. | #### [](#examples-107)Examples Parse Parquet file data into structured objects: ```bloblang root.records = content().parse_parquet() ``` Process Parquet data from a field and extract specific columns: ```bloblang root.users = this.parquet_data.parse_parquet().map_each(row -> {"name": row.name, "email": row.email}) ``` ### [](#parse_url)parse_url Attempts to parse a URL from a string value, returning a structured result that describes the various facets of the URL. The fields returned within the structured result roughly follow [https://pkg.go.dev/net/url#URL](https://pkg.go.dev/net/url#URL), and may be expanded in future in order to present more information. #### [](#examples-108)Examples ```bloblang root.foo_url = this.foo_url.parse_url() # In: {"foo_url":"https://docs.redpanda.com/redpanda-connect/guides/bloblang/about/"} # Out: {"foo_url":{"fragment":"","host":"docs.redpanda.com","opaque":"","path":"/redpanda-connect/guides/bloblang/about/","raw_fragment":"","raw_path":"","raw_query":"","scheme":"https"}} ``` ```bloblang root.username = this.url.parse_url().user.name | "unknown" # In: {"url":"amqp://foo:bar@127.0.0.1:5672/"} # Out: {"username":"foo"} # In: {"url":"redis://localhost:6379"} # Out: {"username":"unknown"} ``` ### [](#parse_xml)parse_xml Parses an XML document into a structured object. Converts XML elements to JSON-like objects following these rules: - Element attributes are prefixed with a hyphen (e.g., `-id` for an `id` attribute) - Elements with both attributes and text content store the text in a `#text` field - Repeated elements become arrays - XML comments, directives, and processing instructions are ignored - Optionally cast numeric and boolean strings to their proper types. #### [](#parameters-77)Parameters | Name | Type | Description | | --- | --- | --- | | cast (optional) | bool | Whether to automatically cast numeric and boolean string values to their proper types. When false, all values remain as strings. | #### [](#examples-109)Examples Parse XML document into object structure: ```bloblang root.doc = this.doc.parse_xml() # In: {"doc":"This is a titleThis is some content"} # Out: {"doc":{"root":{"content":"This is some content","title":"This is a title"}}} ``` Parse XML with type casting enabled to convert strings to numbers and booleans: ```bloblang root.doc = this.doc.parse_xml(cast: true) # In: {"doc":"This is a title123True"} # Out: {"doc":{"root":{"bool":true,"number":{"#text":123,"-id":99},"title":"This is a title"}}} ``` ### [](#parse_yaml)parse_yaml Parses a YAML string into a structured value. #### [](#examples-110)Examples ```bloblang root.doc = this.doc.parse_yaml() # In: {"doc":"foo: bar"} # Out: {"doc":{"foo":"bar"}} ``` ## [](#regular-expressions)Regular expressions ### [](#re_find_all)re_find_all Finds all matches of a regular expression in a string. #### [](#parameters-78)Parameters | Name | Type | Description | | --- | --- | --- | | pattern | string | The pattern to match against. | #### [](#examples-111)Examples ```bloblang root.matches = this.value.re_find_all("a.") # In: {"value":"paranormal"} # Out: {"matches":["ar","an","al"]} ``` ```bloblang root.numbers = this.text.re_find_all("[0-9]+") # In: {"text":"I have 2 apples and 15 oranges"} # Out: {"numbers":["2","15"]} ``` ### [](#re_find_all_object)re_find_all_object Finds all regex matches as objects with named groups. #### [](#parameters-79)Parameters | Name | Type | Description | | --- | --- | --- | | pattern | string | The pattern to match against. | #### [](#examples-112)Examples ```bloblang root.matches = this.value.re_find_all_object("a(?Px*)b") # In: {"value":"-axxb-ab-"} # Out: {"matches":[{"0":"axxb","foo":"xx"},{"0":"ab","foo":""}]} ``` ```bloblang root.matches = this.value.re_find_all_object("(?m)(?P\\w+):\\s+(?P\\w+)$") # In: {"value":"option1: value1\noption2: value2\noption3: value3"} # Out: {"matches":[{"0":"option1: value1","key":"option1","value":"value1"},{"0":"option2: value2","key":"option2","value":"value2"},{"0":"option3: value3","key":"option3","value":"value3"}]} ``` ### [](#re_find_all_submatch)re_find_all_submatch Finds all regex matches with capture groups. #### [](#parameters-80)Parameters | Name | Type | Description | | --- | --- | --- | | pattern | string | The pattern to match against. | #### [](#examples-113)Examples ```bloblang root.matches = this.value.re_find_all_submatch("a(x*)b") # In: {"value":"-axxb-ab-"} # Out: {"matches":[["axxb","xx"],["ab",""]]} ``` ```bloblang root.emails = this.text.re_find_all_submatch("(\\w+)@(\\w+\\.\\w+)") # In: {"text":"Contact: alice@example.com or bob@test.org"} # Out: {"emails":[["alice@example.com","alice","example.com"],["bob@test.org","bob","test.org"]]} ``` ### [](#re_find_object)re_find_object Finds the first regex match as an object with named groups. #### [](#parameters-81)Parameters | Name | Type | Description | | --- | --- | --- | | pattern | string | The pattern to match against. | #### [](#examples-114)Examples ```bloblang root.matches = this.value.re_find_object("a(?Px*)b") # In: {"value":"-axxb-ab-"} # Out: {"matches":{"0":"axxb","foo":"xx"}} ``` ```bloblang root.matches = this.value.re_find_object("(?P\\w+):\\s+(?P\\w+)") # In: {"value":"option1: value1"} # Out: {"matches":{"0":"option1: value1","key":"option1","value":"value1"}} ``` ### [](#re_match)re_match Tests if a string matches a regular expression. #### [](#parameters-82)Parameters | Name | Type | Description | | --- | --- | --- | | pattern | string | The pattern to match against. | #### [](#examples-115)Examples ```bloblang root.matches = this.value.re_match("[0-9]") # In: {"value":"there are 10 puppies"} # Out: {"matches":true} # In: {"value":"there are ten puppies"} # Out: {"matches":false} ``` ### [](#re_replace)re_replace Replaces all regex matches with a replacement string that can reference capture groups using `$1`, `$2`, etc. Use for pattern-based transformations or data reformatting. #### [](#parameters-83)Parameters | Name | Type | Description | | --- | --- | --- | | pattern | string | The pattern to match against. | | value | string | The value to replace with. | ### [](#re_replace_all)re_replace_all Replaces all regex matches with a replacement string. #### [](#parameters-84)Parameters | Name | Type | Description | | --- | --- | --- | | pattern | string | The pattern to match against. | | value | string | The value to replace with. | #### [](#examples-116)Examples ```bloblang root.new_value = this.value.re_replace_all("ADD ([0-9]+)","+($1)") # In: {"value":"foo ADD 70"} # Out: {"new_value":"foo +(70)"} ``` ```bloblang root.masked = this.email.re_replace_all("(\\w{2})\\w+@", "$1***@") # In: {"email":"alice@example.com"} # Out: {"masked":"al***@example.com"} ``` ## [](#sql)SQL ### [](#vector)vector Converts an array of numbers into a vector type suitable for insertion into SQL databases with vector/embedding support. This is commonly used with PostgreSQL’s pgvector extension for storing and querying machine learning embeddings, enabling similarity search and vector operations in your database. #### [](#examples-117)Examples Convert embeddings array to vector for pgvector storage: ```bloblang root.embedding = this.embeddings.vector() root.text = this.text ``` Process ML model output into database-ready vector format: ```bloblang root.doc_id = this.id root.vector_embedding = this.model_output.map_each(num -> num.number()).vector() ``` ## [](#string-manipulation)String manipulation ### [](#capitalize)capitalize Converts a string to title case with Unicode letter mapping. #### [](#examples-118)Examples ```bloblang root.title = this.title.capitalize() # In: {"title":"the foo bar"} # Out: {"title":"The Foo Bar"} ``` ```bloblang root.name = this.name.capitalize() # In: {"name":"alice smith"} # Out: {"name":"Alice Smith"} ``` ### [](#compare_argon2)compare_argon2 Checks whether a string matches a hashed secret using Argon2. #### [](#parameters-85)Parameters | Name | Type | Description | | --- | --- | --- | | hashed_secret | string | The hashed secret to compare with the input. This must be a fully-qualified string which encodes the Argon2 options used to generate the hash. | #### [](#examples-119)Examples ```bloblang root.match = this.secret.compare_argon2("$argon2id$v=19$m=4096,t=3,p=1$c2FsdHktbWNzYWx0ZmFjZQ$RMUMwgtS32/mbszd+ke4o4Ej1jFpYiUqY6MHWa69X7Y") # In: {"secret":"there-are-many-blobs-in-the-sea"} # Out: {"match":true} ``` ```bloblang root.match = this.secret.compare_argon2("$argon2id$v=19$m=4096,t=3,p=1$c2FsdHktbWNzYWx0ZmFjZQ$RMUMwgtS32/mbszd+ke4o4Ej1jFpYiUqY6MHWa69X7Y") # In: {"secret":"will-i-ever-find-love"} # Out: {"match":false} ``` ### [](#compare_bcrypt)compare_bcrypt Checks whether a string matches a hashed secret using bcrypt. #### [](#parameters-86)Parameters | Name | Type | Description | | --- | --- | --- | | hashed_secret | string | The hashed secret value to compare with the input. | #### [](#examples-120)Examples ```bloblang root.match = this.secret.compare_bcrypt("$2y$10$Dtnt5NNzVtMCOZONT705tOcS8It6krJX8bEjnDJnwxiFKsz1C.3Ay") # In: {"secret":"there-are-many-blobs-in-the-sea"} # Out: {"match":true} ``` ```bloblang root.match = this.secret.compare_bcrypt("$2y$10$Dtnt5NNzVtMCOZONT705tOcS8It6krJX8bEjnDJnwxiFKsz1C.3Ay") # In: {"secret":"will-i-ever-find-love"} # Out: {"match":false} ``` ### [](#contains-2)contains Tests if an array or object contains a value. #### [](#parameters-87)Parameters | Name | Type | Description | | --- | --- | --- | | value | unknown | A value to test against elements of the target. | #### [](#examples-121)Examples ```bloblang root.has_foo = this.thing.contains("foo") # In: {"thing":["this","foo","that"]} # Out: {"has_foo":true} # In: {"thing":["this","bar","that"]} # Out: {"has_foo":false} ``` ```bloblang root.has_bar = this.thing.contains(20) # In: {"thing":[10.3,20.0,"huh",3]} # Out: {"has_bar":true} # In: {"thing":[2,3,40,67]} # Out: {"has_bar":false} ``` ```bloblang root.has_foo = this.thing.contains("foo") # In: {"thing":"this foo that"} # Out: {"has_foo":true} # In: {"thing":"this bar that"} # Out: {"has_foo":false} ``` ### [](#escape_html)escape_html Escapes HTML special characters. #### [](#examples-122)Examples ```bloblang root.escaped = this.value.escape_html() # In: {"value":"foo & bar"} # Out: {"escaped":"foo & bar"} ``` ```bloblang root.safe_html = this.user_input.escape_html() # In: {"user_input":""} # Out: {"safe_html":"<script>alert('xss')</script>"} ``` ### [](#escape_url_path)escape_url_path Escapes a string for use in URL paths. #### [](#examples-123)Examples ```bloblang root.escaped = this.value.escape_url_path() # In: {"value":"foo & bar"} # Out: {"escaped":"foo%20&%20bar"} ``` ```bloblang root.url = "https://example.com/docs/" + this.path.escape_url_path() # In: {"path":"my document.pdf"} # Out: {"url":"https://example.com/docs/my%20document.pdf"} ``` ### [](#escape_url_query)escape_url_query Escapes a string for use in URL query parameters. #### [](#examples-124)Examples ```bloblang root.escaped = this.value.escape_url_query() # In: {"value":"foo & bar"} # Out: {"escaped":"foo+%26+bar"} ``` ```bloblang root.url = "https://example.com?search=" + this.query.escape_url_query() # In: {"query":"hello world!"} # Out: {"url":"https://example.com?search=hello+world%21"} ``` ### [](#filepath_join)filepath_join Joins filepath components into a single path. #### [](#examples-125)Examples ```bloblang root.path = this.path_elements.filepath_join() # In: {"path_elements":["/foo/","bar.txt"]} # Out: {"path":"/foo/bar.txt"} ``` ### [](#filepath_split)filepath_split Splits a filepath into directory and filename components. #### [](#examples-126)Examples ```bloblang root.path_sep = this.path.filepath_split() # In: {"path":"/foo/bar.txt"} # Out: {"path_sep":["/foo/","bar.txt"]} # In: {"path":"baz.txt"} # Out: {"path_sep":["","baz.txt"]} ``` ### [](#format)format Formats a value using a specified format string. #### [](#examples-127)Examples ```bloblang root.foo = "%s(%v): %v".format(this.name, this.age, this.fingers) # In: {"name":"lance","age":37,"fingers":13} # Out: {"foo":"lance(37): 13"} ``` ```bloblang root.message = "User %s has %v points".format(this.username, this.score) # In: {"username":"alice","score":100} # Out: {"message":"User alice has 100 points"} ``` ### [](#has_prefix)has_prefix Tests if a string starts with a specified prefix. #### [](#parameters-88)Parameters | Name | Type | Description | | --- | --- | --- | | value | string | The string to test. | #### [](#examples-128)Examples ```bloblang root.t1 = this.v1.has_prefix("foo") root.t2 = this.v2.has_prefix("foo") # In: {"v1":"foobar","v2":"barfoo"} # Out: {"t1":true,"t2":false} ``` ### [](#has_suffix)has_suffix Tests if a string ends with a specified suffix. #### [](#parameters-89)Parameters | Name | Type | Description | | --- | --- | --- | | value | string | The string to test. | #### [](#examples-129)Examples ```bloblang root.t1 = this.v1.has_suffix("foo") root.t2 = this.v2.has_suffix("foo") # In: {"v1":"foobar","v2":"barfoo"} # Out: {"t1":false,"t2":true} ``` ### [](#index_of)index_of Returns the index of the first occurrence of a substring. #### [](#parameters-90)Parameters | Name | Type | Description | | --- | --- | --- | | value | string | A string to search for. | #### [](#examples-130)Examples ```bloblang root.index = this.thing.index_of("bar") # In: {"thing":"foobar"} # Out: {"index":3} ``` ```bloblang root.index = content().index_of("meow") # In: the cat meowed, the dog woofed # Out: {"index":8} ``` ### [](#length-2)length Returns the length of an array, object, or string. #### [](#examples-131)Examples ```bloblang root.foo_len = this.foo.length() # In: {"foo":"hello world"} # Out: {"foo_len":11} ``` ```bloblang root.foo_len = this.foo.length() # In: {"foo":["first","second"]} # Out: {"foo_len":2} # In: {"foo":{"first":"bar","second":"baz"}} # Out: {"foo_len":2} ``` ### [](#lowercase)lowercase Converts all letters in a string to lowercase. #### [](#examples-132)Examples ```bloblang root.foo = this.foo.lowercase() # In: {"foo":"HELLO WORLD"} # Out: {"foo":"hello world"} ``` ```bloblang root.email = this.user_email.lowercase() # In: {"user_email":"User@Example.COM"} # Out: {"email":"user@example.com"} ``` ### [](#quote)quote Wraps a string in double quotes and escapes special characters. #### [](#examples-133)Examples ```bloblang root.quoted = this.thing.quote() # In: {"thing":"foo\nbar"} # Out: {"quoted":"\"foo\\nbar\""} ``` ```bloblang root.literal = this.text.quote() # In: {"text":"hello\tworld"} # Out: {"literal":"\"hello\\tworld\""} ``` ### [](#repeat)repeat Creates a string by repeating the input a specified number of times. #### [](#parameters-91)Parameters | Name | Type | Description | | --- | --- | --- | | count | integer | The number of times to repeat the string. | #### [](#examples-134)Examples ```bloblang root.repeated = this.name.repeat(3) root.not_repeated = this.name.repeat(0) # In: {"name":"bob"} # Out: {"not_repeated":"","repeated":"bobbobbob"} ``` ```bloblang root.separator = "-".repeat(10) # In: {} # Out: {"separator":"----------"} ``` ### [](#replace)replace Replaces all occurrences of a substring with another string. Use for text transformation, cleaning data, or normalizing strings. #### [](#parameters-92)Parameters | Name | Type | Description | | --- | --- | --- | | old | string | A string to match against. | | new | string | A string to replace with. | ### [](#replace_all)replace_all Replaces all occurrences of a substring with another. #### [](#parameters-93)Parameters | Name | Type | Description | | --- | --- | --- | | old | string | A string to match against. | | new | string | A string to replace with. | #### [](#examples-135)Examples ```bloblang root.new_value = this.value.replace_all("foo","dog") # In: {"value":"The foo ate my homework"} # Out: {"new_value":"The dog ate my homework"} ``` ```bloblang root.clean = this.text.replace_all(" ", " ") # In: {"text":"hello world foo"} # Out: {"clean":"hello world foo"} ``` ### [](#replace_all_many)replace_all_many Performs multiple find-and-replace operations in sequence. #### [](#parameters-94)Parameters | Name | Type | Description | | --- | --- | --- | | values | array | An array of values, each even value will be replaced with the following odd value. | #### [](#examples-136)Examples ```bloblang root.new_value = this.value.replace_all_many([ "", "<b>", "", "</b>", "", "<i>", "", "</i>", ]) # In: {"value":"Hello World"} # Out: {"new_value":"<i>Hello</i> <b>World</b>"} ``` ### [](#replace_many)replace_many Performs multiple find-and-replace operations in sequence using an array of `[old, new]` pairs. More efficient than chaining multiple `replace_all` calls. Use for bulk text transformations. #### [](#parameters-95)Parameters | Name | Type | Description | | --- | --- | --- | | values | array | An array of values, each even value will be replaced with the following odd value. | ### [](#reverse)reverse Reverses the order of characters in a string. #### [](#examples-137)Examples ```bloblang root.reversed = this.thing.reverse() # In: {"thing":"backwards"} # Out: {"reversed":"sdrawkcab"} ``` ```bloblang root = content().reverse() # In: {"thing":"backwards"} # Out: }"sdrawkcab":"gniht"{ ``` ### [](#slice-2)slice Extracts a portion of an array or string. #### [](#parameters-96)Parameters | Name | Type | Description | | --- | --- | --- | | low | integer | The low bound, which is the first element of the selection, or if negative selects from the end. | | high (optional) | integer | An optional high bound. | #### [](#examples-138)Examples ```bloblang root.beginning = this.value.slice(0, 2) root.end = this.value.slice(4) # In: {"value":"foo bar"} # Out: {"beginning":"fo","end":"bar"} ``` A negative low index can be used, indicating an offset from the end of the sequence. If the low index is greater than the length of the sequence then an empty result is returned: ```bloblang root.last_chunk = this.value.slice(-4) root.the_rest = this.value.slice(0, -4) # In: {"value":"foo bar"} # Out: {"last_chunk":" bar","the_rest":"foo"} ``` ```bloblang root.beginning = this.value.slice(0, 2) root.end = this.value.slice(4) # In: {"value":["foo","bar","baz","buz","bev"]} # Out: {"beginning":["foo","bar"],"end":["bev"]} ``` A negative low index can be used, indicating an offset from the end of the sequence. If the low index is greater than the length of the sequence then an empty result is returned: ```bloblang root.last_chunk = this.value.slice(-2) root.the_rest = this.value.slice(0, -2) # In: {"value":["foo","bar","baz","buz","bev"]} # Out: {"last_chunk":["buz","bev"],"the_rest":["foo","bar","baz"]} ``` ### [](#slug)slug Converts a string into a URL-friendly slug by replacing spaces with hyphens, removing special characters, and converting to lowercase. Supports multiple languages for proper transliteration of non-ASCII characters. #### [](#parameters-97)Parameters | Name | Type | Description | | --- | --- | --- | | lang (optional) | string | | #### [](#examples-139)Examples Create a URL-friendly slug from a string with special characters: ```bloblang root.slug = this.title.slug() # In: {"title":"Hello World! Welcome to Redpanda Connect"} # Out: {"slug":"hello-world-welcome-to-redpanda-connect"} ``` Create a slug preserving French language rules: ```bloblang root.slug = this.title.slug("fr") # In: {"title":"Café & Restaurant"} # Out: {"slug":"cafe-et-restaurant"} ``` ### [](#split)split Splits a string into an array of substrings. #### [](#parameters-98)Parameters | Name | Type | Description | | --- | --- | --- | | delimiter | string | The delimiter to split with. | | empty_as_null | bool | To treat empty substrings as null values | #### [](#examples-140)Examples ```bloblang root.new_value = this.value.split(",") # In: {"value":"foo,bar,baz"} # Out: {"new_value":["foo","bar","baz"]} ``` ```bloblang root.new_value = this.value.split(",", true) # In: {"value":"foo,,qux"} # Out: {"new_value":["foo",null,"qux"]} ``` ```bloblang root.words = this.sentence.split(" ") # In: {"sentence":"hello world from bloblang"} # Out: {"words":["hello","world","from","bloblang"]} ``` ### [](#strip_html)strip_html Removes HTML tags from a string, returning only the text content. Useful for extracting plain text from HTML documents, sanitizing user input, or preparing content for text analysis. Optionally preserves specific HTML elements while stripping all others. #### [](#parameters-99)Parameters | Name | Type | Description | | --- | --- | --- | | preserve (optional) | unknown | Optional array of HTML element names to preserve (e.g., ["strong", "em", "a"]). All other HTML tags will be removed. | #### [](#examples-141)Examples Extract plain text from HTML content: ```bloblang root.plain_text = this.html_content.strip_html() # In: {"html_content":"

Welcome to Redpanda Connect!

"} # Out: {"plain_text":"Welcome to Redpanda Connect!"} ``` Preserve specific HTML elements while removing others: ```bloblang root.sanitized = this.html.strip_html(["strong", "em"]) # In: {"html":"

Some bold and italic text with a

"} # Out: {"sanitized":"Some bold and italic text with a "} ``` ### [](#trim)trim Removes leading and trailing characters from a string. #### [](#parameters-100)Parameters | Name | Type | Description | | --- | --- | --- | | cutset (optional) | string | An optional string of characters to trim from the target value. | #### [](#examples-142)Examples ```bloblang root.title = this.title.trim("!?") root.description = this.description.trim() # In: {"description":" something happened and its amazing! ","title":"!!!watch out!?"} # Out: {"description":"something happened and its amazing!","title":"watch out"} ``` ### [](#trim_prefix)trim_prefix Removes a specified prefix from the beginning of a string. #### [](#parameters-101)Parameters | Name | Type | Description | | --- | --- | --- | | prefix | string | The leading prefix substring to trim from the string. | #### [](#examples-143)Examples ```bloblang root.name = this.name.trim_prefix("foobar_") root.description = this.description.trim_prefix("foobar_") # In: {"description":"unchanged","name":"foobar_blobton"} # Out: {"description":"unchanged","name":"blobton"} ``` ### [](#trim_suffix)trim_suffix Removes a specified suffix from the end of a string. #### [](#parameters-102)Parameters | Name | Type | Description | | --- | --- | --- | | suffix | string | The trailing suffix substring to trim from the string. | #### [](#examples-144)Examples ```bloblang root.name = this.name.trim_suffix("_foobar") root.description = this.description.trim_suffix("_foobar") # In: {"description":"unchanged","name":"blobton_foobar"} # Out: {"description":"unchanged","name":"blobton"} ``` ### [](#unescape_html)unescape_html Converts HTML entities back to their original characters. #### [](#examples-145)Examples ```bloblang root.unescaped = this.value.unescape_html() # In: {"value":"foo & bar"} # Out: {"unescaped":"foo & bar"} ``` ```bloblang root.text = this.html.unescape_html() # In: {"html":"<p>Hello & goodbye</p>"} # Out: {"text":"

Hello & goodbye

"} ``` ### [](#unescape_url_path)unescape_url_path Unescapes URL path encoding. #### [](#examples-146)Examples ```bloblang root.unescaped = this.value.unescape_url_path() # In: {"value":"foo%20&%20bar"} # Out: {"unescaped":"foo & bar"} ``` ```bloblang root.filename = this.path.unescape_url_path() # In: {"path":"my%20document.pdf"} # Out: {"filename":"my document.pdf"} ``` ### [](#unescape_url_query)unescape_url_query Unescapes URL query parameter encoding. #### [](#examples-147)Examples ```bloblang root.unescaped = this.value.unescape_url_query() # In: {"value":"foo+%26+bar"} # Out: {"unescaped":"foo & bar"} ``` ```bloblang root.search = this.param.unescape_url_query() # In: {"param":"hello+world%21"} # Out: {"search":"hello world!"} ``` ### [](#unicode_segments)unicode_segments Splits text into segments based on Unicode text segmentation rules. Returns an array of strings representing individual graphemes (visual characters), words (including punctuation and whitespace), or sentences. Handles complex Unicode correctly, including emoji with skin tone modifiers and zero-width joiners. #### [](#parameters-103)Parameters | Name | Type | Description | | --- | --- | --- | | segmentation_type | string | Type of segmentation: "grapheme", "word", or "sentence" | #### [](#examples-148)Examples Split text into sentences (preserves trailing spaces): ```bloblang root.sentences = this.text.unicode_segments("sentence") # In: {"text":"Hello world. How are you?"} # Out: {"sentences":["Hello world. ","How are you?"]} ``` Split text into grapheme clusters (handles complex emoji correctly): ```bloblang root.graphemes = this.emoji.unicode_segments("grapheme") # In: {"emoji":"👨‍👩‍👧‍👦❤️"} # Out: {"graphemes":["👨‍👩‍👧‍👦","❤️"]} ``` ### [](#unquote)unquote Removes surrounding quotes and interprets escape sequences. #### [](#examples-149)Examples ```bloblang root.unquoted = this.thing.unquote() # In: {"thing":"\"foo\\nbar\""} # Out: {"unquoted":"foo\nbar"} ``` ```bloblang root.text = this.literal.unquote() # In: {"literal":"\"hello\\tworld\""} # Out: {"text":"hello\tworld"} ``` ### [](#uppercase)uppercase Converts all letters in a string to uppercase. #### [](#examples-150)Examples ```bloblang root.foo = this.foo.uppercase() # In: {"foo":"hello world"} # Out: {"foo":"HELLO WORLD"} ``` ```bloblang root.code = this.product_code.uppercase() # In: {"product_code":"abc-123"} # Out: {"code":"ABC-123"} ``` ## [](#timestamp-manipulation)Timestamp manipulation ### [](#parse_duration)parse_duration Parses a Go-style duration string into nanoseconds. A duration string is a signed sequence of decimal numbers with unit suffixes like "300ms", "-1.5h", or "2h45m". Valid units: "ns", "us" (or "µs"), "ms", "s", "m", "h". #### [](#examples-151)Examples Parse microseconds to nanoseconds: ```bloblang root.delay_for_ns = this.delay_for.parse_duration() # In: {"delay_for":"50us"} # Out: {"delay_for_ns":50000} ``` Parse hours to seconds: ```bloblang root.delay_for_s = this.delay_for.parse_duration() / 1000000000 # In: {"delay_for":"2h"} # Out: {"delay_for_s":7200} ``` ### [](#parse_duration_iso8601)parse_duration_iso8601 Parses an ISO 8601 duration string into nanoseconds. Format: "P\[n\]Y\[n\]M\[n\]DT\[n\]H\[n\]M\[n\]S" or "P\[n\]W". Example: "P3Y6M4DT12H30M5S" means 3 years, 6 months, 4 days, 12 hours, 30 minutes, 5 seconds. Supports fractional seconds with full precision (not just one decimal place). #### [](#examples-152)Examples Parse complex ISO 8601 duration to nanoseconds: ```bloblang root.delay_for_ns = this.delay_for.parse_duration_iso8601() # In: {"delay_for":"P3Y6M4DT12H30M5S"} # Out: {"delay_for_ns":110839937000000000} ``` Parse hours to seconds: ```bloblang root.delay_for_s = this.delay_for.parse_duration_iso8601() / 1000000000 # In: {"delay_for":"PT2H"} # Out: {"delay_for_s":7200} ``` ### [](#ts_add_iso8601)ts_add_iso8601 Adds an ISO 8601 duration to a timestamp with calendar-aware precision for years, months, and days. Useful when you need to add durations that account for variable month lengths or leap years. #### [](#parameters-104)Parameters | Name | Type | Description | | --- | --- | --- | | duration | string | Duration in ISO 8601 format (e.g., "P1Y2M3D" for 1 year, 2 months, 3 days) | #### [](#examples-153)Examples Add one year to a timestamp: ```bloblang root.next_year = this.created_at.ts_add_iso8601("P1Y") # In: {"created_at":"2020-08-14T05:54:23Z"} # Out: {"next_year":"2021-08-14T05:54:23Z"} ``` Add a complex duration with multiple units: ```bloblang root.future_date = this.created_at.ts_add_iso8601("P1Y2M3DT4H5M6S") # In: {"created_at":"2020-01-01T00:00:00Z"} # Out: {"future_date":"2021-03-04T04:05:06Z"} ``` ### [](#ts_format)ts_format Formats a timestamp as a string using Go’s reference time format. Defaults to RFC 3339 if no format specified. The format uses "Mon Jan 2 15:04:05 -0700 MST 2006" as a reference. Accepts unix timestamps (with decimal precision) or RFC 3339 strings. Use ts\_strftime for strftime-style formats. #### [](#parameters-105)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The output format using Go’s reference time. | | tz (optional) | string | Optional timezone (e.g., 'UTC', 'America/New_York'). Defaults to input timezone or local time for unix timestamps. | #### [](#examples-154)Examples Format timestamp with custom format: ```bloblang root.something_at = this.created_at.ts_format("2006-Jan-02 15:04:05") # In: {"created_at":"2020-08-14T11:50:26.371Z"} # Out: {"something_at":"2020-Aug-14 11:50:26"} ``` Format unix timestamp with timezone specification: ```bloblang root.something_at = this.created_at.ts_format(format: "2006-Jan-02 15:04:05", tz: "UTC") # In: {"created_at":1597405526} # Out: {"something_at":"2020-Aug-14 11:45:26"} ``` ### [](#ts_parse)ts_parse Parses a timestamp string using Go’s reference time format and outputs a timestamp object. The format uses "Mon Jan 2 15:04:05 -0700 MST 2006" as a reference - show how this reference time would appear in your format. Use ts\_strptime for strftime-style formats instead. #### [](#parameters-106)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The format of the input string using Go’s reference time. | #### [](#examples-155)Examples Parse a date with abbreviated month name: ```bloblang root.doc.timestamp = this.doc.timestamp.ts_parse("2006-Jan-02") # In: {"doc":{"timestamp":"2020-Aug-14"}} # Out: {"doc":{"timestamp":"2020-08-14T00:00:00Z"}} ``` Parse a custom datetime format: ```bloblang root.parsed = this.timestamp.ts_parse("Jan 2, 2006 at 3:04pm (MST)") # In: {"timestamp":"Aug 14, 2020 at 5:54am (UTC)"} # Out: {"parsed":"2020-08-14T05:54:00Z"} ``` ### [](#ts_round)ts_round Rounds a timestamp to the nearest multiple of the specified duration. Halfway values round up. Accepts unix timestamps (seconds with optional decimal precision) or RFC 3339 formatted strings. #### [](#parameters-107)Parameters | Name | Type | Description | | --- | --- | --- | | duration | integer | A duration measured in nanoseconds to round by. | #### [](#examples-156)Examples Round timestamp to the nearest hour: ```bloblang root.created_at_hour = this.created_at.ts_round("1h".parse_duration()) # In: {"created_at":"2020-08-14T05:54:23Z"} # Out: {"created_at_hour":"2020-08-14T06:00:00Z"} ``` Round timestamp to the nearest minute: ```bloblang root.created_at_minute = this.created_at.ts_round("1m".parse_duration()) # In: {"created_at":"2020-08-14T05:54:23Z"} # Out: {"created_at_minute":"2020-08-14T05:54:00Z"} ``` ### [](#ts_strftime)ts_strftime Formats a timestamp as a string using strptime format specifiers (like %Y, %m, %d). Accepts unix timestamps (with decimal precision) or RFC 3339 strings. Supports %f for microseconds. Use ts\_format for Go-style reference time formats. #### [](#parameters-108)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The output format using strptime specifiers. | | tz (optional) | string | Optional timezone. Defaults to input timezone or local time for unix timestamps. | #### [](#examples-157)Examples Format timestamp with strftime specifiers: ```bloblang root.something_at = this.created_at.ts_strftime("%Y-%b-%d %H:%M:%S") # In: {"created_at":"2020-08-14T11:50:26.371Z"} # Out: {"something_at":"2020-Aug-14 11:50:26"} ``` Format with microseconds using %f directive: ```bloblang root.something_at = this.created_at.ts_strftime("%Y-%b-%d %H:%M:%S.%f", "UTC") # In: {"created_at":"2020-08-14T11:50:26.371Z"} # Out: {"something_at":"2020-Aug-14 11:50:26.371000"} ``` ### [](#ts_strptime)ts_strptime Parses a timestamp string using strptime format specifiers (like %Y, %m, %d) and outputs a timestamp object. Use ts\_parse for Go-style reference time formats instead. #### [](#parameters-109)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The format string using strptime specifiers (e.g., %Y-%m-%d). | #### [](#examples-158)Examples Parse date with abbreviated month using strptime format: ```bloblang root.doc.timestamp = this.doc.timestamp.ts_strptime("%Y-%b-%d") # In: {"doc":{"timestamp":"2020-Aug-14"}} # Out: {"doc":{"timestamp":"2020-08-14T00:00:00Z"}} ``` Parse datetime with microseconds using %f directive: ```bloblang root.doc.timestamp = this.doc.timestamp.ts_strptime("%Y-%b-%d %H:%M:%S.%f") # In: {"doc":{"timestamp":"2020-Aug-14 11:50:26.371000"}} # Out: {"doc":{"timestamp":"2020-08-14T11:50:26.371Z"}} ``` ### [](#ts_sub)ts_sub Calculates the duration in nanoseconds between two timestamps (t1 - t2). Returns a signed integer: positive if t1 is after t2, negative if t1 is before t2. Use .abs() for absolute duration. #### [](#parameters-110)Parameters | Name | Type | Description | | --- | --- | --- | | t2 | timestamp | The timestamp to subtract from the target timestamp. | #### [](#examples-159)Examples Calculate absolute duration between two timestamps: ```bloblang root.between = this.started_at.ts_sub("2020-08-14T05:54:23Z").abs() # In: {"started_at":"2020-08-13T05:54:23Z"} # Out: {"between":86400000000000} ``` Calculate signed duration (can be negative): ```bloblang root.duration_ns = this.end_time.ts_sub(this.start_time) # In: {"start_time":"2020-08-14T10:00:00Z","end_time":"2020-08-14T11:30:00Z"} # Out: {"duration_ns":5400000000000} ``` ### [](#ts_sub_iso8601)ts_sub_iso8601 Subtracts an ISO 8601 duration from a timestamp with calendar-aware precision for years, months, and days. Useful when you need to subtract durations that account for variable month lengths or leap years. #### [](#parameters-111)Parameters | Name | Type | Description | | --- | --- | --- | | duration | string | Duration in ISO 8601 format (e.g., "P1Y2M3D" for 1 year, 2 months, 3 days) | #### [](#examples-160)Examples Subtract one year from a timestamp: ```bloblang root.last_year = this.created_at.ts_sub_iso8601("P1Y") # In: {"created_at":"2020-08-14T05:54:23Z"} # Out: {"last_year":"2019-08-14T05:54:23Z"} ``` Subtract a complex duration with multiple units: ```bloblang root.past_date = this.created_at.ts_sub_iso8601("P1Y2M3DT4H5M6S") # In: {"created_at":"2021-03-04T04:05:06Z"} # Out: {"past_date":"2020-01-01T00:00:00Z"} ``` ### [](#ts_tz)ts_tz Converts a timestamp to a different timezone while preserving the moment in time. Accepts unix timestamps (seconds with optional decimal precision) or RFC 3339 formatted strings. #### [](#parameters-112)Parameters | Name | Type | Description | | --- | --- | --- | | tz | string | The timezone to change to. Use "UTC" for UTC, "Local" for local timezone, or an IANA Time Zone database location name like "America/New_York". | #### [](#examples-161)Examples Convert timestamp to UTC timezone: ```bloblang root.created_at_utc = this.created_at.ts_tz("UTC") # In: {"created_at":"2021-02-03T17:05:06+01:00"} # Out: {"created_at_utc":"2021-02-03T16:05:06Z"} ``` Convert timestamp to a specific timezone: ```bloblang root.created_at_ny = this.created_at.ts_tz("America/New_York") # In: {"created_at":"2021-02-03T16:05:06Z"} # Out: {"created_at_ny":"2021-02-03T11:05:06-05:00"} ``` ### [](#ts_unix)ts_unix Converts a timestamp to a unix timestamp (seconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing seconds. #### [](#examples-162)Examples Convert RFC 3339 timestamp to unix seconds: ```bloblang root.created_at_unix = this.created_at.ts_unix() # In: {"created_at":"2009-11-10T23:00:00Z"} # Out: {"created_at_unix":1257894000} ``` Unix timestamp passthrough returns same value: ```bloblang root.timestamp = this.ts.ts_unix() # In: {"ts":1257894000} # Out: {"timestamp":1257894000} ``` ### [](#ts_unix_micro)ts_unix_micro Converts a timestamp to a unix timestamp with microsecond precision (microseconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing microseconds. #### [](#examples-163)Examples Convert timestamp to microseconds since epoch: ```bloblang root.created_at_unix = this.created_at.ts_unix_micro() # In: {"created_at":"2009-11-10T23:00:00Z"} # Out: {"created_at_unix":1257894000000000} ``` Preserve microsecond precision from timestamp: ```bloblang root.precise_time = this.timestamp.ts_unix_micro() # In: {"timestamp":"2020-08-14T11:45:26.123456Z"} # Out: {"precise_time":1597405526123456} ``` ### [](#ts_unix_milli)ts_unix_milli Converts a timestamp to a unix timestamp with millisecond precision (milliseconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing milliseconds. #### [](#examples-164)Examples Convert timestamp to milliseconds since epoch: ```bloblang root.created_at_unix = this.created_at.ts_unix_milli() # In: {"created_at":"2009-11-10T23:00:00Z"} # Out: {"created_at_unix":1257894000000} ``` Useful for JavaScript timestamp compatibility: ```bloblang root.js_timestamp = this.event_time.ts_unix_milli() # In: {"event_time":"2020-08-14T11:45:26.123Z"} # Out: {"js_timestamp":1597405526123} ``` ### [](#ts_unix_nano)ts_unix_nano Converts a timestamp to a unix timestamp with nanosecond precision (nanoseconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing nanoseconds. #### [](#examples-165)Examples Convert timestamp to nanoseconds since epoch: ```bloblang root.created_at_unix = this.created_at.ts_unix_nano() # In: {"created_at":"2009-11-10T23:00:00Z"} # Out: {"created_at_unix":1257894000000000000} ``` Preserve full nanosecond precision: ```bloblang root.precise_time = this.timestamp.ts_unix_nano() # In: {"timestamp":"2020-08-14T11:45:26.123456789Z"} # Out: {"precise_time":1597405526123456789} ``` ## [](#type-coercion)Type coercion ### [](#array)array Converts a value to an array. #### [](#examples-166)Examples ```bloblang root.my_array = this.name.array() # In: {"name":"foobar bazson"} # Out: {"my_array":["foobar bazson"]} ``` ### [](#bool)bool Converts a value to a boolean with optional fallback. #### [](#parameters-113)Parameters | Name | Type | Description | | --- | --- | --- | | default (optional) | bool | An optional value to yield if the target cannot be parsed as a boolean. | #### [](#examples-167)Examples ```bloblang root.foo = this.thing.bool() root.bar = this.thing.bool(true) ``` ### [](#bytes)bytes Marshals a value into a byte array. #### [](#examples-168)Examples ```bloblang root.first_byte = this.name.bytes().index(0) # In: {"name":"foobar bazson"} # Out: {"first_byte":102} ``` ### [](#not_empty)not_empty Ensures a value is not empty. #### [](#examples-169)Examples ```bloblang root.a = this.a.not_empty() # In: {"a":"foo"} # Out: {"a":"foo"} # In: {"a":""} # Out: Error("failed assignment (line 1): field `this.a`: string value is empty") # In: {"a":["foo","bar"]} # Out: {"a":["foo","bar"]} # In: {"a":[]} # Out: Error("failed assignment (line 1): field `this.a`: array value is empty") # In: {"a":{"b":"foo","c":"bar"}} # Out: {"a":{"b":"foo","c":"bar"}} # In: {"a":{}} # Out: Error("failed assignment (line 1): field `this.a`: object value is empty") ``` ### [](#not_null)not_null Ensures a value is not null. #### [](#examples-170)Examples ```bloblang root.a = this.a.not_null() # In: {"a":"foobar","b":"barbaz"} # Out: {"a":"foobar"} # In: {"b":"barbaz"} # Out: Error("failed assignment (line 1): field `this.a`: value is null") ``` ### [](#number)number Converts a value to a number with optional fallback. #### [](#parameters-114)Parameters | Name | Type | Description | | --- | --- | --- | | default (optional) | float | An optional value to yield if the target cannot be parsed as a number. | #### [](#examples-171)Examples ```bloblang root.foo = this.thing.number() + 10 root.bar = this.thing.number(5) * 10 ``` ### [](#string)string Converts a value to a string representation. #### [](#examples-172)Examples ```bloblang root.nested_json = this.string() # In: {"foo":"bar"} # Out: {"nested_json":"{\"foo\":\"bar\"}"} ``` ```bloblang root.id = this.id.string() # In: {"id":228930314431312345} # Out: {"id":"228930314431312345"} ``` ### [](#timestamp)timestamp Converts a value to a timestamp with optional fallback. #### [](#parameters-115)Parameters | Name | Type | Description | | --- | --- | --- | | default (optional) | timestamp | An optional value to yield if the target cannot be parsed as a timestamp. | #### [](#examples-173)Examples ```bloblang root.foo = this.ts.timestamp() root.bar = this.none.timestamp(1234567890.timestamp()) ``` ### [](#type)type Returns the type of a value as a string. #### [](#examples-174)Examples ```bloblang root.bar_type = this.bar.type() root.foo_type = this.foo.type() # In: {"bar":10,"foo":"is a string"} # Out: {"bar_type":"number","foo_type":"string"} ``` ```bloblang root.type = this.type() # In: "foobar" # Out: {"type":"string"} # In: 666 # Out: {"type":"number"} # In: false # Out: {"type":"bool"} # In: ["foo", "bar"] # Out: {"type":"array"} # In: {"foo": "bar"} # Out: {"type":"object"} # In: null # Out: {"type":"null"} ``` ```bloblang root.type = content().type() # In: foobar # Out: {"type":"bytes"} ``` ```bloblang root.type = this.ts_parse("2006-01-02").type() # In: "2022-06-06" # Out: {"type":"timestamp"} ``` ## [](#deprecated)Deprecated ### [](#format_timestamp)format_timestamp > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Formats a timestamp as a string using Go’s reference time format. Defaults to RFC 3339 if no format specified. The format uses "Mon Jan 2 15:04:05 -0700 MST 2006" as a reference. Accepts unix timestamps (with decimal precision) or RFC 3339 strings. Use ts\_strftime for strftime-style formats. #### [](#parameters-116)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The output format using Go’s reference time. | | tz (optional) | string | Optional timezone (e.g., 'UTC', 'America/New_York'). Defaults to input timezone or local time for unix timestamps. | ### [](#format_timestamp_strftime)format_timestamp_strftime > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Formats a timestamp as a string using strptime format specifiers (like %Y, %m, %d). Accepts unix timestamps (with decimal precision) or RFC 3339 strings. Supports %f for microseconds. Use ts\_format for Go-style reference time formats. #### [](#parameters-117)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The output format using strptime specifiers. | | tz (optional) | string | Optional timezone. Defaults to input timezone or local time for unix timestamps. | ### [](#format_timestamp_unix)format_timestamp_unix > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Converts a timestamp to a unix timestamp (seconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing seconds. ### [](#format_timestamp_unix_micro)format_timestamp_unix_micro > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Converts a timestamp to a unix timestamp with microsecond precision (microseconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing microseconds. ### [](#format_timestamp_unix_milli)format_timestamp_unix_milli > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Converts a timestamp to a unix timestamp with millisecond precision (milliseconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing milliseconds. ### [](#format_timestamp_unix_nano)format_timestamp_unix_nano > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Converts a timestamp to a unix timestamp with nanosecond precision (nanoseconds since epoch). Accepts unix timestamps or RFC 3339 strings. Returns an integer representing nanoseconds. ### [](#parse_timestamp)parse_timestamp > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Parses a timestamp string using Go’s reference time format and outputs a timestamp object. The format uses "Mon Jan 2 15:04:05 -0700 MST 2006" as a reference - show how this reference time would appear in your format. Use ts\_strptime for strftime-style formats instead. #### [](#parameters-118)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The format of the input string using Go’s reference time. | ### [](#parse_timestamp_strptime)parse_timestamp_strptime > ⚠️ **WARNING** > > This method is deprecated and will be removed in a future version. Parses a timestamp string using strptime format specifiers (like %Y, %m, %d) and outputs a timestamp object. Use ts\_parse for Go-style reference time formats instead. #### [](#parameters-119)Parameters | Name | Type | Description | | --- | --- | --- | | format | string | The format string using strptime specifiers (e.g., %Y-%m-%d). | --- # Page 303: Bloblang Walkthrough **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/walkthrough.md --- # Bloblang Walkthrough > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Bloblang Walkthrough latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/bloblang/walkthrough page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/bloblang/walkthrough.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/bloblang/walkthrough.adoc description: A step by step introduction to Bloblang page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- Bloblang is the most advanced mapping language that you’ll learn from this walkthrough (probably). It is designed for readability, the power to shape even the most outrageous input documents, and to easily make erratic schemas bend to your will. Bloblang is the native mapping language of Redpanda Connect, but it has been designed as a general purpose technology ready to be adopted by other tools. In this walkthrough you’ll learn how to make new friends by mapping their documents, and lose old friends as they grow jealous and bitter of your mapping abilities. There are a few ways to execute Bloblang but the way we’ll do it in this guide is to pull a Redpanda Connect docker image and run the command `rpk connect blobl server`, which opens up an interactive Bloblang editor: ```sh docker pull docker.redpanda.com/redpandadata/connect:latest docker run -p 4195:4195 --rm docker.redpanda.com/redpandadata/connect blobl server --no-open --host 0.0.0.0 ``` Next, open your browser at `http://localhost:4195` and you should see an app with three panels, the top-left is where you paste an input document, the bottom is your Bloblang mapping and on the top-right is the output. ## [](#your-first-assignment)Your first assignment The primary goal of a Bloblang mapping is to construct a brand new document by using an input document as a reference, which we achieve through a series of assignments. Bloblang is traditionally used to map JSON documents and that’s mostly what we’ll be doing in this walkthrough. The first mapping you’ll see when you open the editor is a single assignment: ```bloblang root = this # In: {"message":"hello world"} # Out: {"message":"hello world"} ``` On the left-hand side of the assignment is our assignment target, where `root` is a keyword referring to the root of the new document being constructed. On the right-hand side is a query which determines the value to be assigned, where `this` is a keyword that refers to the context of the mapping which begins as the root of the input document. As you can see the input document in the editor begins as a JSON object `{"message":"hello world"}`, and the output panel should show the result as: ```json { "message": "hello world" } ``` This output is a (neatly formatted) replica of the input document. This is the result of our mapping because we assigned the entire input document to the root of our new thing. Let’s create a brand new document by assigning a fresh object to the root: ```bloblang root = {} root.foo = this.message # In: {"message":"hello world"} # Out: {"foo":"hello world"} ``` Bloblang supports a bunch of [literal types](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/about/#literals), and the first line of this mapping assigns an empty object literal to the root. The second line then creates a new field `foo` on that object by assigning it the value of `message` from the input document. You should see that our output has changed to: ```json { "foo": "hello world" } ``` In Bloblang, when the path that we assign to contains fields that are themselves unset then they are created as empty objects. This rule also applies to `root` itself, which means the mapping: ```bloblang root.foo.bar = this.message root.foo."buz me".baz = "I like mapping" # In: {"message":"hello world"} # Out: {"foo":{"bar":"hello world","buz me":{"baz":"I like mapping"}}} ``` Will automatically create the objects required to produce the output document: ```json { "foo": { "bar": "hello world", "buz me": { "baz": "I like mapping" } } } ``` Also note that we can use quotes in order to express path segments that contain symbols or whitespace. Great, let’s move on quick before our self-satisfaction gets in the way of progress. ## [](#basic-methods-and-functions)Basic methods and functions Nothing is ever good enough for you, why should the input document be any different? Usually in our mappings it’s necessary to mutate values whilst we map them over, this is almost always done with methods, of which [there are many](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/). To demonstrate we’re going to change our mapping to [uppercase](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#uppercase) the field `message` from our input document: ```bloblang root.foo.bar = this.message.uppercase() root.foo."buz me".baz = "I like mapping" # In: {"message":"hello world"} # Out: {"foo":{"bar":"HELLO WORLD","buz me":{"baz":"I like mapping"}}} ``` As you can see the syntax for a method is similar to many languages, simply add a dot on the target value followed by the method name and arguments within brackets. With this method added our output document should look like this: ```json { "foo": { "bar": "HELLO WORLD", "buz me": { "baz": "I like mapping" } } } ``` Since the result of any Bloblang query is a value you can use methods on anything, including other methods. For example, we could expand our mapping of `message` to also replace `WORLD` with `EARTH` using the [`replace_all` method](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#replace_all): ```bloblang root.foo.bar = this.message.uppercase().replace_all("WORLD", "EARTH") root.foo."buz me".baz = "I like mapping" # In: {"message":"hello world"} # Out: {"foo":{"bar":"HELLO EARTH","buz me":{"baz":"I like mapping"}}} ``` As you can see this method required some arguments. Methods support both nameless (like above) and named arguments, which are often literal values but can also be queries themselves. For example try out the following mapping using both named style and a dynamic argument: ```bloblang root.foo.bar = this.message.uppercase().replace_all(old: "WORLD", new: this.message.capitalize()) root.foo."buz me".baz = "I like mapping" # In: {"message":"hello world"} # Out: {"foo":{"bar":"HELLO Hello World","buz me":{"baz":"I like mapping"}}} ``` Woah, I think that’s the plot to Inception, let’s move onto functions. Functions are just boring methods that don’t have a target, and there are [plenty of them as well](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/). Functions are often used to extract information unrelated to the input document, such as [environment variables](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#env), or to generate data such as [timestamps](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#now) or [UUIDs](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#uuid_v4). Since we’re completionists let’s add one to our mapping: ```bloblang root.foo.bar = this.message.uppercase().replace_all("WORLD", "EARTH") root.foo."buz me".baz = "I like mapping" root.foo.id = uuid_v4() # In: {"message":"hello world"} ``` Now I can’t tell you what the output looks like since it will be different each time it’s mapped, how fun! ### [](#deletions)Deletions Everything in Bloblang is an expression to be assigned, including deletions, which is a [function `deleted()`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#deleted). To illustrate let’s create a field we want to delete by changing our input to the following: ```json { "name": "fooman barson", "age": 7, "opinions": ["trucks are cool","trains are cool","chores are bad"] } ``` If we wanted a full copy of this document without the field `name` then we can assign `deleted()` to it: ```bloblang root = this root.name = deleted() # In: {"name":"fooman barson","age":7,"opinions":["trucks are cool","trains are cool","chores are bad"]} # Out: {"age":7,"opinions":["trucks are cool","trains are cool","chores are bad"]} ``` And it won’t be included in the output: ```json { "age": 7, "opinions": [ "trucks are cool", "trains are cool", "chores are bad" ] } ``` An alternative way to delete fields is the [method `without`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#without), our above example could be rewritten as a single assignment `root = this.without("name")`. However, `deleted()` is generally more powerful and will come into play more later on. ## [](#variables)Variables Sometimes it’s necessary to capture a value for later, but we might not want it to be added to the resulting document. In Bloblang we can achieve this with variables which are created using the `let` keyword, and can be referenced within subsequent queries with a dollar sign prefix: ```bloblang let id = uuid_v4() root.id_sha1 = $id.hash("sha1").encode("hex") root.id_md5 = $id.hash("md5").encode("hex") # In: {} ``` Variables can be assigned any value type, including objects and arrays. ## [](#unstructured-and-binary-data)Unstructured and binary data So far in all of our examples both the input document and our newly mapped document are structured, but this does not need to be so. Try assigning some literal value types directly to the `root`, such as a string `root = "hello world"`, or a number `root = 5`. You should notice that when a value type is assigned to the root the output is the raw value, and therefore strings are not quoted. This is what makes it possible to output data of any format, including encrypted, encoded or otherwise binary data. Unstructured mapping is not limited to the output. Rather than referencing the input document with `this`, where it must be structured, it is possible to reference it as a binary string with the [function `content`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#content), try changing your mapping to: ```bloblang root = content().uppercase() # In: hello world # Out: HELLO WORLD ``` When you add content to the input panel, it should be the same in the output panel, but in all uppercase. ## [](#conditionals)Conditionals In order to play around with conditionals let’s set our input to something structured: ```json { "pet": { "type": "cat", "is_cute": true, "treats": 5, "toys": 3 } } ``` In Bloblang all conditionals are expressions, this is a core principal of Bloblang and will be important later on when we’re mapping deeply nested structures. ### [](#if-expression)If expression The simplest conditional is the `if` expression, where the boolean condition does not need to be in parentheses. Let’s create a map that modifies the number of treats our pet receives based on a field: ```bloblang root = this root.pet.treats = if this.pet.is_cute { this.pet.treats + 10 } # In: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} # Out: {"pet":{"type":"cat","is_cute":true,"treats":15,"toys":3}} ``` Try that mapping out and you should see the number of treats in the output increased to 15. Now try changing the input field `pet.is_cute` to `false` and the output treats count should go back to the original 5. When a conditional expression doesn’t have a branch to execute then the assignment is skipped entirely, which means when the pet is not cute the value of `pet.treats` is unchanged (and remains the value set in the `root = this` assignment). We can add an `else` block to our `if` expression to remove treats entirely when the pet is not cute: ```bloblang root = this root.pet.treats = if this.pet.is_cute { this.pet.treats + 10 } else { deleted() } # In: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} # Out: {"pet":{"type":"cat","is_cute":true,"treats":15,"toys":3}} ``` This is possible because field deletions are expressed as assigned values created with the `deleted()` function. ### [](#if-statement)If statement The `if` keyword can also be used as a statement in order to conditionally apply a series of mapping assignments, the previous example can be rewritten as: ```bloblang root = this if this.pet.is_cute { root.pet.treats = this.pet.treats + 10 } else { root.pet.treats = deleted() } # In: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} # Out: {"pet":{"type":"cat","is_cute":true,"treats":15,"toys":3}} ``` Converting this mapping to use a statement has resulted in a more verbose mapping as we had to specify `root.pet.treats` multiple times as an assignment target. However, using `if` as a statement can be beneficial when multiple assignments rely on the same logic: ```bloblang root = this if this.pet.is_cute { root.pet.treats = this.pet.treats + 10 root.pet.toys = this.pet.toys + 10 } # In: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} # Out: {"pet":{"type":"cat","is_cute":true,"treats":15,"toys":13}} ``` More treats _and_ more toys! Lucky Spot! ### [](#match-expression)Match expression Another conditional expression is `match` which allows you to list many branches consisting of a condition and a query to execute separated with `=>`, where the first condition to pass is the one that is executed: ```bloblang root = this root.pet.toys = match { this.pet.treats > 5 => this.pet.treats - 5, this.pet.type == "cat" => 3, this.pet.type == "dog" => this.pet.toys - 3, this.pet.type == "horse" => this.pet.toys + 10, _ => 0, } # In: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} # Out: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} ``` Try executing that mapping with different values for `pet.type` and `pet.treats`. Match expressions can also specify a new context for the keyword `this` which can help reduce some of the boilerplate in your boolean conditions. The following mapping is equivalent to the previous: ```bloblang root = this root.pet.toys = match this.pet { this.treats > 5 => this.treats - 5, this.type == "cat" => 3, this.type == "dog" => this.toys - 3, this.type == "horse" => this.toys + 10, _ => 0, } # In: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} # Out: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} ``` Your boolean conditions can also be expressed as value types, in which case the context being matched will be compared to the value: ```bloblang root = this root.pet.toys = match this.pet.type { "cat" => 3, "dog" => 5, "rabbit" => 8, "horse" => 20, _ => 0, } # In: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} # Out: {"pet":{"type":"cat","is_cute":true,"treats":5,"toys":3}} ``` ## [](#error-handling)Error handling Bloblang can simplify handling errors. First, let’s take a look at what happens when errors _aren’t_ handled, change your input to the following: ```json { "palace_guards": 10, "angry_peasants": "I couldn't be bothered to ask them" } ``` And change your mapping to something simple like a number comparison: ```bloblang root.in_trouble = this.angry_peasants > this.palace_guards # In: {"palace_guards":10,"angry_peasants":"I couldn't be bothered to ask them"} ``` Uh oh! It looks like our canvasser was too lazy and our `angry_peasants` count was incorrectly set for this document. You should see an error in the output window that mentions something like `cannot compare types string (from field this.angry_peasants) and number (from field this.palace_guards)`, which means the mapping was abandoned. So what if we want to try and map something, but don’t care if it fails? In this case if we are unable to compare our angry peasants with palace guards then I would still consider us in trouble just to be safe. For that we have a special [method `catch`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#catch), which if we add to any query allows us to specify an argument to be returned when an error occurs. Since methods can be added to any query we can surround our arithmetic with brackets and catch the whole thing: ```bloblang root.in_trouble = (this.angry_peasants > this.palace_guards).catch(true) # In: {"palace_guards":10,"angry_peasants":"I couldn't be bothered to ask them"} # Out: {"in_trouble":true} ``` Now instead of an error we should see an output with `in_trouble` set to `true`. Try changing to value of `angry_peasants` to a few different values, including some numbers. One of the powerful features of `catch` is that when it is added at the end of a series of expressions and methods it will capture errors at any part of the series, allowing you to capture errors at any granularity. For example, the mapping: ```bloblang root.abort_mission = if this.mission.type == "impossible" { !this.user.motives.contains("must clear name") } else { this.mission.difficulty > 10 }.catch(false) # In: {"mission":{"type":"impossible","difficulty":5},"user":{"motives":["must clear name"]}} # Out: {"abort_mission":false} ``` Will catch errors caused by: - `this.mission.type` not being a string - `this.user.motives` not being an array - `this.mission.difficulty` not being a number But will always return `false` if any of those errors occur. Try it out with this input and play around by breaking some of the fields: ```json { "mission": { "type": "impossible", "difficulty": 5 }, "user": { "motives": ["must clear name"] } } ``` Now try out this mapping: ```bloblang root.abort_mission = if (this.mission.type == "impossible").catch(true) { !this.user.motives.contains("must clear name").catch(false) } else { (this.mission.difficulty > 10).catch(true) } # In: {"mission":{"type":"impossible","difficulty":5},"user":{"motives":["must clear name"]}} # Out: {"abort_mission":false} ``` This version is more granular and will capture each of the errors individually, with each error given a unique `true` or `false` fallback. ## [](#validation)Validation Sometimes errors are what we want. Failing a mapping with an error allows us to handle the bad document in other ways, such as routing it to a dead-letter queue or filtering it entirely. You can read about common Redpanda Connect error handling patterns for bad data in the [error handling guide](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/error_handling/), but the first step is to create the error. Luckily, Bloblang has a range of ways of creating errors under certain circumstances, which can be used in order to validate the data being mapped. There are [a few helper methods](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#type-coercion) that make validating and coercing fields nice and easy, try this mapping out: ```bloblang root.foo = this.foo.number() root.bar = this.bar.not_null() root.baz = this.baz.not_empty() # In: {"foo":5,"bar":"hello world","baz":[1,2,3]} # Out: {"foo":5,"bar":"hello world","baz":[1,2,3]} ``` With some of these sample inputs: ```json {"foo":"nope","bar":"hello world","baz":[1,2,3]} {"foo":5,"baz":[1,2,3]} {"foo":10,"bar":"hello world","baz":[]} ``` However, these methods don’t cover all use cases. The general purpose error throwing technique is the [`throw` function](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/functions/#throw), which takes an argument string that describes the error. When it’s called it will throw a mapping error that abandons the mapping. For example, we can check the type of a field with the [method `type`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#type), and then throw an error if it’s not the type we expected: ```bloblang root.foos = if this.user.foos.type() == "array" { this.user.foos } else { throw("foos must be an array, but it ain't, what gives?") } # In: {"user":{"foos":[1,2,3]}} ``` Try this mapping out with a few sample inputs: ```json {"user":{"foos":[1,2,3]}} {"user":{"foos":"1,2,3"}} ``` ## [](#context)Context In Bloblang, when we refer to the context we’re talking about the value returned with the keyword `this`. At the beginning of a mapping the context starts off as a reference to the root of a structured input document, which is why the mapping `root = this` will result in the same document coming out as you put in. However, in Bloblang there are mechanisms whereby the context might change, we’ve already seen how this can happen within a `match` expression. Another useful way to change the context is by adding a bracketed query expression as a method to a query, which looks like this: ```bloblang root = this.foo.bar.(this.baz + this.buz) # In: {"foo":{"bar":{"baz":1,"buz":2}}} # Out: 3 ``` Within the bracketed query expression the context becomes the result of the query that it’s a method of, so within the brackets in the above mapping the value of `this` points to the result of `this.foo.bar`, and the mapping is therefore equivalent to: ```bloblang root = this.foo.bar.baz + this.foo.bar.buz # In: {"foo":{"bar":{"baz":1,"buz":2}}} # Out: 3 ``` With this handy trick the `throw` mapping from the validation section above could be rewritten as: ```bloblang root.foos = this.user.foos.(if this.type() == "array" { this } else { throw("foos must be an array, but it ain't, what gives?") }) # In: {"user":{"foos":[1,2,3]}} # Out: {"foos":[1,2,3]} ``` ### [](#naming-the-context)Naming the context Shadowing the keyword `this` with new contexts can look confusing in your mappings, and it also limits you to only being able to reference one context at any given time. As an alternative, Bloblang supports context capture expressions that look similar to lambda functions from other languages, where you can name the new context with the syntax ` -> `, which looks like this: ```bloblang root = this.foo.bar.(thing -> thing.baz + thing.buz) # In: {"foo":{"bar":{"baz":1,"buz":2}}} # Out: 3 ``` Within the brackets we now have a new field `thing`, which returns the context that would have otherwise been captured as `this`. This also means the value returned from `this` hasn’t changed and will continue to return the root of the input document. ## [](#coalescing)Coalescing Being able to open up bracketed query expressions on fields leads us onto another cool trick in Bloblang referred to as coalescing. It’s very common in the world of document mapping that due to structural deviations a value that we wish to obtain could come from one of multiple possible paths. To illustrate this problem change the input document to the following: ```json { "thing": { "article": { "id": "foo", "contents": "Some people did some stuff" } } } ``` Let’s say we wish to flatten this structure with the following mapping: ```bloblang root.contents = this.thing.article.contents # In: {"thing":{"article":{"id":"foo","contents":"Some people did some stuff"}}} # Out: {"contents":"Some people did some stuff"} ``` But articles are only one of many document types we expect to receive, where the field `contents` remains the same but the field `article` could instead be `comment` or `share`. In this case we could expand our map of `contents` to use a `match` expression where we check for the existence of `article`, `comment`, etc in the input document. However, a much cleaner way of approaching this is with the pipe operator (`|`), which in Bloblang can be used to join multiple queries, where the first to yield a non-null result is selected. Change your mapping to the following: ```bloblang root.contents = this.thing.article.contents | this.thing.comment.contents # In: {"thing":{"article":{"id":"foo","contents":"Some people did some stuff"}}} # Out: {"contents":"Some people did some stuff"} ``` And now try changing the field `article` in your input document to `comment`. You should see that the value of `contents` remains as `Some people did some stuff` in the output document. Now, rather than write out the full path prefix `this.thing` each time we can use a bracketed query expression to change the context, giving us more space for adding other fields: ```bloblang root.contents = this.thing.(this.article | this.comment | this.share).contents # In: {"thing":{"article":{"id":"foo","contents":"Some people did some stuff"}}} # Out: {"contents":"Some people did some stuff"} ``` And by the way, the keyword `this` within queries can be omitted and made implicit, which allows us to reduce this even further: ```bloblang root.contents = this.thing.(article | comment | share).contents # In: {"thing":{"article":{"id":"foo","contents":"Some people did some stuff"}}} # Out: {"contents":"Some people did some stuff"} ``` Finally, we can also add a pipe operator at the end to fallback to a literal value when none of our candidates exists: ```bloblang root.contents = this.thing.(article | comment | share).contents | "nothing" # In: {"thing":{"article":{"id":"foo","contents":"Some people did some stuff"}}} # Out: {"contents":"Some people did some stuff"} ``` Neat. ## [](#advanced-methods)Advanced methods What happens when you need to map all of the elements of an array? Or filter the keys of an object by their values? What if the fellowship just used the eagles to fly to mount doom? Bloblang offers a bunch of advanced methods for [manipulating structured data types](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#object%E2%80%94%E2%80%8Barray-manipulation), let’s take a quick tour of some of the cooler ones. Set your input document to this list of things: ```json { "num_friends": 5, "things": [ { "name": "yo-yo", "quantity": 10, "is_cool": true }, { "name": "dish soap", "quantity": 50, "is_cool": false }, { "name": "scooter", "quantity": 1, "is_cool": true }, { "name": "pirate hat", "quantity": 7, "is_cool": true } ] } ``` Let’s say we wanted to reduce the `things` in our input document to only those that are cool and where we have enough of them to share with our friends. We can do this with a [`filter` method](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#filter): ```bloblang root = this.things.filter(thing -> thing.is_cool && thing.quantity > this.num_friends) # In: {"num_friends":5,"things":[{"name":"yo-yo","quantity":10,"is_cool":true},{"name":"dish soap","quantity":50,"is_cool":false},{"name":"scooter","quantity":1,"is_cool":true},{"name":"pirate hat","quantity":7,"is_cool":true}]} # Out: [{"name":"yo-yo","quantity":10,"is_cool":true},{"name":"pirate hat","quantity":7,"is_cool":true}] ``` Try running that mapping and you’ll see that the output is reduced. What is happening here is that the `filter` method takes an argument that is a query, and that query will be mapped for each individual element of the array (where the context is changed to the element itself). We have captured the context into a field `thing` which allows us to continue referencing the root of the input with `this`. The `filter` method requires the query parameter to resolve to a boolean `true` or `false`, and if it resolves to `true` the element will be present in the resulting array, otherwise it is removed. Being able to express a query argument to be applied to a range in this way is one of the more powerful features of Bloblang, and when mapping complex structured data these advanced methods will likely be a common tool that you’ll reach for. Another such method is [`map_each`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#map_each), which allows you to mutate each element of an array, or each value of an object. Change your input document to the following: ```json { "talking_heads": [ "1:E.T. is a bad film,Pokemon corrupted an entire generation", "2:Digimon ripped off Pokemon,Cats are boring", "3:I'm important", "4:Science is just made up,The Pokemon films are good,The weather is good" ] } ``` Here we have an array of talking heads, where each element is a string containing an identifer, a colon, and a comma separated list of their opinions. We wish to map each string into a structured object, which we can do with the following mapping: ```bloblang root = this.talking_heads.map_each(raw -> { "id": raw.split(":").index(0), "opinions": raw.split(":").index(1).split(",") }) # In: {"talking_heads":["1:E.T. is a bad film,Pokemon corrupted an entire generation","2:Digimon ripped off Pokemon,Cats are boring","3:I'm important","4:Science is just made up,The Pokemon films are good,The weather is good"]} # Out: [{"id":"1","opinions":["E.T. is a bad film","Pokemon corrupted an entire generation"]},{"id":"2","opinions":["Digimon ripped off Pokemon","Cats are boring"]},{"id":"3","opinions":["I'm important"]},{"id":"4","opinions":["Science is just made up","The Pokemon films are good","The weather is good"]}] ``` The argument to `map_each` is a query where the context is the element, which we capture into the field `raw`. The result of the query argument will become the value of the element in the resulting array, and in this case we return an object literal. In order to separate the identifier from opinions we perform a `split` by colon on the raw string element and get the first substring with the `index` method. We then do the split again and extract the remainder, and split that by comma in order to extract all of the opinions to an array field. However, one problem with this mapping is that the split by colon is written out twice and executed twice. A more efficient way of performing the same thing is with the bracketed query expressions we’ve played with before: ```bloblang root = this.talking_heads.map_each(raw -> raw.split(":").(split_string -> { "id": split_string.index(0), "opinions": split_string.index(1).split(",") })) # In: {"talking_heads":["1:E.T. is a bad film,Pokemon corrupted an entire generation","2:Digimon ripped off Pokemon,Cats are boring","3:I'm important","4:Science is just made up,The Pokemon films are good,The weather is good"]} # Out: [{"id":"1","opinions":["E.T. is a bad film","Pokemon corrupted an entire generation"]},{"id":"2","opinions":["Digimon ripped off Pokemon","Cats are boring"]},{"id":"3","opinions":["I'm important"]},{"id":"4","opinions":["Science is just made up","The Pokemon films are good","The weather is good"]}] ``` > 📝 **NOTE: Challenge!** > > Challenge! > > Try updating that map so that only opinions that mention Pokemon are kept To find more methods for manipulating structured data types check out the [methods page](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#object%E2%80%94%E2%80%8Barray-manipulation). ## [](#reusable-mappings)Reusable mappings Bloblang has cool methods, sure, but there’s nothing cooler than methods you’ve made yourself. When the going gets tough in the mapping world the best solution is often to create a named mapping, which you can do with the keyword `map`: ```bloblang map parse_talking_head { let split_string = this.split(":") root.id = $split_string.index(0) root.opinions = $split_string.index(1).split(",") } root = this.talking_heads.map_each(raw -> raw.apply("parse_talking_head")) # In: {"talking_heads":["1:E.T. is a bad film,Pokemon corrupted an entire generation","2:Digimon ripped off Pokemon,Cats are boring","3:I'm important","4:Science is just made up,The Pokemon films are good,The weather is good"]} # Out: [{"id":"1","opinions":["E.T. is a bad film","Pokemon corrupted an entire generation"]},{"id":"2","opinions":["Digimon ripped off Pokemon","Cats are boring"]},{"id":"3","opinions":["I'm important"]},{"id":"4","opinions":["Science is just made up","The Pokemon films are good","The weather is good"]}] ``` The body of a named map, encapsulated with squiggly brackets, is a totally isolated mapping where `root` now refers to a new value being created for each invocation of the map, and `this` refers to the root of the context provided to the map. Named maps are executed with the [method `apply`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#apply), which has a string parameter identifying the map to execute, this means it’s possible to dynamically select the target map. As you can see in the above example we were able to use a custom map in order to create our talking head objects without the object literal. Within a named map we can also create variables that exist only within the scope of the map. A nice feature of named mappings is that they can invoke themselves recursively, allowing you to define mappings that walk deeply nested structures. The following mapping will scrub all values from a document that contain the word "Voldemort" (case insensitive): ```bloblang map remove_naughty_man { root = match { this.type() == "object" => this.map_each(item -> item.value.apply("remove_naughty_man")), this.type() == "array" => this.map_each(ele -> ele.apply("remove_naughty_man")), this.type() == "string" => if this.lowercase().contains("voldemort") { deleted() }, this.type() == "bytes" => if this.lowercase().contains("voldemort") { deleted() }, _ => this, } } root = this.apply("remove_naughty_man") # In: {"summer_party":{"theme":"the woman in black","guests":["Emma Bunton","the seal I spotted in Trebarwith","Voldemort","The cast of Swiss Army Man","Richard"],"notes":{"lisa":"I don't think voldemort eats fish","monty":"Seals hate dance music"}},"crushes":["Richard is nice but he hates pokemon","Victoria Beckham but I think she's taken","Charlie but they're totally into Voldemort"]} ``` Try running that mapping with the following input document: ```json { "summer_party": { "theme": "the woman in black", "guests": [ "Emma Bunton", "the seal I spotted in Trebarwith", "Voldemort", "The cast of Swiss Army Man", "Richard" ], "notes": { "lisa": "I don't think voldemort eats fish", "monty": "Seals hate dance music" } }, "crushes": [ "Richard is nice but he hates pokemon", "Victoria Beckham but I think she's taken", "Charlie but they're totally into Voldemort" ] } ``` ## [](#unit-testing)Unit testing Redpanda Connect has it’s own [unit testing capabilities](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/unit_testing/) that you can also use for your mappings. To start with save a mapping into a file called something like `naughty_man.blobl`, we can use the example above from the reusable mappings section: ```bloblang map remove_naughty_man { root = match { this.type() == "object" => this.map_each(item -> item.value.apply("remove_naughty_man")), this.type() == "array" => this.map_each(ele -> ele.apply("remove_naughty_man")), this.type() == "string" => if this.lowercase().contains("voldemort") { deleted() }, this.type() == "bytes" => if this.lowercase().contains("voldemort") { deleted() }, _ => this, } } root = this.apply("remove_naughty_man") ``` Next, we can define our unit tests in an accompanying YAML file in the same directory, let’s call this `naughty_man_test.yaml`: ```yaml tests: - name: test naughty man scrubber target_mapping: './naughty_man.blobl' environment: {} input_batch: - content: | { "summer_party": { "theme": "the woman in black", "guests": [ "Emma Bunton", "the seal I spotted in Trebarwith", "Voldemort", "The cast of Swiss Army Man", "Richard" ] } } output_batches: - - json_equals: { "summer_party": { "theme": "the woman in black", "guests": [ "Emma Bunton", "the dolphin I spotted in Trebarwith", "The cast of Swiss Army Man", "Richard" ] } } ``` As you can see we’ve defined a single test, where we point to our mapping file which will be executed in our test. We then specify an input message which is a reduced version of the document we tried out before, and finally we specify output predicates, which is a JSON comparison against the output document. We can execute these tests with `rpk connect test ./naughty_man_test.yaml`, Redpanda Connect will also automatically find our tests if you simply run `rpk connect test ./…​`. You should see an output something like: ```text Test 'naughty_man_test.yaml' failed Failures: --- naughty_man_test.yaml --- test naughty man scrubber [line 2]: batch 0 message 0: json_equals: JSON content mismatch { "summer_party": { "guests": [ "Emma Bunton", "the seal I spotted in Trebarwith" => "the dolphin I spotted in Trebarwith", "The cast of Swiss Army Man", "Richard" ], "theme": "the woman in black" } } ``` Because in actual fact our expected output is wrong, I’ll leave it to you to spot the error. Once the test is fixed you should see: ```text Test 'naughty_man_test.yaml' succeeded ``` And now our mapping, should we need to expand it in the future, is better protected against regressions. You can read more about the Redpanda Connect unit test specification, including alternative output predicates, in [this document](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/unit_testing/). --- # Page 304: Amazon Web Services **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/aws.md --- # Amazon Web Services > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Amazon Web Services latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/cloud/aws page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/cloud/aws.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/cloud/aws.adoc description: Find out about AWS components in Redpanda Connect. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- There are many components within Redpanda Connect which utilize AWS services. You will find that each of these components contains a configuration section under the field `credentials`, of the format: ```yml credentials: profile: "" id: "" secret: "" token: "" role: "" role_external_id: "" ``` This section contains many fields and it isn’t immediately clear which of them are compulsory and which aren’t. This document aims to make it clear what each field is responsible for and how it might be used. ## [](#credentials)Credentials By explicitly setting the credentials you are using at the component level it’s possible to connect to components using different accounts within the same Redpanda Connect process. If you are using long term credentials for your account you only need to set the fields `id` and `secret`: ```yml credentials: id: foo # aws_access_key_id secret: bar # aws_secret_access_key ``` If you are using short term credentials then you will also need to set the field `token`: ```yml credentials: id: foo # aws_access_key_id secret: bar # aws_secret_access_key token: baz # aws_session_token ``` ## [](#assume-a-role)Assume a role It’s also possible to configure Redpanda Connect to [assume a role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use.html) using your credentials by setting the field `role` to your target role ARN. ```yml credentials: role: fooarn # Role ARN ``` This does NOT require explicit credentials, but it’s possible to use both. > 📝 **NOTE** > > Redpanda Cloud pipelines cannot assume roles through the default credential chain (IAM Roles for Service Accounts), and you cannot grant Redpanda-managed IAM roles permission to assume your roles. To assume a role from a Cloud pipeline, set explicit `id` and `secret` credentials for an IAM identity that you own and that has permission to assume the target role, and make sure the target role’s trust policy trusts that identity. Store these credentials as [secrets](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/). If you need to assume a role owned by another organization they might require you to [provide an external ID](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_create_for-user_externalid.html), in which case place it in the field `role_external_id`: ```yml credentials: role: fooarn # Role ARN role_external_id: bar_id ``` --- # Page 305: Ingest Real-Time Sensor Telemetry with the HTTP Gateway **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/gateway.md --- # Ingest Real-Time Sensor Telemetry with the HTTP Gateway > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Ingest Real-Time Sensor Telemetry with the HTTP Gateway latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/cloud/gateway page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/cloud/gateway.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/cloud/gateway.adoc description: Learn how to stream sensor telemetry data into Redpanda Cloud using the gateway input in Redpanda Connect. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-05-26" --- In this guide, you’ll build a pipeline that uses the `gateway` input to receive real-time telemetry data from sensors over HTTP. Each incoming message is normalized, published to a Redpanda topic, and acknowledged back to the sender. This setup is ideal for IoT, mobile, and embedded systems that need to stream data to Redpanda Cloud without using a Kafka client. The `gateway` input exposes a secure HTTP endpoint, simplifying ingestion from devices. Because HTTP is universally supported, it’s easier to integrate on constrained devices, microcontrollers, or languages that don’t support Kafka natively. Additional benefits: - **Simplified security**: Devices authenticate with Redpanda Cloud API tokens (using Bearer headers). No need to embed Kafka credentials, manage TLS, or expose brokers publicly. - **Operational flexibility**: Devices are decoupled from Kafka internals like topics or schemas. You can evolve pipeline logic without touching device code. - **Automatic provisioning**: Redpanda Cloud generates a secure endpoint URL when you deploy the pipeline. ## [](#prerequisites)Prerequisites - A Redpanda Cloud cluster (Serverless, Dedicated, or BYOC) - cURL or another compatible HTTP client ## [](#create-a-sensor-user-in-redpanda-cloud)Create a sensor user in Redpanda Cloud A sensor user is required to securely authenticate and manage access to the `sensor.telemetry` topic, ensuring that only authorized devices can produce messages to the topic. 1. [Log in to Redpanda Cloud](https://cloud.redpanda.com). 2. Go to **Topics** and create a topic named `sensor.telemetry`. This topic will be used to store incoming telemetry messages. 3. Go to **Security** and create a user with the following details: - **Username**: `sensor-sasl-user` - **Password**: `` (choose a secure password) - **SASL Mechanism**: `SCRAM-SHA-256` 4. Copy the password and save it securely for the next step. 5. Go to **Secrets Store** and create a new secret named `SENSOR_SASL_PASSWORD` with the value of the password you set for the user. - Set the scope of the secret to Redpanda Cluster and Redpanda Connect. 6. Go to **Security > ACLs** and create an access policy for the `sensor-sasl-user` user. This policy should allow the user to produce messages to the `sensor.telemetry` topic. ## [](#create-a-service-account)Create a service account The service account is used to authenticate requests to the gateway endpoint. It provides a secure way to manage access to the gateway without embedding sensitive credentials in your devices. 1. [Create a new service account](https://cloud.redpanda.com/service-accounts/new) in Redpanda Cloud named `sensor-ingest` and give it a description like "Service account for sensor telemetry ingestion". 2. Copy the client ID and secret. 3. Request a new API token for the service account. This token will be used to authenticate requests to the gateway. ```bash curl --request POST \ --url 'https://auth.prd.cloud.redpanda.com/oauth/token' \ --header 'content-type: application/x-www-form-urlencoded' \ --data grant_type=client_credentials \ --data client_id= \ --data client_secret= \ --data audience=cloudv2-production.redpanda.cloud ``` Replace `` and `` with the values you copied from the service account. The request response provides an access token that remains **valid for one hour**. 4. Set the access token as an environment variable: ```bash export CLOUD_API_TOKEN= ``` ## [](#create-a-redpanda-cloud-pipeline)Create a Redpanda Cloud pipeline 1. Go to **Connect** and click **Create Pipeline**. 2. Name the pipeline `sensor-telemetry-ingest` and give it a description like "Ingest real-time sensor telemetry data". 3. Paste the following pipeline configuration into the editor: ```yaml input: gateway: rate_limit: "limit" rate_limit_resources: - label: limit local: count: 100 interval: 1s pipeline: processors: - bloblang: | root.sensor_id = this.sensor_id root.type = this.type root.value = this.value root.unit = this.unit root.received_at = now() output: broker: pattern: fan_out_sequential outputs: - redpanda: seed_brokers: - ${REDPANDA_BROKERS} topic: sensor.telemetry tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: sensor-sasl-user password: ${secrets.SENSOR_SASL_PASSWORD} - sync_response: processors: - mapping: | root = { "status": "ok", "received_at": now() } ``` This pipeline listens for incoming telemetry messages over HTTP and processes each one in real time. Here’s what each section does: - `input.gateway`: Defines the input source. It exposes a secure HTTP endpoint that devices can post to. The optional `rate_limit` named `limit` is applied to protect the pipeline from overload. - `rate_limit_resources.limit`: Limits traffic to 100 requests per second. If this rate is exceeded, HTTP requests are rejected with a 429 response. - `pipeline.processors.bloblang`: Normalizes the incoming message by copying fields and adding a `received_at` timestamp (using the current time). - `output.broker`: Uses a `fan_out_sequential` pattern to send each message to two outputs: - The first output publishes the normalized message to the `sensor.telemetry` Redpanda topic. - The second output sends a synchronous JSON response back to the sender confirming receipt. 4. Click **Start**. The pipeline starts deploying. When the state changes to "Running", the pipeline is ready to accept incoming messages. 5. Click the pipeline to view its details. When the pipeline is deployed, a URL is displayed. This is the HTTP endpoint to which you’ll post sensor data. 6. Copy the URL. ## [](#send-sensor-data)Send sensor data Send test data using cURL. Replace `` with the URL provided by Redpanda Cloud when you deployed the pipeline. ```bash curl -X POST \ -H "Authorization: Bearer $CLOUD_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sensor_id": "thermo-42", "type": "temperature", "value": 21.7, "unit": "C" }' ``` Expected response: ```json { "received_at":"2025-06-17T09:48:50.986719231Z", "sensor_id":"thermo-42", "type":"temperature", "unit":"C", "value":21.7 } ``` You can verify that the message was successfully ingested by checking the `sensor.telemetry` topic in Redpanda Cloud. To verify that the rate limit is working, try sending more than 100 requests per second. You should receive a 429 response with a `Retry-After` header indicating when to retry. ```bash seq 1 300 | xargs -n1 -P50 -I{} curl -s -o /dev/null -w "%{http_code}\n" \ -X POST \ -H "Authorization: Bearer $CLOUD_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sensor_id":"test", "value": 42}' ``` You should see a mixture of `200` and `429` responses, indicating that the rate limit is being enforced. ## [](#monitor-the-pipeline)Monitor the pipeline You can monitor the pipeline’s logs in the Redpanda Cloud UI. 1. Go to **Connect** and select the `sensor-telemetry-ingest` pipeline. 2. Click on the **Logs** tab to view real-time logs of the pipeline’s activity. You can see any errors that occur during processing. ## [](#next-steps)Next steps - Filter or enrich events with conditional Bloblang. - Route messages by `sensor.type` to different topics. ## [](#suggested-reading)Suggested reading - [`gateway` input reference](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gateway/) - [Bloblang functions](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/interpolation/) - [Redpanda Cloud API authentication](https://docs.redpanda.com/api/doc/cloud-dataplane/authentication) --- # Page 306: Google Cloud Platform **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/gcp.md --- # Google Cloud Platform > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Google Cloud Platform latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/cloud/gcp page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/cloud/gcp.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/cloud/gcp.adoc description: Find out about GCP components in Redpanda Connect. page-git-created-date: "2024-09-09" page-git-modified-date: "2026-08-11" --- There are many components within Redpanda Connect which utilize Google Cloud Platform (GCP) services. You will find that each of these components require valid credentials. When running Redpanda Connect inside a Google Cloud environment that has a [default service account](https://cloud.google.com/iam/docs/service-accounts#default), it can automatically retrieve the service account credentials to call Google Cloud APIs through a library called Application Default Credentials (ADC). Otherwise, if your application runs outside Google Cloud environments that provide a default service account, you need to manually create one. Once you have a service account set up which has the required permissions, you can [create](https://console.cloud.google.com/apis/credentials/serviceaccountkey) a new Service Account Key and download it as a JSON file. Then all you need to do set the path to this JSON file in the `GOOGLE_APPLICATION_CREDENTIALS` environment variable. Please refer to [this document](https://cloud.google.com/docs/authentication/production) for details. --- # Page 307: Migrate to the Unified Redpanda Migrator **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/migrate-unified-redpanda-migrator.md --- # Migrate to the Unified Redpanda Migrator > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Migrate to the Unified Redpanda Migrator latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/migrate-unified-redpanda-migrator page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/migrate-unified-redpanda-migrator.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/migrate-unified-redpanda-migrator.adoc description: Learn how to migrate from legacy migrator components to the unified `redpanda_migrator` input/output pair in Redpanda Connect 4.67.5+. page-git-created-date: "2025-10-24" page-git-modified-date: "2026-05-26" --- > ❗ **IMPORTANT** > > This page is about migrating to a newer version of Redpanda Connect. For information about migrating your data using Redpanda Migrator, see [Redpanda Migrator](https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/redpanda_migrator/). This guide explains how to migrate from legacy migrator components (`redpanda_migrator_bundle`, `legacy_redpanda_migrator` and `legacy_redpanda_migrator_offsets`) to the unified `redpanda_migrator` input/output pair introduced in Redpanda Connect 4.67.5+. The unified migrator consolidates all migration logic into a single input/output pair, simplifying configuration and improving reliability. ## [](#overview)Overview | Available in | Redpanda Connect 4.67.5+ | | --- | --- | | Legacy status | Deprecated in 4.67.5, removed in 4.85.0 | | Compatibility | Not backward-compatible | | Configuration model | One input and one output, paired by label | | Primary control | All migration logic resides in the output component | Key concepts: - Components are paired by matching `label` values. - The input defines the source cluster and schema registry. - The output defines the destination cluster, schema registry, and migration behavior. - Topic mapping and consumer group migration are configured in the output. ## [](#architectural-changes)Architectural changes ### [](#legacy-architecture)Legacy architecture A complex bundle (`redpanda_migrator_bundle`) that managed three subcomponents: - `redpanda_migrator`: Data transfer - `schema_registry`: Schema synchronization - `redpanda_migrator_offsets`: Consumer group offsets This design required complex internal routing and sequencing. ### [](#unified-architecture)Unified architecture A single `redpanda_migrator` input/output pair replaces the bundle: - **Input**: Consumes from the source Kafka cluster. - **Output**: Handles topic creation, schema synchronization, ACLs, and consumer group offsets. Benefits: - Simplified setup: all configuration consolidated in one output component. - Improved coordination: no internal routing or wrapper logic. - Enhanced control: fine-grained schema and topic options, improved offset handling. ## [](#migration-steps)Migration steps Follow this checklist in order to ensure a safe, low-risk migration. - Back up your existing configurations. - Add new `input.redpanda_migrator` and `output.redpanda_migrator` components with matching labels. - Move source Kafka and Schema Registry settings to the input. - Move destination Kafka and Schema Registry settings to the output. - Replace `topic_prefix` with `topic` using interpolation syntax. - Move offset settings to `output.redpanda_migrator.consumer_groups`. - Remove deprecated fields. - Validate configuration with `rpk connect lint`. - Test using non-production topics first. - Monitor logs and performance during migration. - Remove legacy configuration after successful migration. ## [](#field-mapping-reference)Field mapping reference ### [](#bundle-wrapper-redpanda_migrator_bundle)Bundle wrapper (`redpanda_migrator_bundle`) #### [](#input-mapping)Input mapping | Legacy Field | New Location | Status | Notes | | --- | --- | --- | --- | | redpanda_migrator | input.redpanda_migrator | Moved | Source cluster connection | | schema_registry | input.redpanda_migrator.schema_registry | Moved | Source schema registry | | migrate_schemas_before_data | - | Removed | Controlled by output schema interval | | consumer_group_offsets_poll_interval | output.redpanda_migrator.consumer_groups.interval | Moved | Now controls sync frequency | #### [](#output-mapping)Output mapping | Legacy Field | New Location | Status | Notes | | --- | --- | --- | --- | | redpanda_migrator | output.redpanda_migrator | Moved | Destination cluster configuration | | schema_registry | output.redpanda_migrator.schema_registry | Moved | Destination schema registry | | translate_schema_ids | output.redpanda_migrator.schema_registry.translate_ids | Moved | Schema ID translation | | input_bundle_label | label | Replaced | Input and output paired by label | ### [](#data-migration-fields)Data migration fields | Legacy Field | New Location | Status | Notes | | --- | --- | --- | --- | | All (*) | input.redpanda_migrator.* | Moved | Direct mapping | | topics (explicit list) | input.redpanda_migrator.topics | Unchanged | Still supported for explicit lists | | regexp_topics: true | input.redpanda_migrator.regexp_topics_include, regexp_topics_exclude | Deprecated | Use include/exclude arrays for pattern-based selection | | topic_prefix | output.redpanda_migrator.topic | Replaced | Use interpolation, for example 'prefix_${! @kafka_topic }' | | replication_factor_override, replication_factor | output.redpanda_migrator.topic_replication_factor | Replaced | Unified field | | input_resource | label | Replaced | Label pairing replaces internal routing | | - | output.redpanda_migrator.provenance_header | New | Optional header for tracking message source cluster | ### [](#schema-migration-fields)Schema migration fields | Legacy Field | New Location | Status | Notes | | --- | --- | --- | --- | | Connection fields | input.redpanda_migrator.schema_registry.* | Moved | Source schema registry | | subject_filter | output.redpanda_migrator.schema_registry.include, exclude | Replaced | Use regex lists for filtering | | include_deleted | output.redpanda_migrator.schema_registry.include_deleted | Moved | Configured on destination | | backfill_dependencies | output.redpanda_migrator.schema_registry.versions | Replaced | Choose all or latest | ### [](#consumer-group-offset-migration)Consumer group offset migration The `redpanda_migrator_offsets` pair is replaced by the `consumer_groups` block in the output. | Legacy Component | New Location | Status | Notes | | --- | --- | --- | --- | | redpanda_migrator_offsets (input/output) | output.redpanda_migrator.consumer_groups | Replaced | Unified control block | ## [](#migration-example)Migration example The following example demonstrates a complete migration from legacy to unified components. Legacy configuration ```yaml input: label: "source_cluster" redpanda_migrator_bundle: legacy_redpanda_migrator: seed_brokers: [ "source-kafka:9092" ] topics: [ "orders", "payments" ] consumer_group: "migration_group" schema_registry: url: "http://source-registry:8081" migrate_schemas_before_data: false consumer_group_offsets_poll_interval: 30s output: redpanda_migrator_bundle: legacy_redpanda_migrator: seed_brokers: [ "destination-redpanda:9092" ] topic_prefix: "migrated_" schema_registry: url: "http://destination-registry:8081" translate_schema_ids: true input_bundle_label: "source_cluster" ``` Unified configuration ```yaml input: label: "migration_pipeline" (1) redpanda_migrator: # Source Kafka settings seed_brokers: [ "source-kafka:9092" ] # Pattern-based topic selection (for migrating all topics except system topics) # Note: You can still use explicit lists: topics: [ "orders", "payments" ] regexp_topics_include: [ '.' ] (2) regexp_topics_exclude: [ '^_' ] (3) consumer_group: "migration_group" # Source Schema Registry settings schema_registry: url: "http://source-registry:8081" output: label: "migration_pipeline" (4) redpanda_migrator: # Destination Redpanda settings seed_brokers: [ "destination-redpanda:9092" ] # Topic mapping (replaces topic_prefix) topic: 'migrated_${! @kafka_topic }' (5) # Add source cluster tracking header provenance_header: "x-source-cluster" (6) # Destination Schema Registry and migration settings schema_registry: url: "http://destination-registry:8081" translate_ids: true # Rename subjects subject: 'migrated_${! metadata("schema_registry_subject") }' # Consumer group migration settings consumer_groups: enabled: true interval: 30s (7) ``` | 1 | Labels are now used for pairing input and output. | | --- | --- | | 2 | Match all topics using regex pattern. | | 3 | Exclude internal/system topics starting with underscore. | | 4 | Matching label pairs the input and output components. | | 5 | Use interpolation syntax to replicate topic_prefix behavior. | | 6 | Adds a header to track which cluster messages originated from, useful for debugging and auditing. | | 7 | Replaces consumer_group_offsets_poll_interval. | ## [](#validation)Validation Before running, validate your configuration: ```bash rpk connect lint config.yaml ``` Then test on a small set of topics before running full migrations. ## [](#troubleshooting)Troubleshooting | Problem | Likely Cause | Solution | | --- | --- | --- | | Labels do not match | Input and output labels differ | Use identical, case-sensitive labels. | | Topic interpolation errors | Incorrect syntax | Use topic: 'prefix_${! @kafka_topic }' with quotes and !. | | Schema registry connection fails | Incorrect registry placement | The source registry must be in the input. The destination registry must be in the output. | | Consumer group migration not working | Missing consumer_groups.enabled: true | Ensure consumer group migration is explicitly enabled. | ## [](#after-migration)After migration After verifying that the new migrator works as expected: - Remove legacy configuration files. - Update internal documentation and runbooks. - Train your team on the new configuration model. - See the [`redpanda_migrator` output](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_migrator/) reference for advanced configuration options. --- # Page 308: Synchronous Responses **URL**: https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/sync_responses.md --- # Synchronous Responses > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Synchronous Responses latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: connect/guides/sync_responses page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: connect/guides/sync_responses.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/connect/guides/sync_responses.adoc description: Understand synchronous response handling in Redpanda Connect, ensuring reliable and efficient data processing. page-git-created-date: "2025-06-25" page-git-modified-date: "2026-08-11" --- In a regular Redpanda Connect pipeline, messages flow in one direction and acknowledgements in the other: ```text ----------- Message -------------> Input (AMQP) -> Processors -> Output (AMQP) <------- Acknowledgement --------- ``` However, Redpanda Connect supports bidirectional protocols like HTTP and WebSocket, which allow responses to be returned directly from the pipeline. For example, HTTP is a request/response protocol, and inputs like `http_server` (Self-Managed) or `gateway` (Redpanda Cloud) support returning response payloads to the requester. ```text --------- Request Body --------> Input (HTTP) -> Processors -> Output (Sync Response) <--- Response Body (and ack) --- ``` ## [](#routing-processed-messages-back)Routing processed messages back To return a processed response, use the [`sync_response`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/sync_response/) output. Use the `gateway` input in Redpanda Cloud: ```yaml input: gateway: {} pipeline: processors: - mapping: | root = { city: json("location"), forecast: "Clear skies with light winds", temperature_c: 22 } output: sync_response: {} ``` Sending this request: ```json { "location": "Berlin" } ``` Returns: ```json { "city": "Berlin", "forecast": "Clear skies with light winds", "temperature_c": 22 } ``` ## [](#combine-with-other-outputs)Combine with other outputs You can route processed messages to storage and return a response using a [`broker`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/broker/) output. ```yaml input: gateway: {} output: broker: pattern: fan_out outputs: - redpanda: seed_brokers: - ${REDPANDA_BROKERS} topic: weather.requests tls: enabled: true sasl: - mechanism: SCRAM-SHA-256 username: ${secrets.USERNAME} password: ${secrets.PASSWORD} - sync_response: processors: - mapping: | root = { status: "received", received_at: now() } ``` ## [](#returning-partially-processed-messages)Returning partially processed messages You can return a response before the message is fully processed by using the [`sync_response` processor](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/sync_response/). This allows continued processing after the response is set. ```yaml pipeline: processors: - mapping: root = "Received weather report for %s".format(json("location")) - sync_response: {} - mapping: root.reported_at = now() ``` This returns `"Received weather report for Berlin"` to the client, but continues modifying the message before storing or forwarding it. > 📝 **NOTE** > > Due to delivery guarantees, the response is not sent until all downstream processing and acknowledgements are complete. --- # Page 309: Consume Data **URL**: https://docs.redpanda.com/cloud-data-platform/develop/consume-data.md --- # Consume Data > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Consume Data latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: consume-data/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: consume-data/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/consume-data/index.adoc description: Learn about consumer offsets and follower fetching. page-git-created-date: "2024-07-25" page-git-modified-date: "2024-08-01" --- - [Consumer Offsets](consumer-offsets/) Redpanda uses an internal topic, `__consumer_offsets`, to store committed offsets from each Kafka consumer that is attached to Redpanda. - [Follower Fetching](follower-fetching/) Learn about follower fetching and how to configure a Redpanda consumer to fetch records from the closest replica. - [Paginate Messages in Redpanda Console](paginate-messages-events/) Retrieve more than the default batch of messages in Redpanda Console by paging through larger result sets. --- # Page 310: Consumer Offsets **URL**: https://docs.redpanda.com/cloud-data-platform/develop/consume-data/consumer-offsets.md --- # Consumer Offsets > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Consumer Offsets latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: consume-data/consumer-offsets page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: consume-data/consumer-offsets.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/consume-data/consumer-offsets.adoc description: Redpanda uses an internal topic, __consumer_offsets, to store committed offsets from each Kafka consumer that is attached to Redpanda. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- In Redpanda, all messages are organized by [topic](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#topic) and distributed across multiple partitions, based on a [partition strategy](https://www.redpanda.com/guides/kafka-tutorial-kafka-partition-strategy). For example, when using the round robin strategy, a producer writing to a topic with five partitions would distribute approximately 20% of the messages to each [partition](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#partition). Within a partition, each message (once accepted and acknowledged by the partition leader) is permanently assigned a unique sequence number called an [offset](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#offset). Offsets enable consumers to resume processing from a specific point, such as after an application outage. If an outage prevents your application from receiving events, you can use the consumer offset to retrieve only the events that occurred during the downtime. By default, the first message in a partition is assigned offset 0, the next is offset 1, and so on. You can manually specify a specific start value for offsets if needed. Once assigned, offsets are immutable, ensuring that the order of messages within a partition is preserved. ## [](#how-consumers-use-offsets)How consumers use offsets As a consumer reads messages from Redpanda, it can save its progress by “committing the offset” (known as an [offset commit](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#offset-commit)), an action initiated by the consumer, not Redpanda. Kafka client libraries provide an API for committing offsets, which communicates with Redpanda using the [consumer group](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#consumer-group) API. Each committed offset is stored as a message in the `__consumer_offsets` topic, which is a private Redpanda topic that stores committed offsets from each Kafka consumer attached to Redpanda, allowing the consumer to resume processing from the last committed point. Redpanda exposes the `__consumer_offsets` key to enable the many tools in the Kafka ecosystem that rely on this value for their operation, providing greater ecosystem interoperability with environments and applications. When a consumer group works together to consume data from topics, the partitions are divided among the consumers in the group. For example, if a topic has 12 partitions, and there are two consumers, each consumer would be assigned six partitions to consume. If a new consumer starts later and joins this consumer group, a rebalance occurs, such that each consumer ends up with four partitions to consume. You specify a consumer group by setting the `group.id` property to a unique name for the group. Kafka tracks the maximum offset it has consumed in each partition and can commit offsets to ensure it can resume processing from the same point in the event of a restart. Kafka allows offsets for a consumer group to be stored on a designated broker, known as the group coordinator. All consumers in the group send their offset commits and fetch requests to this group coordinator. > 📝 **NOTE** > > More advanced consumers can read data from Redpanda without using a consumer group by requesting to read a specific topic, partition, and offset range. This pattern is often used by stream processing systems such as Apache Spark and Apache Flink, which have their own mechanisms for assigning work to consumers. Redpanda Console derives its consumer group lists and lag information from committed consumer group offsets, so these consumers appear in Console only if they also commit offsets to a consumer group. For example, Spark Structured Streaming does not commit offsets, so its consumers never appear. Flink commits offsets only when it is configured to do so on checkpoint completion. When the group coordinator receives an OffsetCommitRequest, it appends the request to the [compacted](https://kafka.apache.org/documentation/#compaction) Kafka topic `__consumer_offsets`. The broker sends a successful offset commit response to the consumer only after all the replicas of the offsets topic receive the offsets. If the offsets fail to replicate within a configurable timeout, the offset commit fails and the consumer may retry the commit after backing off. The brokers periodically compact the `__consumer_offsets` topic, because it only needs to maintain the most recent offset commit for each partition. The coordinator also caches the offsets in an in-memory table to serve offset fetches quickly. ## [](#commit-strategies)Commit strategies There are several strategies for managing offset commits: ### [](#automatic-offset-commit)Automatic offset commit Auto commit is the default commit strategy, where the client automatically commits offsets at regular intervals. This is set with the `enable.auto.commit` property. The client then commits offsets every `auto.commit.interval.ms` milliseconds. The primary advantage of the auto commit approach is its simplicity. After it is configured, the consumer requires no additional effort. Commits are managed in the background. However, the consumer is unaware of what was committed or when. As a result, after an application restart, some messages may be reprocessed (since consumption resumes from the last committed offset, which may include already-processed messages). The strategy guarantees at-least-once delivery. > 📝 **NOTE** > > If your consume configuration is set up to consume and write to another data store, and the write to that datastore fails, the consumer might not recover when it is auto-committed. It may not only duplicate messages, but could also drop messages intended to be in another datastore. Make sure you understand the trade-off possibilities associated with this default behavior. ### [](#manual-offset-commit)Manual offset commit The manual offset commit strategy gives consumers greater control over when commits occur. This approach is typically used when a consumer needs to align commits with an external system, such as database transactions in an RDBMS. The main advantage of manual commits is that they allow you to decide exactly when a record is considered consumed. You can use two API calls for this: `commitSync` and `commitAsync`, which differ in their blocking behavior. #### [](#synchronous-commit)Synchronous commit The advantage of synchronous commits is that consumers can take appropriate action before continuing to consume messages, albeit at the expense of increased latency (while waiting for the commit to return). The commit (`commitSync`) will also retry automatically, until it either succeeds or receives an unrecoverable error. The following example shows a synchronous commit: ```java consumer.subscribe(Arrays.asList("foo", "bar")); while (true) { ConsumerRecords records = consumer.poll(100); for (ConsumerRecord record : records) { // process records here ... // ... and at the appropriate point, call commit (not after every message) consumer.commitSync(); } } ``` #### [](#asynchronous-commit)Asynchronous commit The advantage of asynchronous commits is lower latency, because the consumer does not pause to wait for the commit response. However, there is no automatic retry of the commit (`commitAsync`) if it fails. There is also increased coding complexity (due to the asynchronous callbacks). The following example shows an asynchronous commit in which the consumer will not block. Instead, the commit call registers a callback, which is executed once the commit returns: ```java void callback() { // executed when the commit returns } consumer.subscribe(Arrays.asList("foo", "bar")); while (true) { ConsumerRecords records = consumer.poll(100); for (ConsumerRecord record : records) { // process records here ... // ... and at the appropriate point, call commit consumer.commitAsync(callback); } } ``` ### [](#external-offset-management)External offset management The external offset management strategy allows consumers to manage offsets independently of Redpanda. In this approach: - Consumers bypass the consumer group API and directly assign partitions instead of subscribing to a topic. - Offsets are not committed to Redpanda, but are instead stored in an external storage system. Because consumers that use this strategy do not commit offsets to Redpanda, they do not appear in the consumer group lists in Redpanda Console. This is expected behavior for systems such as Apache Spark Structured Streaming, which manages offsets entirely in its own checkpoint storage. Systems such as Apache Flink can also commit offsets to Kafka when a checkpoint completes. In that configuration, they follow the hybrid strategy described in the next section and do appear in Console. To implement an external offset management strategy: 1. Set `enable.auto.commit` to `false`. 2. Use `assign(Collection)` to assign partitions. 3. Use the offset provided with each ConsumerRecord to save your position. 4. Upon restart, use `seek(TopicPartition, long)` to restore the position of the consumer. ### [](#hybrid-offset-management)Hybrid offset management The hybrid offset management strategy allows consumers to handle their own consumer rebalancing while still leveraging Redpanda’s offset commit functionality. In this approach: - Consumers bypass the consumer group API and directly assign partitions instead of subscribing to a topic. - Offsets are committed to Redpanda. ## [](#offset-commit-best-practices)Offset commit best practices Follow these best practices to optimize offset commits. ### [](#avoid-over-committing)Avoid over-committing The purpose of a commit is to save consumer progress. More frequent commits reduce the amount of data to re-read after an application restart, as the commit interval directly affects the recovery point objective (RPO). Because a lower RPO is desirable, application designers may believe that committing frequently is a good design choice. However, committing too frequently can result in adverse consequences. While individually small, each commit still results in a message being written to the `__consumer_offsets` topic, because the position of the consumer against every partition must be recorded. At high commit rates, this workload can become a bottleneck for both the client and the server. Additionally, many Kafka client implementations do not coalesce offset commits, meaning redundant commits in a backlog still need to be processed. In many Kafka client implementations, offset commits aren’t coalesced at the client; so if a backlog of commits forms (when using the asynchronous commit API), the earlier commits still need to be processed, even though they are effectively redundant. **Best practice**: Monitor commit latency to ensure commits are timely. If you notice performance issues, commit less frequently. ### [](#use-unique-consumer-groups)Use unique consumer groups Like many topics, the consumer group topic has multiple partitions to help with performance. When writing commit messages, Redpanda groups all of the commits for a consumer group into a specific partition to maintain ordering. Reusing a consumer group across multiple applications, even for different topics, forces all commits to use a single partition, negating the benefits of partitioning. **Best practice**: Assign a unique consumer group to each application to distribute the commit load across all partitions. ### [](#tune-the-consumer-group)Tune the consumer group In highly parallel applications, frequent consumer group heartbeats can create unnecessary overhead. For example, 3,200 consumers checking every 500 milliseconds generate 6,400 heartbeats per second. You can optimize this behavior by increasing the `heartbeat.interval.ms` (along with `session.timeout.ms`). **Best practice**: Adjust heartbeat and session timeout settings to reduce unnecessary overhead in large-scale applications. --- # Page 311: Follower Fetching **URL**: https://docs.redpanda.com/cloud-data-platform/develop/consume-data/follower-fetching.md --- # Follower Fetching > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Follower Fetching latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: consume-data/follower-fetching page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: consume-data/follower-fetching.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/consume-data/follower-fetching.adoc description: Learn about follower fetching and how to configure a Redpanda consumer to fetch records from the closest replica. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- Learn about follower fetching and how to configure a Redpanda consumer to fetch records from the closest replica. ## [](#about-follower-fetching)About follower fetching **Follower fetching** enables a consumer to fetch records from the closest replica of a topic partition, regardless of whether it’s a leader or a follower. For a Redpanda cluster deployed across different data centers and availability zones (AZs), restricting a consumer to fetch only from the leader of a partition can incur greater costs and have higher latency than fetching from a follower that is geographically closer to the consumer. With follower fetching (proposed in [KIP-392](https://cwiki.apache.org/confluence/display/KAFKA/KIP-392%3A+Allow+consumers+to+fetch+from+closest+replica)), the fetch protocol is extended to support a consumer fetching from any replica. This includes [Remote Read Replicas](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/remote-read-replicas/). The first fetch from a consumer is processed by a Redpanda leader broker. The leader checks for a replica (itself or a follower) that has a rack ID that matches the consumer’s rack ID. If a replica with a matching rack ID is found, the fetch request returns records from that replica. Otherwise, the fetch is handled by the leader. ## [](#configure-follower-fetching)Configure follower fetching Redpanda decides which replica a consumer fetches from. If the consumer configures its `client.rack` property, Redpanda by default selects a replica from the same rack as the consumer, if available. For each consumer, set the `client.rack` property to a rack ID. Rack awareness is pre-enabled for cloud-based clusters in multi-AZ environments. ## [](#suggested-videos)Suggested videos - [YouTube - Redpanda Office Hour: Follower Fetching (52 mins)](https://www.youtube.com/watch?v=wV6gH5_yVaw&ab_channel=RedpandaData) --- # Page 312: Paginate Messages in Redpanda Console **URL**: https://docs.redpanda.com/cloud-data-platform/develop/consume-data/paginate-messages-events.md --- # Paginate Messages in Redpanda Console > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Paginate Messages in Redpanda Console latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: consume-data/paginate-messages-events page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: consume-data/paginate-messages-events.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/consume-data/paginate-messages-events.adoc description: Retrieve more than the default batch of messages in Redpanda Console by paging through larger result sets. page-git-created-date: "2026-04-30" page-git-modified-date: "2026-05-26" --- By default, the **Messages** tab on a topic returns the number of records set in **Max results**. Enable **Continuous Pagination** when you need to inspect a topic beyond that cap. ## [](#browse-all-messages-in-a-topic)Browse all messages in a topic 1. Go to **Topics** and select a topic. 2. Open the **Messages** tab. 3. (Optional) Set **Start offset** and **Max results**, or apply filters, to narrow the records you want to inspect. See [Programmable Push Filters](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/programmable-push-filters/) and [Deserialization](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/record-deserialization/). 4. Enable the **Continuous Pagination** toggle. 5. Scroll the message list. Redpanda Cloud keeps loading records until you reach the end of the topic. When continuous pagination is on, the max results cap no longer limits the browsing session. ## [](#performance-considerations)Performance considerations Retrieving large result sets increases load on the Redpanda Cloud backend and the cluster. To keep responses fast: - Narrow the result set with filters or a bounded offset range before enabling continuous pagination. - Use [JavaScript push filters](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/programmable-push-filters/) to match only the records you need. - Leave continuous pagination off and rely on max results when you only need a sample. --- # Page 313: Data Transforms **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms.md --- # Data Transforms > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Data Transforms latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/index.adoc description: Learn about WebAssembly data transforms within Redpanda Cloud. page-git-created-date: "2025-04-08" page-git-modified-date: "2025-04-08" --- - [How Data Transforms Work](how-transforms-work/) Learn how Redpanda data transforms work. - [Develop Data Transforms](build/) Learn how to initialize a data transforms project and write transform functions in your chosen language. - [Configure Data Transforms](configure/) Learn how to configure data transforms in Redpanda, including editing the `transform.yaml` file, environment variables, and memory settings. This topic covers both the configuration of transform functions and the WebAssembly (Wasm) engine's environment. - [Deploy Data Transforms](deploy/) Learn how to build, deploy, share, and troubleshoot data transforms in Redpanda. - [Write Integration Tests for Transform Functions](test/) Learn how to write integration tests for data transform functions in Redpanda, including setting up unit tests and using testcontainers for integration tests. - [Monitor Data Transforms](monitor/) This topic provides guidelines on how to monitor the health of your data transforms and view logs. - [Manage Data Transforms](data-transforms/) You can monitor the status and performance metrics of your transform functions. You can also view detailed logs and delete transform functions when they are no longer needed. --- # Page 314: Develop Data Transforms **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build.md --- # Develop Data Transforms > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Develop Data Transforms latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/build page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/build.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/build.adoc description: Learn how to initialize a data transforms project and write transform functions in your chosen language. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- > 📝 **NOTE** > > Data transforms are supported on BYOC and Dedicated clusters running Redpanda version 24.3 and later. > 💡 **TIP: When to use Redpanda Connect instead** > > Data transforms do not access external networks or disks, and are best for lightweight data preparation (filtering, scrubbing, schema/format conversion). Use [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/) when you need any of the following: > > - External integration (HTTP services, databases, cloud storage) for enrichment or fan-out to third-party systems > > - Batching or windowed processing for grouping/aggregation > > - Prebuilt processors and connectors to reduce custom code Learn how to initialize a data transforms project and write transform functions in your chosen language. After reading this page, you will be able to: - Initialize a data transforms project using the rpk CLI - Build transform functions that process records and write to output topics - Implement multi-topic routing patterns with Schema Registry integration ## [](#prerequisites)Prerequisites You must have the following development tools installed on your host machine: - The [`rpk` command-line client](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) installed. - For Golang projects, you must have at least version 1.20 of [Go](https://go.dev/doc/install). - For Rust projects, you must have the latest stable version of [Rust](https://rustup.rs/). - For JavaScript and TypeScript projects, you must have the [latest long-term-support release of Node.js](https://nodejs.org/en/download/package-manager). ## [](#enable-data-transforms)Enable data transforms Data transforms are disabled on all clusters by default. Before you can deploy data transforms to a cluster, you must first enable the feature with the `rpk` command-line tool. To enable data transforms, set the [`data_transforms_enabled`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#data_transforms_enabled) cluster property to `true`: ```bash rpk cluster config set data_transforms_enabled true ``` > 📝 **NOTE** > > This property requires a rolling restart, and it can take several minutes for the update to complete. ## [](#init)Initialize a data transforms project To initialize a data transforms project, use the following command to set up the project files in your current directory. This command adds the latest version of the [SDK](https://docs.redpanda.com/cloud-data-platform/reference/data-transforms/sdks/) as a project dependency: ```bash rpk transform init --language= --name= ``` If you do not include the `--language` flag, the command prompts you for the language. Supported languages include: - `tinygo-no-goroutines` (does not include [Goroutines](https://golangdocs.com/goroutines-in-golang)) - `tinygo-with-goroutines` - `rust` - `javascript` - `typescript` For example, if you choose `tinygo-no-goroutines`, `rpk` creates the following project files: . ├── go.mod ├── go.sum ├── README.md ├── transform.go └── transform.yaml The `transform.go` file contains a boilerplate transform function. The `transform.yaml` file specifies the configuration settings for the transform function. See also: [Configure Data Transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/) ## [](#build-transform-functions)Build transform functions You can develop your transform logic with one of the available SDKs that allow your transform code to interact with a Redpanda cluster. #### Go All transform functions must register a callback with the `OnRecordWritten()` method. You should run any initialization steps in the `main()` function because it’s only run once when the transform function is first deployed. You can also use the standard predefined [`init()` function](https://go.dev/doc/effective_go#init). ```go package main import ( "github.com/redpanda-data/redpanda/src/transform-sdk/go/transform" ) func main() { // Register your transform function. // This is a good place to perform other setup too. transform.OnRecordWritten(myTransform) } // myTransform is where you read the record that was written, and then you can // output new records that will be written to the destination topic func myTransform(event transform.WriteEvent, writer transform.RecordWriter) error { return writer.Write(event.Record()) } ``` #### Rust All transform functions must register a callback with the `on_record_written()` method. You should run any initialization steps in the `main()` function because it’s only run once when the transform function is first deployed. ```rust use redpanda_transform_sdk::*; fn main() { // Register your transform function. // This is a good place to perform other setup too. on_record_written(my_transform); } // my_transform is where you read the record that was written, and then you can // return new records that will be written to the output topic fn my_transform(event: WriteEvent, writer: &mut RecordWriter) -> Result<(), Box> { writer.write(event.record)?; Ok(()) } ``` #### JavaScript All transform functions must register a callback with the `onRecordWritten()` method. You should run any initialization steps outside of the callback so that they are only run once when the transform function is first deployed. ```js // src/index.js import { onRecordWritten } from "@redpanda-data/transform-sdk"; // This is a good place to perform setup steps. // Register your transform function. onRecordWritten((event, writer) => { // This is where you read the record that was written, and then you can // output new records that will be written to the destination topic writer.write(event.record); }); ``` If you need to use Node.js standard modules in your transform function, you must configure the [`polyfillNode` plugin](https://github.com/cyco130/esbuild-plugin-polyfill-node) for [esbuild](https://esbuild.github.io/). This plugin allows you to polyfill Node.js APIs that are not natively available in the Redpanda JavaScript runtime environment. `esbuild.js` ```js import * as esbuild from 'esbuild'; import { polyfillNode } from 'esbuild-plugin-polyfill-node'; await esbuild.build({ plugins: [ polyfillNode({ globals: { buffer: true, // Allow a global Buffer variable if referenced. process: false, // Don't inject the process global, the Redpanda JavaScript runtime does that. }, polyfills: { crypto: true, // Enable crypto polyfill // Add other polyfills as needed }, }), ], }); ``` ### [](#errors)Error handling By distinguishing between recoverable and critical errors, you can ensure that your transform functions are both resilient and robust. Handling recoverable errors internally helps maintain continuous operation, while allowing critical errors to escape ensures that the system can address severe issues effectively. Redpanda tracks the offsets of records that transform functions have processed. If an error escapes the Wasm virtual machine (VM), the VM will fail. When the Wasm engine detects this failure and starts a new VM, the transform function retries processing the input topics from the last processed offset, potentially leading to repeated failures if the underlying issue is not resolved. Handling errors internally by logging them and continuing to process subsequent records can help maintain continuous operation. However, this approach can result in silently discarding problematic records, which may lead to unnoticed data loss if the logs are not monitored closely. #### Go ```go package main import ( "log" "github.com/redpanda-data/redpanda/src/transform-sdk/go/transform" ) func main() { transform.OnRecordWritten(myTransform) } func myTransform(event transform.WriteEvent, writer transform.RecordWriter) error { record := event.Record() if record.Key == nil { // Handle the error internally by logging it log.Println("Error: Record key is nil") // Skip this record and continue to process other records return nil } // Allow errors with writes to escape return writer.Write(record) } ``` #### Rust ```rust use redpanda_transform_sdk::*; use log::error; fn main() { // Set up logging env_logger::init(); on_record_written(my_transform); } fn my_transform(event: WriteEvent, writer: &mut RecordWriter) -> anyhow::Result<()> { let record = event.record; if record.key().is_none() { // Handle the error internally by logging it error!("Error: Record key is nil"); // Skip this record and continue to process other records return Ok(()); } // Allow errors with writes to escape return writer.write(record) } ``` #### JavaScript ```js import { onRecordWritten } from "@redpanda-data/transform-sdk"; // Register your transform function. onRecordWritten((event, writer) => { const record = event.record; if (!record.key) { // Handle the error internally by logging it console.error("Error: Record key is nil"); // Skip this record and continue to process other records return; } // Allow errors with writes to escape writer.write(record); }); ``` When you deploy this transform function, and produce a message without a key, you’ll get the following in the logs: ```js { "body": { "stringValue": "2024/06/20 08:17:33 Error: Record key is nil\n" }, "timeUnixNano": 1718871455235337000, "severityNumber": 13, "attributes": [ { "key": "transform_name", "value": { "stringValue": "test" } }, { "key": "node", "value": { "intValue": 0 } } ] } ``` You can view logs for transform functions using the `rpk transform logs ` command. To ensure that you are notified of any errors or issues in your data transforms, Redpanda provides metrics that you can use to monitor the state of your data transforms. See also: - [View logs for transform functions](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/monitor/#logs) - [Monitor data transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/monitor/) - [Configure transform logging](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/#log) - [`rpk transform logs` reference](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-transform/rpk-transform-logs/) ### [](#avoid-state-management)Avoid state management Relying on in-memory state across transform invocations can lead to inconsistencies and unpredictable behavior. Data transforms operate with at-least-once semantics, meaning a transform function might be executed more than once for a given record. Redpanda may also restart a transform function at any point, which causes its state to be lost. ### [](#env-vars)Access environment variables You can access both [built-in and custom environment variables](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/#environment-variables) in your transform function. In this example, environment variables are checked once during initialization: #### Go ```go package main import ( "fmt" "os" "github.com/redpanda-data/redpanda/src/transform-sdk/go/transform" ) func main() { // Check environment variables before registering the transform function. outputTopic1, ok := os.LookupEnv("REDPANDA_OUTPUT_TOPIC_1") if ok { fmt.Printf("Output topic 1: %s\n", outputTopic1) } else { fmt.Println("Only one output topic is set") } // Register your transform function. transform.OnRecordWritten(myTransform) } func myTransform(event transform.WriteEvent, writer transform.RecordWriter) error { return writer.Write(event.Record()) } ``` #### Rust ```rust use redpanda_transform_sdk::*; use std::env; use log::error; fn main() { // Set up logging env_logger::init(); // Check environment variables before registering the transform function. match env::var("REDPANDA_OUTPUT_TOPIC_1") { Ok(output_topic_1) => println!("Output topic 1: {}", output_topic_1), Err(_) => println!("Only one output topic is set"), } // Register your transform function. on_record_written(my_transform); } fn my_transform(_event: WriteEvent, _writer: &mut RecordWriter) -> anyhow::Result<()> { Ok(()) } ``` #### JavaScript ```js import { onRecordWritten } from "@redpanda-data/transform-sdk"; // Check environment variables before registering the transform function. const outputTopic1 = process.env.REDPANDA_OUTPUT_TOPIC_1; if (outputTopic1) { console.log(`Output topic 1: ${outputTopic1}`); } else { console.log("Only one output topic is set"); } // Register your transform function. onRecordWritten((event, writer) => { return writer.write(event.record); }); ``` ### [](#write-to-specific-output-topics)Write to specific output topics You can configure your transform function to write records to specific output topics based on message content, enabling powerful routing and fan-out patterns. This capability is useful for: - Filtering messages by criteria and routing to different topics - Fan-out patterns that distribute data from one input topic to multiple output topics - Event routing based on message type or schema - Data distribution for downstream consumers Wasm transforms provide a simpler alternative to external connectors like Kafka Connect for in-broker data routing, with lower latency and no additional infrastructure to manage. #### [](#basic-json-validation-example)Basic JSON validation example The following example shows a filter that outputs only valid JSON from the input topic into the output topic. The transform writes invalid JSON to a different output topic. ##### Go ```go import ( "encoding/json" "github.com/redpanda-data/redpanda/src/transform-sdk/go/transform" ) func main() { transform.OnRecordWritten(filterValidJson) } func filterValidJson(event transform.WriteEvent, writer transform.RecordWriter) error { if json.Valid(event.Record().Value) { return writer.Write(event.Record()) } // Send invalid records to separate topic return writer.Write(event.Record(), transform.ToTopic("invalid-json")) } ``` ##### Rust ```rust use anyhow::Result; use redpanda_transform_sdk::*; fn main() { on_record_written(filter_valid_json); } fn filter_valid_json(event: WriteEvent, writer: &mut RecordWriter) -> Result<()> { let value = event.record.value().unwrap_or_default(); if serde_json::from_slice::(value).is_ok() { writer.write(event.record)?; } else { // Send invalid records to separate topic writer.write_with_options(event.record, WriteOptions::to_topic("invalid-json"))?; } Ok(()) } ``` ##### JavaScript The JavaScript SDK does not support writing records to a specific output topic. #### [](#multi-topic-fanout)Multi-topic fan-out with Schema Registry This example shows how to route batched updates from a single input topic to multiple output topics based on a routing field in each message. Messages are encoded with the [Schema Registry wire format](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/#wire-format) for validation against the output topic schema. Consider using this pattern with Iceberg-enabled topics to fan out data directly into lakehouse tables. Input message example ```json { "updates": [ {"table": "orders", "data": {"order_id": "123", "amount": 99.99}}, {"table": "inventory", "data": {"product_id": "P456", "quantity": 50}}, {"table": "customers", "data": {"customer_id": "C789", "name": "Jane"}} ] } ``` [Configure the transform](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/) with multiple output topics: ```yaml name: event-router input_topic: events output_topics: - orders - inventory - customers ``` The transform extracts each update and routes it to the appropriate topic based on the `table` field. Schemas are registered dynamically in the `main()` function using the Schema Registry client, which returns the schema IDs needed for encoding messages in the wire format. > 📝 **NOTE** > > In this example, it is assumed that you have created the output topics and have the schema definitions ready. The transform registers the schemas dynamically on startup using the `{topic-name}-value` naming convention for schema subjects (for example, `orders-value`, `inventory-value`). ##### Go `go.mod` ```go module fanout-example go 1.20 require github.com/redpanda-data/redpanda/src/transform-sdk/go/transform v1.1.0 // v1.1.0+ required ``` `transform.go`: ```go package main import ( "encoding/binary" "encoding/json" "log" "github.com/redpanda-data/redpanda/src/transform-sdk/go/transform" "github.com/redpanda-data/redpanda/src/transform-sdk/go/transform/sr" ) // Input message structure with array of updates type BatchMessage struct { Updates []TableUpdate `json:"updates"` } // Individual table update with routing field type TableUpdate struct { Table string `json:"table"` // Routing field - determines output topic Data json.RawMessage `json:"data"` // The actual data to write } // Schema IDs for each output topic, registered dynamically at startup var schemaIDs = make(map[string]int) func main() { // Create Schema Registry client client := sr.NewClient() // Define schemas for each output topic schemas := map[string]string{ "orders": `{"type":"record","name":"Order","fields":[{"name":"order_id","type":"string"},{"name":"amount","type":"double"}]}`, "inventory": `{"type":"record","name":"Inventory","fields":[{"name":"product_id","type":"string"},{"name":"quantity","type":"int"}]}`, "customers": `{"type":"record","name":"Customer","fields":[{"name":"customer_id","type":"string"},{"name":"name","type":"string"}]}`, } // Register schemas and store their IDs for topic, schemaStr := range schemas { subject := topic + "-value" schema := sr.Schema{ Schema: schemaStr, Type: sr.TypeAvro, } result, err := client.CreateSchema(subject, schema) if err != nil { log.Fatalf("Failed to register schema for %s: %v", topic, err) } schemaIDs[topic] = result.ID log.Printf("Registered schema for %s with ID %d", topic, result.ID) } log.Printf("Starting fanout transform with schema IDs: %v", schemaIDs) transform.OnRecordWritten(routeUpdates) } func routeUpdates(event transform.WriteEvent, writer transform.RecordWriter) error { var batch BatchMessage if err := json.Unmarshal(event.Record().Value, &batch); err != nil { log.Printf("Failed to parse batch message: %v", err) return nil // Skip invalid records } // Process each update in the batch for i, update := range batch.Updates { schemaID, exists := schemaIDs[update.Table] if !exists { log.Printf("Unknown table in update %d: %s", i, update.Table) continue } if err := writeUpdate(update, schemaID, writer, event); err != nil { log.Printf("Failed to write update %d to %s: %v", i, update.Table, err) } } return nil } func writeUpdate(update TableUpdate, schemaID int, writer transform.RecordWriter, event transform.WriteEvent) error { // Create Schema Registry wire format: [magic_byte, schema_id (4 bytes BE), data...] value := make([]byte, 5) value[0] = 0 // magic byte binary.BigEndian.PutUint32(value[1:5], uint32(schemaID)) value = append(value, update.Data...) record := transform.Record{ Key: event.Record().Key, Value: value, } return writer.Write(record, transform.ToTopic(update.Table)) } ``` ##### Rust `Cargo.toml` ```toml [package] name = "fanout-rust-example" version = "0.1.0" edition = "2021" [dependencies] redpanda-transform-sdk = "1.1.0" # v1.1.0+ required for WriteOptions API redpanda-transform-sdk-sr = "1.1.0" serde = { version = "1", features = ["derive"] } serde_json = "1" log = "0.4" env_logger = "0.11" [profile.release] opt-level = "z" lto = true strip = true ``` `src/main.rs`: ```rust use redpanda_transform_sdk::*; use redpanda_transform_sdk_sr::{SchemaRegistryClient, Schema, SchemaFormat}; use serde::Deserialize; use std::collections::HashMap; use std::error::Error; use std::sync::OnceLock; use log::{info, error}; #[derive(Deserialize)] struct BatchMessage { updates: Vec, } #[derive(Deserialize)] struct TableUpdate { table: String, data: serde_json::Value, } // Schema IDs for each output topic, registered dynamically at startup static SCHEMA_IDS: OnceLock> = OnceLock::new(); fn main() { // Initialize logging env_logger::init(); // Create Schema Registry client let mut client = SchemaRegistryClient::new(); // Define schemas for each output topic let schemas = [ ("orders", r#"{"type":"record","name":"Order","fields":[{"name":"order_id","type":"string"},{"name":"amount","type":"double"}]}"#), ("inventory", r#"{"type":"record","name":"Inventory","fields":[{"name":"product_id","type":"string"},{"name":"quantity","type":"int"}]}"#), ("customers", r#"{"type":"record","name":"Customer","fields":[{"name":"customer_id","type":"string"},{"name":"name","type":"string"}]}"#), ]; let mut schema_ids = HashMap::new(); // Register schemas and store their IDs for (topic, schema_str) in schemas { let subject = format!("{}-value", topic); let schema = Schema::new(schema_str.to_string(), SchemaFormat::Avro, vec![]); match client.create_schema(&subject, schema) { Ok(result) => { let id = result.id(); // SchemaId type schema_ids.insert(topic.to_string(), id.0); // Extract i32 from SchemaId wrapper info!("Registered schema for {} with ID {}", topic, id.0); } Err(e) => { error!("Failed to register schema for {}: {}", topic, e); panic!("Schema registration failed"); } } } let _ = SCHEMA_IDS.set(schema_ids); info!("Starting fanout transform with schema IDs"); on_record_written(route_updates); } fn write_update( update: &TableUpdate, schema_id: i32, writer: &mut RecordWriter, event: &WriteEvent, ) -> Result<(), Box> { // Create Schema Registry wire format: [magic_byte, schema_id (4 bytes BE), data...] let mut value = vec![0u8; 5]; value[0] = 0; // magic byte value[1..5].copy_from_slice(&schema_id.to_be_bytes()); let data_bytes = serde_json::to_vec(&update.data)?; value.extend_from_slice(&data_bytes); let key = event.record.key().map(|k| k.to_vec()); let record = BorrowedRecord::new(key.as_deref(), Some(&value)); writer.write_with_options(record, WriteOptions::to_topic(&update.table))?; Ok(()) } fn route_updates(event: WriteEvent, writer: &mut RecordWriter) -> Result<(), Box> { let batch: BatchMessage = serde_json::from_slice(event.record.value().unwrap_or_default())?; let schema_ids = SCHEMA_IDS.get().unwrap(); for update in batch.updates.iter() { if let Some(&schema_id) = schema_ids.get(&update.table) { write_update(update, schema_id, writer, &event)?; } } Ok(()) } ``` ##### JavaScript The JavaScript SDK does not support writing records to specific output topics. For multi-topic fan-out, use the Go or Rust SDK. ### [](#connect-to-the-schema-registry)Connect to the Schema Registry You can use the Schema Registry client library to read and write schemas as well as serialize and deserialize records. This client library is useful when working with schema-based topics in your data transforms. See also: - [Redpanda Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/) - [Go Schema Registry client reference](https://docs.redpanda.com/cloud-data-platform/reference/data-transforms/golang-sdk/) - [Rust Schema Registry client reference](https://docs.redpanda.com/cloud-data-platform/reference/data-transforms/rust-sdk/) - [JavaScript Schema Registry client reference](https://docs.redpanda.com/cloud-data-platform/reference/data-transforms/js/js-sdk-sr/) ## [](#next-steps)Next steps [Configure Data Transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/) ## [](#suggested-reading)Suggested reading - [How Data Transforms Work](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/how-transforms-work/) - [Data Transforms SDKs](https://docs.redpanda.com/cloud-data-platform/reference/data-transforms/sdks/) - [`rpk transform` commands](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-transform/rpk-transform/) --- # Page 315: Configure Data Transforms **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure.md --- # Configure Data Transforms > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Data Transforms latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/configure page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/configure.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/configure.adoc description: Learn how to configure data transforms in Redpanda, including editing the transform.yaml file, environment variables, and memory settings. This topic covers both the configuration of transform functions and the WebAssembly (Wasm) engine's environment. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- Learn how to configure data transforms in Redpanda, including editing the `transform.yaml` file, environment variables, and memory settings. This topic covers both the configuration of transform functions and the WebAssembly (Wasm) engine’s environment. ## [](#configure-transform-functions)Configure transform functions This section covers how to configure transform functions using the `transform.yaml` configuration file, command-line overrides, and environment variables. ### [](#config-file)Transform configuration file When you [initialize](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build/#init) a data transforms project, a `transform.yaml` file is generated in the provided directory. You can use this configuration file to configure the transform function with settings, including input and output topics, the language used for the data transform, and any environment variables. - `name`: The name of the transform function. - `description`: A description of what the transform function does. - `input-topic`: The topic from which data is read. - `output-topics`: A list of up to eight topics to which the transformed data is written. - `language`: The language used for the transform function. The language is set to the one you defined during [initialization](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build/#init). - `env`: A dictionary of custom environment variables that are passed to the transform function. Do not prefix keys with `REDPANDA_`. Check the list of all [limitations](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/how-transforms-work/#limitations). Here is an example of a transform.yaml file: ```yaml name: redpanda-example description: | This transform function is an example to demonstrate how to configure data transforms in Redpanda. input-topic: example-input-topic output-topics: - example-output-topic-1 - example-output-topic-2 language: tinygo-no-goroutines env: DATA_TRANSFORMS_ARE_AWESOME: 'true' ``` ### [](#cl)Override configurations with command-line options You can set the name of the transform function, environment variables, and input and output topics on the command-line when you deploy the transform. These command-line settings take precedence over those specified in the `transform.yaml` file. See [Deploy Data Transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/deploy/) ### [](#built-in)Built-In environment variables As well as custom environment variables set in either the [command-line](#cl) or the [configuration file](#config-file), Redpanda makes some built-in environment variables available to your transform functions. These variables include: - `REDPANDA_INPUT_TOPIC`: The input topic specified. - `REDPANDA_OUTPUT_TOPIC_0..REDPANDA_OUTPUT_TOPIC_N`: The output topics in the order specified on the command line or in the configuration file. For example, `REDPANDA_OUTPUT_TOPIC_0` is the first variable, `REDPANDA_OUTPUT_TOPIC_1` is the second variable, and so on. Transform functions are isolated from the broker’s internal environment variables to maintain security and encapsulation. Each transform function only uses the environment variables explicitly provided to it. ## [](#configure-the-wasm-engine)Configure the Wasm engine This section covers how to configure the Wasm engine environment using Redpanda cluster configuration properties. ### [](#enable-transforms)Enable data transforms To use data transforms, you must enable it for a Redpanda cluster using the [`data_transforms_enabled`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#data_transforms_enabled) property. ### [](#log)Configure transform logging The following properties configure logging for data transforms: - [`data_transforms_logging_line_max_bytes`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#data_transforms_logging_line_max_bytes): Increase this value if your log messages are frequently truncated. Setting this value too low may truncate important log information. ## [](#next-steps)Next steps [Deploy Data Transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/deploy/) --- # Page 316: Manage Data Transforms **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/data-transforms.md --- # Manage Data Transforms > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Data Transforms latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/data-transforms page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/data-transforms.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/data-transforms.adoc description: You can monitor the status and performance metrics of your transform functions. You can also view detailed logs and delete transform functions when they are no longer needed. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- You can monitor the status and performance metrics of your transform functions. You can also view detailed logs and delete transform functions when they are no longer needed. ## [](#prerequisites)Prerequisites Before you begin, ensure that you have the following: - [Data transforms enabled](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/#enable-transforms) in your Redpanda cluster. - At least one transform function deployed to your Redpanda cluster. ## [](#monitor)Monitor transform functions To monitor transform functions: 1. Navigate to the **Transforms** menu. 2. Click the name of a transform function to view detailed information: - The partitions that the function is running on - The broker (node) ID - Any lag (the amount of pending records on the input topic that have yet to be processed by the transform) ## [](#logs)View logs To view logs for a transform function: 1. Navigate to the **Transforms** menu. 2. Click on the name of a transform function. 3. Click the **Logs** tab to see the logs. Redpanda Cloud displays a limited number of logs for transform functions. To view the full history of logs, use the [`rpk` command-line tool](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/monitor/#logs). ## [](#delete)Delete transform functions To delete a transform function: 1. Navigate to the **Transforms** menu. 2. Find the transform function you want to delete from the list. 3. Click the delete icon at the end of the row. 4. Confirm the deletion when prompted. Deleting a transform function will remove it from the cluster and stop any further processing. ## [](#suggested-reading)Suggested reading - [How Data Transforms Work](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/how-transforms-work/) - [Deploy Data Transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/deploy/) - [Monitor Data Transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/monitor/) --- # Page 317: Deploy Data Transforms **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/deploy.md --- # Deploy Data Transforms > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Deploy Data Transforms latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/deploy page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/deploy.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/deploy.adoc description: Learn how to build, deploy, share, and troubleshoot data transforms in Redpanda. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- Learn how to build, deploy, share, and troubleshoot data transforms in Redpanda. ## [](#prerequisites)Prerequisites Before you begin, ensure that you have the following: - [Data transforms enabled](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/#enable-transforms) in your Redpanda cluster. - The [`rpk` command-line client](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). - A [data transform](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build/) project. ## [](#build)Build the Wasm binary To build a Wasm binary: 1. Ensure your project directory contains a `transform.yaml` file. 2. Build the Wasm binary using the [`rpk transform build`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-transform/rpk-transform-build/) command. ```bash rpk transform build ``` You should now have a Wasm binary named `.wasm`, where `` is the name specified in your `transform.yaml` file. This binary is your data transform function, ready to be deployed to a Redpanda cluster or hosted on a network for others to use. ## [](#deploy)Deploy the Wasm binary You can deploy your transform function using the [`rpk transform deploy`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-transform/rpk-transform-deploy/) command. 1. Validate your setup against the pre-deployment checklist: - Do you meet the [Prerequisites](#prerequisites)? - Does your transform function access any environment variables? If so, make sure to set them in the `transform.yaml` file or in the command-line when you deploy the binary. - Do your configured input and output topics already exist? Input and output topics must exist in your Redpanda cluster before you deploy the Wasm binary. 2. Deploy the Wasm binary: ```bash rpk transform deploy ``` When the transform function reaches Redpanda, it starts processing new records that are written to the input topic. ### [](#reprocess)Reprocess records In some cases, you may need to reprocess records from an input topic that already contains data. Processing existing records can be useful, for example, to process historical data into a different format for a new consumer, to re-create lost data from a deleted topic, or to resolve issues with a previous version of a transform that processed data incorrectly. To reprocess records, you can specify the starting point from which the transform function should process records in each partition of the input topic. The starting point can be either a partition offset or a timestamp. > 📝 **NOTE** > > The `--from-offset` flag is only effective the first time you deploy a transform function. On subsequent deployments of the same function, Redpanda resumes processing from the last committed offset. To reprocess existing records using an existing function, [delete the function](#delete) and redeploy it with the `--from-offset` flag. To deploy a transform function and start processing records from a specific partition offset, use the following syntax: ```bash rpk transform deploy --from-offset +/- ``` In this example, the transform function will start processing records from the beginning of each partition of the input topic: ```bash rpk transform deploy --from-offset +0 ``` To deploy a transform function and start processing records from a specific timestamp, use the following syntax: ```bash rpk transform deploy --from-timestamp @ ``` In this example, the transform function will start processing from the first record in each partition of the input topic that was committed after the given timestamp: ```bash rpk transform deploy --from-timestamp @1617181723 ``` ### [](#share-wasm-binaries)Share Wasm binaries You can also deploy data transforms on a Redpanda cluster by providing an addressable path to the Wasm binary. This is useful for sharing transform functions across multiple clusters or teams within your organization. For example, if the Wasm binary is hosted at `https://my-site/my-transform.wasm`, use the following command to deploy it: ```bash rpk transform deploy --file=https://my-site/my-transform.wasm ``` ## [](#edit-existing-transform-functions)Edit existing transform functions To make changes to an existing transform function: 1. [Make your changes to the code](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build/). 2. [Rebuild](#build) the Wasm binary. 3. [Redeploy](#deploy) the Wasm binary to the same Redpanda cluster. When you redeploy a Wasm binary with the same name, it will resume processing from the last offset it had previously processed. If you need to [reprocess existing records](#reprocess), you must delete the transform function, and redeploy it with the `--from-offset` flag. Deploy-time configuration overrides must be provided each time you redeploy a Wasm binary. Otherwise, they will be overwritten by default values or the configuration file’s contents. ## [](#delete)Delete a transform function To delete a transform function, use the following command: ```bash rpk transform delete ``` For more details about this command, see [rpk transform delete](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-transform/rpk-transform-delete/). > 💡 **TIP** > > You can also delete transform functions in Redpanda Cloud. ## [](#troubleshoot)Troubleshoot This section provides guidance on how to diagnose and troubleshoot issues with building or deploying data transforms. ### [](#invalid-transform-environment)Invalid transform environment This error means that one or more of your configured custom environment variables are invalid. Check your custom environment variables against the list of [limitations](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/how-transforms-work/#limitations). ### [](#invalid-webassembly)Invalid WebAssembly This error indicates that the binary is missing a required callback function: Invalid WebAssembly - the binary is missing required transform functions. Check the broker support for the version of the data transforms SDK being used. All transform functions must register a callback with the `OnRecordWritten()` method. For more details, see [Develop Data Transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build/). ## [](#next-steps)Next steps [Set up monitoring](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/monitor/) for data transforms. --- # Page 318: How Data Transforms Work **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/how-transforms-work.md --- # How Data Transforms Work > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: How Data Transforms Work latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/how-transforms-work page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/how-transforms-work.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/how-transforms-work.adoc description: Learn how Redpanda data transforms work. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- > 📝 **NOTE** > > Data transforms are supported on BYOC and Dedicated clusters running Redpanda version 24.3 and later. Redpanda provides the framework to build and deploy inline transformations (data transforms) on data written to Redpanda topics, delivering processed and validated data to consumers in the format they expect. Redpanda does this directly inside the broker, eliminating the need to manage a separate stream processing environment or use third-party tools. ![Data transforms in a broker](https://docs.redpanda.com/cloud-data-platform/shared/_images/wasm1.png) Data transforms let you run common data streaming tasks, like filtering, scrubbing, and transcoding, within Redpanda. For example, you may have consumers that require you to redact credit card numbers or convert JSON to Avro. Data transforms can also interact with the Redpanda Schema Registry to work with encoded data types. To learn how to build and deploy data transforms, see [How Data Transforms Work](./). ## [](#data-transforms-with-webassembly)Data transforms with WebAssembly Data transforms use [WebAssembly](https://webassembly.org/) (Wasm) engines inside a Redpanda broker, allowing Redpanda to control the entire transform lifecycle. For example, Redpanda can stop and start transforms when partitions are moved or to free up system resources for other tasks. Data transforms take data from an input topic and map it to one or more output topics. For each topic partition, a leader is responsible for handling the data. Redpanda runs a Wasm virtual machine (VM) on the same CPU core (shard) as these partition leaders to execute the transform function. Transform functions are the specific implementations of code that carry out the transformations. They read data from input topics, apply the necessary processing logic, and write the transformed data to output topics. To execute a transform function, Redpanda uses just-in-time (JIT) compilation to compile the bytecode in memory, write it to an executable space, then run the directly translated machine code. This JIT compilation ensures efficient execution of the machine code, as it is tailored to the specific hardware it runs on. When you deploy a data transform to a Redpanda broker, it stores the Wasm bytecode and associated metadata, such as input and output topics and environment variables. The broker then replicates this data across the cluster using internal Kafka topics. When the data is distributed, each shard runs its own instance of the transform function. This process includes several resource management features: - Each shard can run only one instance of the transform function at a time to ensure efficient resource utilization and prevent overload. - CPU time is dynamically allocated to the Wasm runtime to ensure that the code does not run forever and cannot block the broker from handling traffic or doing other work, such as Tiered Storage uploads. ## [](#flow-of-data-transforms)Flow of data transforms When a shard becomes the leader of a given partition on the input topic of one or more active transforms, Redpanda does the following: 1. Spins up a Wasm VM using the JIT-compiled Wasm module. 2. Pushes records from the input partition into the Wasm VM. 3. Writes the output. The output partition may exist on the same broker or on another broker in the cluster. Within Redpanda, a single Raft controller manages cluster information, including data transforms. On every shard, Redpanda knows what data transforms exist in the cluster, as well as metadata about the transform function, such as input and output topics and environment variables. ![Wasm architecture in Redpanda](https://docs.redpanda.com/cloud-data-platform/shared/_images/wasm_architecture.png) Each transform function reads from a specified input topic and writes to a specified output topic. The transform function processes every record produced to an input topic and returns zero or more records that are then produced to the specified output topic. Data transforms are applied to all partitions on an input topic. A record is processed after it has been successfully written to disk on the input topic. Because the transform happens in the background after the write finishes, the transform doesn’t affect the original produced record, doesn’t block writes to the input topic, and doesn’t block produce and consume requests. A new transform function reads the input topic from the latest offset. That is, it only reads new data produced to the input topic: it does not read records produced to the input topic before the transform was deployed. If a partition leader moves from one broker to another, then the instance of the transform function assigned to that partition moves with it. When a partition replica [loses leadership](https://docs.redpanda.com/cloud-data-platform/get-started/architecture/#partition-leadership-elections), the broker hosting that partition replica stops the instance of the transform function running on the same shard. The broker that is now hosting the partition’s new leader starts the transform function on the same shard as that leader, and the transform function resumes from the last committed offset. If the previous instance of the transform function failed to commit its latest offsets before moving with the partition leader (for example, if the broker crashed), then it’s likely that the new instance will reprocess some events. For broker failures, transform functions have at-least-once semantics, because records are retried from the committed last offset, and offsets are committed periodically. For more information, see [How Data Transforms Work](./). ## [](#limitations)Limitations This section outlines the limitations of data transforms. These constraints are categorized into general limitations affecting the overall functionality and specific limitations related to giving data transforms access to custom environment variables. ### [](#general)General - **No external access**: Transform functions have no external access to disk or network resources. - **Single message transforms**: Only single record transforms are supported, but multiple output records from a single input record are supported. For aggregations, joins, or complex transformations, consider using [Redpanda Connect](https://docs.redpanda.com/connect/get-started/about/) or [Apache Flink](https://flink.apache.org/). - **Output topic limit**: Up to eight output topics are supported. - **Delivery semantics**: Transform functions have at-least-once delivery. - **Transactions API**: When clients use the Kafka Transactions API on partitions of an input topic, transform functions process only committed records. ### [](#javascript)JavaScript - **No native extensions**: Native Node.js extensions are not supported. Packages that require compiling native code or interacting with low-level system features cannot be used. - **Limited Node.js standard modules**: Only modules that can be polyfilled by the [esbuild plugin](https://www.npmjs.com/package/esbuild-plugin-polyfill-node#implemented-polyfills) can be used. Even if a module can be polyfilled, certain functionalities, such as network connections, will not work because the necessary browser APIs are not exposed in the Redpanda JavaScript runtime environment. For example, while the plugin can provide stubs for some Node.js modules such as `http` and `process`, these stubs will not work in the Redpanda JavaScript runtime environment. - **No write options**: The JavaScript SDK does not support write options, such as specifying which output topic to write to. ### [](#environment-variables)Environment variables - **Maximum number of variables**: You can set up to 128 custom environment variables. - **Reserved prefix**: Variable keys must not start with `REDPANDA_`. This prefix is reserved for [built-in environment variables](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/#built-in). - **Key length**: Each key must be less than 128 bytes in length. - **Total value length**: The combined length of all values for the environment variables must be less than 2000 bytes. - **Encoding**: All keys and values must be encoded in UTF-8. - **Control characters**: Keys and values must not contain any control characters, such as null bytes. ## [](#suggested-reading)Suggested reading - [Golang SDK for Data Transforms](https://docs.redpanda.com/cloud-data-platform/reference/data-transforms/golang-sdk/) - [Rust SDK for Data Transforms](https://docs.redpanda.com/cloud-data-platform/reference/data-transforms/rust-sdk/) - [`rpk transform` commands](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-transform/rpk-transform/) --- # Page 319: Monitor Data Transforms **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/monitor.md --- # Monitor Data Transforms > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Monitor Data Transforms latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/monitor page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/monitor.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/monitor.adoc description: This topic provides guidelines on how to monitor the health of your data transforms and view logs. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- This topic provides guidelines on how to monitor the health of your data transforms and view logs. ## [](#prerequisites)Prerequisites [Set up monitoring](https://docs.redpanda.com/cloud-data-platform/manage/monitor-cloud/) for your cluster. ## [](#performance)Performance You can identify performance bottlenecks by monitoring latency and CPU usage: - [`redpanda_transform_execution_latency_sec`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_transform_execution_latency_sec) - [`redpanda_wasm_engine_cpu_seconds_total`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_wasm_engine_cpu_seconds_total) If latency is high, investigate the transform logic for inefficiencies or consider scaling the resources. High CPU usage might indicate the need for optimization in the code or an increase in [allocated CPU resources](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/). ## [](#reliability)Reliability Tracking execution errors and error states helps in maintaining the reliability of your data transforms: - [`redpanda_transform_execution_errors`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_transform_execution_errors) - [`redpanda_transform_failures`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_transform_failures) - [`redpanda_transform_state`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_transform_state) Make sure to [implement robust error handling and logging](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build/#errors) within your transform functions to help with troubleshooting. ## [](#resource-usage)Resource usage Monitoring memory usage metrics and total execution time ensures that the Wasm engine does not exceed allocated resources, helping in efficient resource management: - [`redpanda_wasm_engine_memory_usage`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_wasm_engine_memory_usage) - [`redpanda_wasm_engine_max_memory`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_wasm_engine_max_memory) - [`redpanda_wasm_binary_executable_memory_usage`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_wasm_binary_executable_memory_usage) If memory usage is consistently high or exceeds the maximum allocated memory: - Review and optimize your transform functions to reduce memory consumption. This step can involve optimizing data structures, reducing memory allocations, and ensuring efficient handling of records. ## [](#throughput)Throughput Keeping track of read and write bytes and processor lag helps in understanding the data flow through your transforms, enabling better capacity planning and scaling: - [`redpanda_transform_read_bytes`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_transform_read_bytes) - [`redpanda_transform_write_bytes`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_transform_write_bytes) - [`redpanda_transform_processor_lag`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_transform_processor_lag) If there is a significant lag or low throughput, investigate potential bottlenecks in the data flow or consider scaling your infrastructure to handle higher throughput. ## [](#logs)View logs for data transforms Runtime logs for transform functions are written to an internal topic called `_redpanda.transform_logs`. You can read these logs by using the [`rpk transform logs`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-transform/rpk-transform-logs/) command. ```bash rpk transform logs ``` Replace `` with the [configured name](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/) of the transform function. > 💡 **TIP** > > You can also view logs in the UI. By default, Redpanda provides several settings to manage logging for data transforms, such as buffer capacity, flush interval, and maximum log line length. These settings ensure that logging operates efficiently without overwhelming the system. However, you may need to adjust these settings based on your specific requirements and workloads. For information on how to configure logging, see the [Configure transform logging](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/configure/#log) section of the configuration guide. ## [](#suggested-reading)Suggested reading - [Data transforms metrics](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#data_transform_metrics) --- # Page 320: Write Integration Tests for Transform Functions **URL**: https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/test.md --- # Write Integration Tests for Transform Functions > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Write Integration Tests for Transform Functions latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: data-transforms/test page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: data-transforms/test.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/data-transforms/test.adoc description: Learn how to write integration tests for data transform functions in Redpanda, including setting up unit tests and using testcontainers for integration tests. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- Learn how to write integration tests for data transform functions in Redpanda, including setting up unit tests and using testcontainers for integration tests. This guide covers how to write both unit tests and integration tests for your transform functions. While unit tests focus on testing individual components in isolation, integration tests verify that the components work together as expected in a real environment. ## [](#unit-tests)Unit tests You can create unit tests for transform functions by mocking the interfaces injected into the transform function and asserting that the input and output work correctly. This typically includes mocking the `WriteEvent` and `RecordWriter` interfaces. ```go package main import ( "testing" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/mock" "github.com/redpanda-data/redpanda/src/transform-sdk/go/transform" ) // MockWriteEvent is a mock implementation of the WriteEvent interface. type MockWriteEvent struct { mock.Mock } func (m *MockWriteEvent) Record() transform.Record { args := m.Called() return args.Get(0).(transform.Record) } // MockRecordWriter is a mock implementation of the RecordWriter interface. type MockRecordWriter struct { mock.Mock } func (m *MockRecordWriter) Write(record transform.Record) error { args := m.Called(record) return args.Error(0) } // copyRecord copies the record to the output topic. func copyRecord(event transform.WriteEvent, writer transform.RecordWriter) error { record := event.Record() return writer.Write(record) } // TestCopyRecord tests the copyRecord function. func TestCopyRecord(t *testing.T) { // Create mocks for the WriteEvent and RecordWriter event := new(MockWriteEvent) writer := new(MockRecordWriter) // Set up the expected behavior record := transform.Record{Value: []byte("test")} event.On("Record").Return(record) writer.On("Write", record).Return(nil) // Call the function under test err := copyRecord(event, writer) // Assert that no error occurred and that the expectations were met assert.NoError(t, err) event.AssertExpectations(t) writer.AssertExpectations(t) } ``` To run your unit tests, use the following command: ```bash go test ``` This will execute all tests in the current directory. ## [](#integration-tests)Integration tests Integration tests verify that your transform functions work correctly in a real Redpanda environment. You can use [testcontainers](https://github.com/testcontainers/testcontainers-go/tree/main) to set up and manage a Redpanda instance for testing. For more detailed examples and helper code for setting up integration tests, refer to the SDK integration tests on [GitHub](https://github.com/redpanda-data/redpanda/tree/dev/src/transform-sdk/tests). --- # Page 321: Use Redpanda with the HTTP Proxy API **URL**: https://docs.redpanda.com/cloud-data-platform/develop/http-proxy.md --- # Use Redpanda with the HTTP Proxy API > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use Redpanda with the HTTP Proxy API latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: http-proxy page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: http-proxy.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/http-proxy.adoc description: HTTP Proxy exposes a REST API to list topics, produce events, and subscribe to events from topics using consumer groups. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- Redpanda HTTP Proxy (`pandaproxy`) allows access to your data through a REST API. For example, you can list topics or brokers, get events, produce events, subscribe to events from topics using consumer groups, and commit offsets for a consumer. See the [HTTP Proxy API reference](https://docs.redpanda.com/api/doc/http-proxy/) for a full list of available endpoints. > 📝 **NOTE** > > The HTTP Proxy API is supported for BYOC and Dedicated clusters only. ## [](#prerequisites)Prerequisites ### [](#start-redpanda)Start Redpanda To log in to your Redpanda Cloud account, run `rpk cloud login`. HTTP Proxy is enabled by default on port 30082. For clusters with private connectivity (AWS PrivateLink, GCP Private Service Connect, and Azure Private Link) enabled, the default seed port for HTTP Proxy is 30282. You can find the HTTP Proxy endpoint on the **How to connect** section of the cluster overview in the Cloud UI. > 📝 **NOTE** > > The rest of this guide assumes that the HTTP Proxy port is `30082`. ## [](#authenticate-with-http-proxy)Authenticate with HTTP Proxy HTTP Proxy supports authentication using SCRAM credentials or OIDC tokens. The authentication method depends on the cluster’s [`http_authentication`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#http_authentication) settings. ### [](#scram-authentication)SCRAM Authentication If HTTP Proxy is configured to support SASL, you can provide the SCRAM username and password as part of the Basic Authentication header in your request. For example, to list topics as an authenticated user: #### curl ```bash curl -s -u ":" "http://:30082/topics" ``` #### NodeJS ```javascript let options = { auth: { username: "", password: "" }, }; axios .get("http://:30082/topics", options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` #### Python ```python auth = ("", "") res = requests.get("http://:30082/topics", auth=auth).json() pretty(res) ``` ### [](#oidc-authentication)OIDC Authentication If HTTP Proxy is configured to support OIDC, you can provide an OIDC token in the Authorization header. For example: #### curl ```bash curl -s -H "Authorization: Bearer " "http://:30082/topics" ``` #### NodeJS ```javascript let options = { headers: { Authorization: `Bearer ` }, }; axios .get("http://:30082/topics", options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` #### Python ```python headers = {"Authorization": "Bearer "} res = requests.get("http://:30082/topics", headers=headers).json() pretty(res) ``` ## [](#set-up-libraries)Set up libraries You need an app that calls the HTTP Proxy endpoint. This app can be curl (or a similar CLI), or it could be your own custom app written in any language. Below are curl, JavaScript and Python examples. > 📝 **NOTE** > > In the examples, `` refers to your Redpanda cluster’s hostname or IP address. All following examples use a `base_uri` variable that combines the protocol, host, and port for consistency across curl, JavaScript, and Python examples. ### curl Curl is likely already installed on your system. If not, see [curl download instructions](https://curl.se/download.html). Set the base URI for your HTTP Proxy: ```bash base_uri="http://:30082" ``` ### NodeJS > 📝 **NOTE** > > This is based on the assumption that you’re in the root directory of an existing NodeJS project. See [Build a Chat Room Application with Redpanda and Node.js](https://docs.redpanda.com/labs/clients/docker-nodejs/) for an example of a NodeJS project. In a terminal window, run: ```bash npm install axios ``` Import the library into your code: ```javascript const axios = require('axios'); const base_uri = 'http://:30082'; ``` ### Python In a terminal window, run: ```bash pip install requests ``` Import the library into your code: ```python import requests import json def pretty(text): print(json.dumps(text, indent=2)) base_uri = "http://:30082" ``` ## [](#create-a-topic)Create a topic To create a test topic for this guide, use [`rpk`](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). You can configure `rpk` for your Redpanda deployment, using [profiles](https://docs.redpanda.com/cloud-data-platform/manage/rpk/config-rpk-profile/), flags, or [environment variables](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-x-options/#environment-variables). To create a topic named `test_topic` with three partitions, run: ```bash rpk topic create test_topic -p 3 ``` For more information, see the [rpk topic create](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-topic/rpk-topic-create/) reference. ## [](#access-your-data)Access your data Here are some sample commands to produce and consume streams: ### [](#get-list-of-topics)Get list of topics #### curl ```bash curl -s "$base_uri/topics" ``` #### NodeJS ```javascript axios .get(`${base_uri}/topics`) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` Run the application. If your file name is `index.js` for example, you would run the following command: ```bash node index.js ``` #### Python ```python res = requests.get(f"{base_uri}/topics").json() pretty(res) ``` Expected output: ```bash ["test_topic"] ``` ### [](#send-events-to-a-topic)Send events to a topic Use POST to send events in the REST endpoint query. The header must include the following line: Content-Type:application/vnd.kafka.json.v2+json The following commands show how to send events to `test_topic`: #### curl ```bash curl -s \ -X POST \ "$base_uri/topics/test_topic" \ -H "Content-Type: application/vnd.kafka.json.v2+json" \ -d '{ "records":[ { "value":"Redpanda", "partition":0 }, { "value":"HTTP proxy", "partition":1 }, { "value":"Test event", "partition":2 } ] }' ``` #### NodeJS ```javascript let payload = { records: [ { "value":"Redpanda", "partition": 0 }, { "value":"HTTP proxy", "partition": 1 }, { "value":"Test event", "partition": 2 } ]}; let options = { headers: { "Content-Type" : "application/vnd.kafka.json.v2+json" }}; axios .post(`${base_uri}/topics/test_topic`, payload, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` Run the application: ```bash node index.js ``` #### Python ```python res = requests.post( url=f"{base_uri}/topics/test_topic", data=json.dumps( dict(records=[ dict(value="Redpanda", partition=0), dict(value="HTTP Proxy", partition=1), dict(value="Test Event", partition=2) ])), headers={"Content-Type": "application/vnd.kafka.json.v2+json"}).json() pretty(res) ``` Expected output (may be formatted differently depending on the chosen application): ```bash {"offsets":[{"partition":0,"offset":0},{"partition":2,"offset":0},{"partition":1,"offset":0}]} ``` ### [](#get-events-from-a-topic)Get events from a topic After events have been sent to the topic, you can retrieve these same events. #### curl ```bash curl -s \ "$base_uri/topics/test_topic/partitions/0/records?offset=0&timeout=1000&max_bytes=100000"\ -H "Accept: application/vnd.kafka.json.v2+json" ``` #### NodeJS ```javascript let options = { headers: { accept: "application/vnd.kafka.json.v2+json" }, params: { offset: 0, timeout: "1000", max_bytes: "100000", }, }; axios .get(`${base_uri}/topics/test_topic/partitions/0/records`, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` Run the application: ```bash node index.js ``` #### Python ```python res = requests.get( url=f"{base_uri}/topics/test_topic/partitions/0/records", params={"offset": 0, "timeout":1000,"max_bytes":100000}, headers={"Accept": "application/vnd.kafka.json.v2+json"}).json() pretty(res) ``` Expected output: ```bash [{"topic":"test_topic","key":null,"value":"Redpanda","partition":0,"offset":0}] ``` ### [](#get-list-of-brokers)Get list of brokers #### curl ```bash curl "$base_uri/brokers" ``` #### NodeJS ```javascript axios .get(`${base_uri}/brokers`) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` #### Python ```python res = requests.get(f"{base_uri}/brokers").json() pretty(res) ``` Expected output: ```bash {brokers: [0]} ``` ### [](#create-a-consumer)Create a consumer To retrieve events from a topic using consumers, you must create a consumer and a consumer group, and then subscribe the consumer instance to a topic. Each action involves a different endpoint and method. The first endpoint is: `/consumers/`. For this REST call, the payload is the group information. #### curl ```bash curl -s \ -X POST \ "$base_uri/consumers/test_group" \ -H "Content-Type: application/vnd.kafka.v2+json" \ -d '{ "format":"json", "name":"test_consumer", "auto.offset.reset":"earliest", "auto.commit.enable":"false", "fetch.min.bytes": "1", "consumer.request.timeout.ms": "10000" }' ``` #### NodeJS ```javascript let payload = { "name": "test_consumer", "format": "json", "auto.offset.reset": "earliest", "auto.commit.enable": "false", "fetch.min.bytes": "1", "consumer.request.timeout.ms": "10000" }; let options = { headers: { "Content-Type": "application/vnd.kafka.v2+json" }}; axios .post(`${base_uri}/consumers/test_group`, payload, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` Run the application: ```bash node index.js ``` #### Python ```python res = requests.post( url=f"{base_uri}/consumers/test_group", data=json.dumps({ "name": "test_consumer", "format": "json", "auto.offset.reset": "earliest", "auto.commit.enable": "false", "fetch.min.bytes": "1", "consumer.request.timeout.ms": "10000" }), headers={"Content-Type": "application/vnd.kafka.v2+json"}).json() pretty(res) ``` Expected output: ```bash {"instance_id":"test_consumer","base_uri":"http://:30082/consumers/test_group/instances/test_consumer"} ``` > 📝 **NOTE** > > - Consumers expire after five minutes of inactivity. To prevent this from happening, try consuming events within a loop. If the consumer has expired, you can create a new one with the same name. > > - The output `base_uri` is the full URL path for this specific consumer instance and differs from the `base_uri` variable used in the code examples. ### [](#subscribe-to-the-topic)Subscribe to the topic After creating the consumer, subscribe to the topic that you created. #### curl ```bash curl -s -o /dev/null -w "%{http_code}" \ -X POST \ "$base_uri/consumers/test_group/instances/test_consumer/subscription"\ -H "Content-Type: application/vnd.kafka.v2+json" \ -d '{ "topics": [ "test_topic" ] }' ``` #### NodeJS ```javascript let payload = { topics: ["test_topic"]}; let options = { headers: { "Content-Type": "application/vnd.kafka.v2+json" }}; axios .post(`${base_uri}/consumers/test_group/instances/test_consumer/subscription`, payload, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` Run the application: ```bash node index.js ``` #### Python ```python res = requests.post( url=f"{base_uri}/consumers/test_group/instances/test_consumer/subscription", data=json.dumps({"topics": ["test_topic"]}), headers={"Content-Type": "application/vnd.kafka.v2+json"}) ``` Expected response is an HTTP 204, without a body. Now you can get the events from `test_topic`. ### [](#retrieve-events)Retrieve events Retrieve the events from the topic: #### curl ```bash curl -s \ "$base_uri/consumers/test_group/instances/test_consumer/records?timeout=1000&max_bytes=100000"\ -H "Accept: application/vnd.kafka.json.v2+json" ``` #### NodeJS ```javascript let options = { headers: { Accept: "application/vnd.kafka.json.v2+json" }, params: { timeout: "1000", max_bytes: "100000", }, }; axios .get(`${base_uri}/consumers/test_group/instances/test_consumer/records`, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` Run the application: ```bash node index.js ``` #### Python ```python res = requests.get( url=f"{base_uri}/consumers/test_group/instances/test_consumer/records", params={"timeout":1000,"max_bytes":100000}, headers={"Accept": "application/vnd.kafka.json.v2+json"}).json() pretty(res) ``` Expected output: ```bash [{"topic":"test_topic","key":null,"value":"Redpanda","partition":0,"offset":0},{"topic":"test_topic","key":null,"value":"HTTP proxy","partition":1,"offset":0},{"topic":"test_topic","key":null,"value":"Test event","partition":2,"offset":0}] ``` ### [](#get-offsets-from-consumer)Get offsets from consumer #### curl ```bash curl -s \ -X 'GET' \ curl -s -o /dev/null -w "%{http_code}" \ -X 'POST' \ "$base_uri/consumers/test_group/instances/test_consumer/offsets" \ -H 'accept: application/vnd.kafka.v2+json' \ -H 'accept: application/vnd.kafka.v2+json' \ -H 'Content-Type: application/vnd.kafka.v2+json' \ -d '{ "partitions": [ { "topic": "test_topic", "partition": 0 }, { "topic": "test_topic", "partition": 1 }, { "topic": "test_topic", "partition": 2 } ] }' ``` #### Python ```python res = requests.get( url=f"{base_uri}/consumers/test_group/instances/test_consumer/offsets", data=json.dumps( dict(partitions=[ dict(topic="test_topic", partition=p) for p in [0, 1, 2] ])), headers={"Content-Type": "application/vnd.kafka.v2+json"}).json() pretty(res) ``` Expected output: ```bash { "offsets": [{ "topic": "test_topic", "partition": 0, "offset": 0, "metadata": "" },{ "topic": "test_topic", "partition": 1, "offset": 0, "metadata": "" }, { "topic": "test_topic", "partition": 2, "offset": 0, "metadata": "" }] } ``` ### [](#commit-offsets-for-consumer)Commit offsets for consumer After events have been handled by a consumer, the offsets can be committed, so that the consumer group won’t retrieve them again. #### curl ```bash curl -s -o /dev/null -w "%{http_code}" \ -X 'POST' \ "$base_uri/consumers/test_group/instances/test_consumer/offsets" \ -H 'accept: application/vnd.kafka.v2+json' \ -H 'Content-Type: application/vnd.kafka.v2+json' \ -d '{ "partitions": [ { "topic": "test_topic", "partition": 0, "offset": 0 }, { "topic": "test_topic", "partition": 1, "offset": 0 }, { "topic": "test_topic", "partition": 2, "offset": 0 } ] }' ``` #### NodeJS ```javascript let options = { headers: { accept: "application/vnd.kafka.v2+json", "Content-Type": "application/vnd.kafka.v2+json", } }; let payload = { partitions: [ { topic: "test_topic", partition: 0, offset: 0 }, { topic: "test_topic", partition: 1, offset: 0 }, { topic: "test_topic", partition: 2, offset: 0 }, ]}; axios .post(`${base_uri}/consumers/test_group/instances/test_consumer/offsets`, payload, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` Run the application: ```bash node index.js ``` #### Python ```python res = requests.post( url=f"{base_uri}/consumers/test_group/instances/test_consumer/offsets", data=json.dumps( dict(partitions=[ dict(topic="test_topic", partition=p, offset=0) for p in [0, 1, 2] ])), headers={"Content-Type": "application/vnd.kafka.v2+json"}) ``` Expected output: none. ### [](#delete-a-consumer)Delete a consumer To remove a consumer from a group, send a DELETE request as shown below: #### curl ```bash curl -s -o /dev/null -w "%{http_code}" \ -X 'DELETE' \ "$base_uri/consumers/test_group/instances/test_consumer" \ -H 'Content-Type: application/vnd.kafka.v2+json' ``` #### NodeJS ```javascript let options = { headers: { "Content-Type": "application/vnd.kafka.v2+json" }}; axios .delete(`${base_uri}/consumers/test_group/instances/test_consumer`, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` #### Python ```python res = requests.delete( url=f"{base_uri}/consumers/test_group/instances/test_consumer", headers={"Content-Type": "application/vnd.kafka.v2+json"}) ``` ## [](#authenticate-with-http-proxy-2)Authenticate with HTTP Proxy HTTP Proxy supports authentication using SCRAM credentials or OIDC tokens. The authentication method depends on the cluster’s [`http_authentication`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#http_authentication) settings. ### [](#scram-authentication-2)SCRAM Authentication If HTTP Proxy is configured to support SASL, you can provide the SCRAM username and password as part of the Basic Authentication header in your request. For example, to list topics as an authenticated user: #### curl ```bash curl -s -u ":" ":8082/topics" ``` #### NodeJS ```javascript let options = { auth: { username: "", password: "" }, }; axios .get(`${base_uri}/topics`, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` #### Python ```python auth = ("", "") res = requests.get(f"{base_uri}/topics", auth=auth).json() pretty(res) ``` ### [](#oidc-authentication-2)OIDC Authentication If HTTP Proxy is configured to support OIDC, you can provide an OIDC token in the Authorization header. For example: #### curl ```bash curl -s -H "Authorization: Bearer " ":8082/topics" ``` #### NodeJS ```javascript let options = { headers: { Authorization: `Bearer ` }, }; axios .get(`${base_uri}/topics`, options) .then(response => console.log(response.data)) .catch(error => console.error(error)); ``` #### Python ```python headers = {"Authorization": "Bearer "} res = requests.get(f"{base_uri}/topics", headers=headers).json() pretty(res) ``` ## [](#use-swagger-with-http-proxy)Use Swagger with HTTP Proxy You can use Swagger UI to test and interact with Redpanda HTTP Proxy endpoints. Use Docker to start Swagger UI: ```bash docker run -p 80:8080 -d swaggerapi/swagger-ui ``` Verify that the Swagger container is available: ```bash docker ps ``` Verify that the Docker container has been added and is running: `swaggerapi/swagger-ui` with `Up…` status In a browser, enter `` in the address bar to open the Swagger console. Change the URL to `[http://:30082/v1](http://\:30082/v1)`, and click `Explore` to update the page with Redpanda HTTP Proxy endpoints. You can call the endpoints in any application and language that supports web interactions. --- # Page 322: Kafka Compatibility **URL**: https://docs.redpanda.com/cloud-data-platform/develop/kafka-clients.md --- # Kafka Compatibility > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Kafka Compatibility latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: kafka-clients page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: kafka-clients.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/kafka-clients.adoc description: Kafka clients, version 0.11 or later, are compatible with Redpanda. Validations and exceptions are listed. page-topic-type: reference personas: developer learning-objective-1: Identify which Kafka clients are validated with Redpanda learning-objective-2: Identify unsupported Kafka features when integrating with Redpanda page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-29" --- Apache Kafka® clients developed for Kafka protocol version 0.11 or later work with Redpanda with minimal or no changes to your application. This page identifies which clients are validated and calls out any exceptions. Use this reference to: - Identify which Kafka clients are validated with Redpanda - Identify unsupported Kafka features when integrating with Redpanda ## [](#kafka-client-compatibility)Kafka client compatibility Redpanda validates the Apache Kafka Java client and a set of widely used non-Java clients, at their current versions that support Kafka 4.x, using the ducktape and chaos test suites. Validation confirms connectivity and correctness across core Kafka APIs, such as produce, consume, and transaction operations, at current client versions. Modern clients auto-negotiate protocol versions or use an earlier protocol version accepted by Redpanda brokers. > 💡 **TIP** > > Always use the latest supported version of a Kafka client. The following clients have been validated with Redpanda. | Language | Client | | --- | --- | | Java | Apache Kafka Java Client | | C/C++ | librdkafka | | Go | franz-goconfluent-kafka-goSarama | | Python | kafka-pythonconfluent-kafka-python | | Rust | kafka-rust | | Node.js | KafkaJSconfluent-kafka-javascript | Clients that have not been validated by Redpanda Data, but use the Kafka protocol, remain compatible with Redpanda subject to the limitations in the next section (particularly those based on librdkafka, such as confluent-kafka-dotnet). If you find a client that does not work with Redpanda, reach out in the [Redpanda community Slack](https://redpanda.com/slack). ## [](#compatibility-exceptions)Compatibility exceptions Redpanda is compatible with the Kafka protocol, with the following exceptions: - Multiple SCRAM mechanisms simultaneously for SASL users are not supported. For example, a user cannot have both a `SCRAM-SHA-256` and a `SCRAM-SHA-512` credential. Redpanda supports only one SASL/SCRAM mechanism per user: either `SCRAM-SHA-256` or `SCRAM-SHA-512`. For details, see [Authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/). - HTTP Proxy (`pandaproxy`): Unlike other REST proxy implementations in the Kafka ecosystem, Redpanda HTTP Proxy does not support topic and ACLs CRUD through the HTTP Proxy. HTTP Proxy is designed for clients producing and consuming data that do not perform administrative functions. - The `delete.retention.ms` topic configuration in Kafka is not supported for Tiered Storage topics. Cloud Topics and local storage topics support Tombstone marker deletion using `delete.retention.ms`, but in Tiered Storage topics, Tombstone markers are only removed in accordance with normal topic retention, and only if the cleanup policy is `delete` or `compact, delete`. - The Kafka request rate quota (`request_percentage`), which limits the share of broker request-handling capacity a client can consume, is not supported. Redpanda supports byte-rate (`producer_byte_rate`, `consumer_byte_rate`) and topic-mutation (`controller_mutation_rate`) quotas, which you can apply [per user, per client, or per client group](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/#client-throughput-limits). - [KIP-890](https://cwiki.apache.org/confluence/display/KAFKA/KIP-890) (Transactions Server-Side Defense): Redpanda does not implement the server-side portion of KIP-890, which addresses transaction errors specific to Kafka’s replication model. Redpanda’s implementation of transactions is not susceptible to this class of errors. When connecting to Redpanda, Kafka 4.x clients detect that Transactions V2 is unsupported and fall back to the original transaction protocol (per-transaction epoch bumping is part of V2 and does not apply). If you find an unsupported feature or incompatibility, [file an issue](https://github.com/redpanda-data/redpanda/issues/new) with the Redpanda team. --- # Page 323: Kafka Connect **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors.md --- # Kafka Connect > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Kafka Connect latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/index.adoc description: Use Kafka Connect to stream data into and out of Redpanda. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-07" --- Use Kafka Connect to integrate your Redpanda data with different data systems. As managed solutions, connectors offer a simpler way to integrate your data than manually creating a solution with the Kafka API. You can set up and manage these connectors for BYOC and Dedicated clusters in the Redpanda Cloud UI or Cloud API. > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. Each connector is either a source or a sink: - A source connector imports data from a source system into a Redpanda cluster. The source connector’s main task is to fetch data from these sources and convert them into a format suitable for Redpanda. - A sink connector exports data from a Redpanda cluster and pushes it into a target system. Sink connectors read the data from Redpanda and transform it into a format that the target system can use. These sources and sinks work together to create a data pipeline that can move and transform data from one system to another. > ⚠️ **WARNING** > > Modifying the properties of topics that are created and managed by Redpanda applications can cause unexpected errors. This may lead to connector and cluster failures. - [Converters and Serialization](converters-and-serialization/) Use converters to handle the serialization and deserialization of data between a Redpanda topic and an external system with Kafka Connect. - [Monitor Kafka Connect](monitor-connectors/) Use metrics to monitor the health of Kafka Connect. - [Disable Kafka Connect](disable-kc/) Learn how to disable Kafka Connect using the Cloud API. - [Single Message Transforms](transforms/) Single Message Transforms (SMTs) let you modify the data and its characteristics as it passes through a connector. - [Sizing Connectors](sizing-connectors/) How to choose number of tasks to set for a connector. - [Create an S3 Sink Connector](create-s3-sink-connector/) Use the Redpanda Cloud UI to create an AWS S3 Sink Connector. - [Create a Google BigQuery Sink Connector](create-gcp-bigquery-connector/) Use the Redpanda Cloud UI to create a Google BigQuery Sink Connector. - [Create a GCS Sink Connector](create-gcs-connector/) Use the Redpanda Cloud UI to create a GCS Sink Connector. - [Create an Iceberg Sink Connector](create-iceberg-sink-connector/) Use the Redpanda Cloud UI to create an Iceberg Sink Connector. - [Create a JDBC Sink Connector](create-jdbc-sink-connector/) Use the Redpanda Cloud UI to create a JDBC Sink Connector. - [Create a JDBC Source Connector](create-jdbc-source-connector/) Use the Redpanda Cloud UI to create a JDBC Source Connector. - [Create a MirrorMaker2 Source Connector](create-mmaker-source-connector/) Use the Redpanda Cloud UI to create a MirrorMaker2 Source Connector. - [Create a MirrorMaker2 Checkpoint Connector](create-mmaker-checkpoint-connector/) Use the Redpanda Cloud UI to create a MirrorMaker2 Checkpoint Connector. - [Create a MirrorMaker2 Heartbeat Connector](create-mmaker-heartbeat-connector/) Use the Redpanda Cloud UI to create a MirrorMaker2 Heartbeat Connector. - [Create a MongoDB Sink Connector](create-mongodb-sink-connector/) Use the Redpanda Cloud UI to create a MongoDB Sink Connector. - [Create a MongoDB Source Connector](create-mongodb-source-connector/) Use the Redpanda Cloud UI to create a MongoDB Source Connector. - [Create a MySQL (Debezium) Source Connector](create-mysql-source-connector/) Use the Redpanda Cloud UI to create a MySQL (Debezium) Source Connector. - [Create a PostgreSQL (Debezium) Source Connector](create-postgresql-connector/) Use the Redpanda Cloud UI to create a PostgreSQL (Debezium) Source Connector. - [Create a SQL Server (Debezium) Source Connector](create-sqlserver-connector/) Use the Redpanda Cloud UI to create a SQL Server (Debezium) Source Connector. - [Create a Snowflake Sink Connector](create-snowflake-connector/) Use the Redpanda Cloud UI to create a Snowflake Sink Connector. --- # Page 324: Converters and Serialization **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/converters-and-serialization.md --- # Converters and Serialization > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Converters and Serialization latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/converters-and-serialization page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/converters-and-serialization.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/converters-and-serialization.adoc description: Use converters to handle the serialization and deserialization of data between a Redpanda topic and an external system with Kafka Connect. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-09-26" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. Connectors are a translation layer working between Redpanda and the remote system. For **sink** connectors the translation happens in the following phases: 1. Converter deserializes data from Redpanda message format (for example JSON or Avro) to a universal in-memory connect data format. 2. The in-memory connect data structure is translated by the connector to the data model of the remote system. For **source** connectors it is vice versa, the phases are: 1. Connector translates the data model from remote system format to the in-memory connect data structure. 2. Converter serializes the data from a universal in-memory connect format to a Redpanda message. Each Redpanda message is a key and value record. Record key and value converters are configured separately with the `Redpanda message key format` and `Redpanda message value format` properties. Key and value converters can be different. > 📝 **NOTE** > > If an external system requires structured data (like BigQuery or a SQL database), then you must provide data with a schema. Use the Avro, Protobuf, or JSON converter with a schema. ## [](#bytearray-converter)ByteArray converter The ByteArray converter is the most primitive and high-throughput converter. Schema is ignored. This is the default converter type for managed connectors. To use the converter, select the `ByteArray` option as a key or value message format. ## [](#string-converter)String converter The String converter is a high-throughput converter. Schema is ignored. All data is converted to a string. To use the converter, select the `String` option as a key or value message format. ## [](#json-converter)JSON converter The JSON converter supports a JSON schema embedded in the message, where each message contains a schema. It results in a bigger message size. The connector needs a message schema to check message format. To use the converter, select the `JSON` option as a key or value message format. Example JSON message with embedded schema: ```json { "schema": { "type": "struct", "fields": [ { "type": "int64", "optional": false, "field": "person_id" }, { "type": "string", "optional": false, "field": "name" } ] }, "payload": { "person_id": 1, "name": "Redpanda" } } ``` If you consume JSON data with no message schema, the schema check for the connector must be disabled with the `Message key JSON contains schema` or `Message value JSON contains schema` option. ## [](#avro-converter)Avro converter The Avro converter requires a schema in Schema Registry. Avro supports primitive types and complex types, like records, enums, arrays, maps, and unions. To specify a timestamp in an Avro schema for use with Kafka Connect, use: ```json { "name": "time1", "type": [ "null", { "type": "long", "connect.version": 1, "connect.name": "org.apache.kafka.connect.data.Timestamp", "logicalType": "timestamp-millis" } ], "default": null } ``` See also: - [Redpanda Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/) - [Avro specification](https://avro.apache.org/docs/1.11.1/specification) ## [](#cloudevents-converter)CloudEvents converter The CloudEvents converter is specific to Debezium PostgreSQL and MySQL source connectors. See also: [CloudEvents Converter documentation](https://debezium.io/documentation/reference/2.2/integrations/cloudevents.html) ## [](#protobuf-converter)Protobuf converter ![Beta](https://img.shields.io/badge/Beta-red.svg) The Protobuf converter requires a schema in Schema Registry. The converter only supports sink connectors. Source connectors are not supported. To use the converter, select the `Protobuf` option as a key or value message format. See also: [Redpanda Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/) ## [](#set-property-keys)Set property keys Kafka Connect connectors use a set of `=` to set up properties. For example if you want to set the property `topic.creation.enable` to `true`, use `topic.creation.enable=true` in the property settings page. --- # Page 325: Create a Google BigQuery Sink Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-gcp-bigquery-connector.md --- # Create a Google BigQuery Sink Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a Google BigQuery Sink Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-gcp-bigquery-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-gcp-bigquery-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-gcp-bigquery-connector.adoc description: Use the Redpanda Cloud UI to create a Google BigQuery Sink Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. The Google BigQuery Sink connector enables you to stream any structured data from Redpanda to BigQuery for advanced analytics. ## [](#prerequisites)Prerequisites Before you can create a Google BigQuery Sink connector in the Redpanda Cloud, you must: 1. Create a [Google Cloud](https://cloud.google.com/) account. 2. In the **Google home** page: 1. [Select an existing project](https://cloud.google.com/resource-manager/docs/creating-managing-projects#get_an_existing_project) or [create a new one](https://cloud.google.com/resource-manager/docs/creating-managing-projects#creating_a_project). 2. [Create a new dataset](https://cloud.google.com/bigquery/docs/datasets) for the project. 3. (_Optional if your data has a schema_) After creating the dataset, [create a new table](https://cloud.google.com/bigquery/docs/tables) to hold the data you intend to stream from Redpanda Cloud topics. Specify a structure for the table using schema values that align with your Redpanda topic data. > 📝 **NOTE** > > This step is mandatory only if the data in Redpanda does not have a schema. If the data in Redpanda includes a schema, then the connector automatically creates the tables in BigQuery. 3. Create a [custom role](https://cloud.google.com/iam/docs/creating-custom-roles). The role must have the following permissions: bigquery.datasets.get bigquery.tables.create bigquery.tables.get bigquery.tables.getData bigquery.tables.list bigquery.tables.update bigquery.tables.updateData 4. Create a [service account](https://cloud.google.com/iam/docs/service-accounts-create). 5. [Add the custom role to your service account](https://cloud.google.com/iam/docs/granting-changing-revoking-access). 6. [Create a service account key](https://cloud.google.com/iam/docs/keys-create-delete), and then download it. ## [](#limitations)Limitations The Google BigQuery Sink connector doesn’t support schemas with recursion. ## [](#create-a-google-bigquery-sink-connector)Create a Google BigQuery Sink connector To create the Google BigQuery Sink connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Export to Google BigQuery**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to export | topics | A comma-separated list of the cluster topics you want to replicate to Google BigQuery. | | Topics regex | topics.regex | A Java regular expression of topics to replicate. For example: specify .* to replicate all available topics in the cluster. Applicable only when Use regular expressions is selected. | | Credentials JSON | keyfile | A JSON key with BigQuery service account credentials. | | Project | project | The BigQuery project to which topic data will be written. | | Default dataset | defaultDataset | The default Google BigQuery dataset to be used. | | Kafka message value format | value.converter | The format of the value in the Redpanda topic. The default is JSON. | | Max Tasks | tasks.max | Maximum number of tasks to use for this connector. The default is 1. Each task replicates exclusive set of partitions assigned to it. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-google-bigquery-sink-connector-configuration)Advanced Google BigQuery Sink connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require any additional property settings (for example, automatically create BigQuery tables or map topics to tables), then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Auto create tables | autoCreateTables | Automatically create BigQuery tables if they don’t already exist. If the table does not exist, then it is created based on the record schema. | | Topic to table map | topic2TableMap | Map of topics to tables. Format: comma-separated tuples, for example topic1:table1,topic2:table2. | | Allow new BigQuery fields | allowNewBigQueryFields | If true, new fields can be added to BigQuery tables during subsequent schema updates. | | Allow BigQuery required field relaxation | allowBigQueryRequiredFieldRelaxation | If true, fields in the BigQuery schema can be changed from REQUIRED to NULLABLE. | | Upsert enabled | upsertEnabled | Enables upsert functionality on the connector. | | Delete enabled | deleteEnabled | Enable delete functionality on the connector. | | Kafka key field name | kafkaKeyFieldName | The name of the BigQuery table field for the Kafka key. Must be set when upsert or delete is enabled. | | Time partitioning type | timePartitioningType | The time partitioning type to use when creating tables. | | BigQuery retry attempts | bigQueryRetry | The number of retry attempts made for each BigQuery request that fails with a backend or quota exceeded error. | | BigQuery retry attempts interval | bigQueryRetryWait | The minimum amount of time, in milliseconds, to wait between BigQuery backend or quota exceeded error retry attempts. | | Error tolerance | errors.tolerance | Error tolerance response during connector operation. Default value is none and signals that any error will result in an immediate connector task failure. Value of all changes the behavior to skip over problematic records. | | Dead letter queue topic name | errors.deadletterqueue.topic.name | The name of the topic to be used as the dead letter queue (DLQ) for messages that result in an error when processed by this sink connector, its transformations, or converters. The topic name is blank by default, which means that no messages are recorded in the DLQ. | | Dead letter queue topic replication factor | errors.deadletterqueue.topic .replication.factor | Replication factor used to create the dead letter queue topic when it doesn’t already exist. | | Enable error context headers | errors.deadletterqueue.context .headers.enable | When true, adds a header containing error context to the messages written to the dead letter queue. To avoid clashing with headers from the original record, all error context header keys, start with __connect.errors. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - `JSON` (`org.apache.kafka.connect.json.JsonConverter`) when your messages are JSON-encoded. Select `Message JSON contains schema`, with the `schema` and `payload` fields. If your messages do not contain schema, manually create tables in BigQuery. - `AVRO` (`io.confluent.connect.avro.AvroConverter`) when your messages contain AVRO-encoded messages, with schema stored in the Schema Registry. ## [](#topic-name-to-table-name-mapping)Topic name to table name mapping By default, the table name is the name of the topic. Use the `Topic to table map` (`topic2TableMap`) configuration property to remap topic names. For example, `topic1:table1,topic2:table2`. ## [](#test-the-connection)Test the connection After the connector is created, go to your BigQuery worksheets and query your table: ```sql SELECT * FROM `project.dataset.table` ``` It may take a couple of minutes for the records to be visible in BigQuery. ## [](#troubleshoot)Troubleshoot Google credentials are checked for validity during connector creation, upon clicking **Finish**. In cases where there are invalid credentials, the connector is not created. Other issues are reported using a failed task error message. Select **Show Logs** to view error details. | Message | Action | | --- | --- | | Not found: Project invalid-project-name | Check to make sure Project contains a valid BigQuery project. | | Not found: Dataset project:invalid-dataset | Check to make sure Default dataset contains a valid BigQuery dataset. | | An unexpected error occurred while validating credentials for BigQuery: Failed to create credentials from input stream | The credentials given as a JSON file in the Credentials JSON property are incorrect. Copy a valid key from the Google Cloud service account. | | JsonConverter with schemas.enable requires "schema" and "payload" fields | The connector encountered an incorrect message format when reading from a topic. | | JsonParseException: Unrecognized token 'test': was expecting JSON | During reading from a topic the connector encountered a message that is invalid JSON. | | Streaming to metadata partition of column-based partitioning table {table_name} is disallowed. | Check to confirm that the bigQueryPartitionDecorator property is set to false. You can check the property in the connector configuration JSON view. | | Caused by: table: GenericData{classInfo=…​ insertion failed for the following rows:…​ no such field: | The Redpanda message contains a property that does not exist in a BigQuery table schema. | | BigQueryConnectException …​ insertion failed for the following rows: …​ [row index 0] (location fieldname[0], reason: invalid): This field: fieldname is not a record. | The Redpanda message contains an array of records, but the BigQuery table expects an array of strings. | | BigQueryConnectException: Failed to unionize schemas of records for the table…​ Could not convert to BigQuery schema with a batch of tombstone records. | The Redpanda message does not contain a schema, so the connector cannot create a BigQuery table. Create the BigQuery table manually. | --- # Page 326: Create a GCS Sink Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-gcs-connector.md --- # Create a GCS Sink Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a GCS Sink Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-gcs-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-gcs-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-gcs-connector.adoc description: Use the Redpanda Cloud UI to create a GCS Sink Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. The Google Cloud Storage (GCS) Sink connector stores Redpanda messages in a Google Cloud Storage bucket. ## [](#prerequisites)Prerequisites Before you can create a GCS Sink connector in the Redpanda Cloud, you must: 1. Create a [Google Cloud](https://cloud.google.com/) account. 2. [Create a service account](https://cloud.google.com/iam/docs/service-accounts-create) that will be used to connect to the GCS service. 3. [Create a service account key](https://cloud.google.com/iam/docs/keys-create-delete) and download it. 4. Create a [custom role](https://cloud.google.com/iam/docs/creating-custom-roles), which must have the following permissions: - `storage.objects.create` to create items in the GCS bucket - `storage.objects.delete` to overwrite items in the GCS bucket 5. [Create a GCS bucket](https://cloud.google.com/storage/docs/creating-buckets) to which to send data. 6. [Grant permissions](https://cloud.google.com/storage/docs/access-control/using-iam-permissions) to the bucket your created for your service account. Use the role created in step 4. ## [](#limitations)Limitations The GCS Sink connector has the following limitations: - You can use only the `STRING` and `BYTES` input formats for `CSV` output format. - You can use only the `PARQUET` format when your messages contain schema. ## [](#create-a-gcs-sink-connector)Create a GCS Sink connector To create the GCS Sink connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Export to Google Cloud Storage**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to export | topics | Comma-separated list of the cluster topics you want to replicate to GCS. | | Topics regex | topics.regex | Java regular expression of topics to replicate. For example: specify .* to replicate all available topics in the cluster. Applicable only when Use regular expressions is selected. | | GCS Credentials JSON | gcs.credentials.json | JSON object with GCS credentials. | | GCS bucket name | gcs.bucket.name | Name of an existing GCS bucket to store output files in. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. Use BYTES for no conversion. | | Kafka message value format | value.converter | Format of the value in the Redpanda topic. Use BYTES for no conversion. | | GCS file format | format.output.type | Format of the files created in GCS: CSV (the default), JSON, JSONL AVRO, or PARQUET. You can use the CSV format output only with BYTES and STRING. | | Avro codec | avro.codec | The Avro compression codec to be used for Avro output files. Available values: null (the default), deflate, snappy, and bzip2. | | Max Tasks | tasks.max | Maximum number of tasks to use for this connector. The default is 1. Each task replicates exclusive set of partitions assigned to it. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-gcs-sink-connector-configuration)Advanced GCS Sink connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require any additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | File name template | file.name.template | The template for file names on GCS. Supports {{ variable }} placeholders for substituting variables. Supported placeholders are:topicpartitionstart_offset (the offset of the first record in the file)timestamp:unit=yyyy|MM|dd|HH (the timestamp of the record)key (when used, other placeholders are not substituted) | | File name prefix | file.name.prefix | The prefix to be added to the name of each file put in GCS. | | Output fields | format.output.fields | Fields to place into output files. Supported values are: 'key', 'value', 'offset', 'timestamp', and 'headers'. | | Value field encoding | format.output.fields.value.encoding | The type of encoding to be used for the value field. Supported values are: 'none' and 'base64'. | | Envelope for primitives | format.output.envelope | Specifies whether or not to enable additional JSON object wrapping of the actual value. | | Output file compression | file.compression.type | The compression type to be used for files put into GCS. Supported values are: 'none', 'gzip', 'snappy', and 'zstd'. | | Max records per file | file.max.records | The maximum number of records to put in a single file. Must be a non-negative number. 0 is interpreted as "unlimited", which is the default. In this case files are only flushed after file.flush.interval.ms. | | File flush interval milliseconds | file.flush.interval.ms | The time interval to periodically flush files and commit offsets. Value specified must be a non-negative number. Default is 60 seconds. 0 indicates that it is disabled. In this case, files are only flushed after reaching file.max.records record size. | | GCS bucket check | gcs.bucket.check | If set to true, the connector will attempt to put a test file to the GCS bucket to validate access. Default is true. | | GCS retry backoff initial delay milliseconds | gcs.retry.backoff.initial.delay.ms | Initial retry delay in milliseconds. The default value is 1000. | | GCS retry backoff max delay milliseconds | gcs.retry.backoff.max.delay.ms | Maximum retry delay in milliseconds. The default value is 32000. | | GCS retry backoff delay multiplier | gcs.retry.backoff.delay.multiplier | Retry delay multiplier. The default value is 2.0. | | GCS retry backoff max attempts | gcs.retry.backoff.max.attempts | Retry max attempts. The default value is 6. | | GCS retry backoff total timeout milliseconds | gcs.retry.backoff.total.timeout.ms | Retry total timeout in milliseconds. The default value is 50000. | | Retry back-off | kafka.retry.backoff.ms | Retry backoff in milliseconds. In case of transient exceptions, useful for performing recovery. Maximum value is 86400000 (24 hours). | | Error tolerance | errors.tolerance | Error tolerance response during connector operation. Default value is none and signals that any error will result in an immediate connector task failure. Value of all changes the behavior to skip over problematic records. | | Dead letter queue topic name | errors.deadletterqueue.topic.name | The name of the topic to be used as the dead letter queue (DLQ) for messages that result in an error when processed by this sink connector, its transformations, or converters. The topic name is blank by default, which means that no messages are recorded in the DLQ. | | Dead letter queue topic replication factor | errors.deadletterqueue.topic .replication.factor | Replication factor used to create the dead letter queue topic when it doesn’t already exist. | | Enable error context headers | errors.deadletterqueue.context .headers.enable | When true, adds a header containing error context to the messages written to the dead letter queue. To avoid clashing with headers from the original record, all error context header keys, start with __connect.errors. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - `JSON` (`org.apache.kafka.connect.json.JsonConverter`) when your messages are JSON-encoded. Select `Message JSON contains schema`, with the `schema` and `payload` fields. - `AVRO` (`io.confluent.connect.avro.AvroConverter`) when your messages contain AVRO-encoded messages, with schema stored in the Schema Registry. - `STRING` (`org.apache.kafka.connect.storage.StringConverter`) when your messages contain textual data. - `BYTES` (`org.apache.kafka.connect.converters.ByteArrayConverter`) when your messages contain arbitrary data. You can also select the output data format for your GCS files as follows: - `CSV` to produce data in the `CSV` format. For `CSV` only, you can set `STRING` and `BYTES` input formats. - `JSON` to produce data in the `JSON` format as an array of record objects. - `JSONL` to produce data in the `JSON` format, each message as a separate JSON, one per line. - `PARQUET` to produce data in the `PARQUET` format when your messages contain schema. - `AVRO` to produce data in the `AVRO` format when your messages contain schema. ## [](#test-the-connection)Test the connection After the connector is created, check the GCS bucket for a new file. Files should appear after the file flush interval (default is 60 seconds). ## [](#troubleshoot)Troubleshoot If there are any connection issues, an error message is returned. Depending on the `GCS bucket check` property value, the error results in a failed connector (`GCS bucket check = true`) or a failed task (`GCS bucket check = false`). Select **Show Logs** to view error details. Additional errors and corrective actions follow. | Message | Action | | --- | --- | | Failed to read credentials from JSON string | The credentials given as JSON file in the GCS credentials JSON property are incorrect. Copy a valid key from the Google Cloud service account. | | The specified bucket does not exist | Create the bucket if the bucket does not exist, or correct the bucket name if the bucket exists, but the specified GCS bucket name value is incorrect. | | No files in the GCS bucket | Be sure to wait until the connector performs the first file flush (default is 60 seconds). | --- # Page 327: Create an Iceberg Sink Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-iceberg-sink-connector.md --- # Create an Iceberg Sink Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create an Iceberg Sink Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-iceberg-sink-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-iceberg-sink-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-iceberg-sink-connector.adoc description: Use the Redpanda Cloud UI to create an Iceberg Sink Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2026-03-31" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use the Iceberg Sink connector to accomplish the following: - Write data into Iceberg tables - Commit coordination for centralized Iceberg commits - Exactly-once delivery semantics - Multi-table fan-out - Row mutations (update/delete rows), upsert mode - Automatic table creation and schema evolution - Field name mapping via Iceberg’s column mapping functionality ## [](#prerequisites)Prerequisites Before you can create an Iceberg Sink connector in Redpanda Cloud, you must: 1. [Set up an Iceberg catalog](https://iceberg.apache.org/concepts/catalog/). 2. Create the Iceberg connector control topic, which cannot be used by other connectors. For details, see [Create a Topic](https://docs.redpanda.com/cloud-data-platform/develop/topics/create-topic/). ## [](#limitations)Limitations - Each Iceberg sink connector must have its own control topic, which you should create before creating the connector. ## [](#create-an-iceberg-sink-connector)Create an Iceberg Sink connector To create the Iceberg Sink connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu and then click **Create Connector**. 2. Select **Export to Iceberg**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to export | topics | Comma-separated list of the cluster topics you want to replicate. | | Topics regex | topics.regex | Java regular expression of topics to replicate. For example: specify .* to replicate all available topics in the cluster. Applicable only when Use regular expressions is selected. | | Iceberg control topic | iceberg.control.topic | The name of the control topic. You must create this topic before creating the Iceberg connector. It cannot be used by other Iceberg connectors. | | Iceberg catalog type | iceberg.catalog.type | The type of Iceberg catalog. Allowed options are: REST, HIVE, HADOOP. | | Iceberg tables | iceberg.tables | Comma-separated list of Iceberg table names, which are specified using the format {namespace}.{table}. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-iceberg-sink-connector-configuration)Advanced Iceberg Sink connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Iceberg commit timeout | iceberg.control.commit.timeout-ms | Commit timeout interval in ms. The default is 30000 (30 sec). | | Iceberg tables route field | iceberg.tables.route-field | For multi-table fan-out, the name of the field used to route records to tables. | | Iceberg tables CDC field | iceberg.tables.cdc-field | Name of the field containing the CDC operation, I, U, or D. Default is none. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - `JSON` when your messages are JSON-encoded. Select `Message JSON contains schema` with the `schema` and `payload` fields. If your messages do not contain schema, create Iceberg tables manually. - `AVRO` when your messages contain AVRO-encoded messages, with schema stored in the Schema Registry. An Iceberg table’s schema is a list of named columns. All data types are either primitives or nested types, which are maps, lists, or structs. A table schema is also a struct type. See also: [Schemas and Data Types](https://iceberg.apache.org/spec/#schemas-and-data-types) ## [](#sinking-data-produced-by-debezium-source-connector)Sinking data produced by Debezium source connector Debezium connectors produce data in CDC format. The message structure can be flattened by using Debezium built-in New Record State Extraction Single Message Transformation (SMT). Add the following properties to the Debezium connector configuration to make it produce flat messages: ```json { ... "transforms", "unwrap", "transforms.unwrap.type", "io.debezium.transforms.ExtractNewRecordState", "transforms.unwrap.drop.tombstones", "false", ... } ``` Depending on your particular use case, you can apply the SMT to a Debezium connector, or to a sink connector that consumes messages that the Debezium connector produces. To enable Apache Kafka to retain the Debezium change event messages in their original format, configure the SMT for a sink connector. See also: [Debezium New Record State Extraction SMT](https://debezium.io/documentation/reference/stable/transformations/event-flattening.html) ## [](#use-analytical-tools-with-iceberg)Use analytical tools with Iceberg Iceberg serves as a single storage solution for analytical data. It is inexpensive to read from various tools such as AWS Athena, Snowflake, or Apache Spark. Traditionally, data import involved pushing data to every tool, incurring high costs for data transfer and storage. Alternatively, you could use plain S3 buckets with Avro or CSV files, but this struggles with schema evolution. [Apache Iceberg](https://iceberg.apache.org) addresses all of these challenges: cost of data transfer, multiple data copies in storage, and support for schema evolution. ![Iceberg sink connector diagram](https://docs.redpanda.com/cloud-data-platform/shared/_images/iceberg_sink_connector_diagram.png) The following example uses: - Iceberg REST catalog - AWS S3 bucket as the storage for Iceberg files - Apache Spark, which reads the Iceberg data from an S3 bucket ```yaml version: '3' services: redpanda: image: docker.redpanda.com/redpandadata/redpanda:latest command: - redpanda start - --smp 1 - --overprovisioned - --node-id 0 - --reserve-memory 0M - --check=false - --set redpanda.auto_create_topics_enabled=false - --kafka-addr PLAINTEXT://0.0.0.0:29092,OUTSIDE://0.0.0.0:9092 - --advertise-kafka-addr PLAINTEXT://redpanda:29092,OUTSIDE://localhost:9092 - --pandaproxy-addr 0.0.0.0:8082 - --advertise-pandaproxy-addr localhost:8082 ports: - 8081:8081 - 8082:8082 - 9092:9092 - 9644:9644 - 29092:29092 console: image: docker.redpanda.com/redpandadata/console:latest restart: on-failure entrypoint: /bin/sh command: -c "echo \"$$CONSOLE_CONFIG_FILE\" > /tmp/config.yml; /app/console" environment: CONFIG_FILEPATH: /tmp/config.yml CONSOLE_CONFIG_FILE: | kafka: brokers: ["redpanda:29092"] schemaRegistry: enabled: true urls: ["http://redpanda:8081"] connect: enabled: true clusters: - name: connectors url: http://connect:8083 ports: - "8090:8080" depends_on: - redpanda connect: image: docker.redpanda.com/redpandadata/connectors:latest hostname: connect depends_on: - redpanda - spark-iceberg ports: - "8083:8083" - "9404:9404" environment: CONNECT_CONFIGURATION: | key.converter=org.apache.kafka.connect.converters.ByteArrayConverter value.converter=org.apache.kafka.connect.converters.ByteArrayConverter group.id=connectors-cluster offset.storage.topic=_internal_connectors_offsets config.storage.topic=_internal_connectors_configs status.storage.topic=_internal_connectors_status config.storage.replication.factor=-1 offset.storage.replication.factor=-1 status.storage.replication.factor=-1 producer.linger.ms=1 producer.batch.size=131072 config.providers=file config.providers.file.class=org.apache.kafka.common.config.provider.FileConfigProvider CONNECT_BOOTSTRAP_SERVERS: redpanda:29092 SCHEMA_REGISTRY_URL: http://redpanda:8081 CONNECT_GC_LOG_ENABLED: "false" CONNECT_HEAP_OPTS: -Xms512M -Xmx512M CONNECT_LOG_LEVEL: info CONNECT_TOPIC_LOG_ENABLED: "true" CONNECT_PLUGIN_PATH: "/opt/kafka/connect-plugins" spark-iceberg: image: tabulario/spark-iceberg:3.4.1_1.3.1 build: spark/ depends_on: - rest volumes: - ./warehouse:/home/iceberg/warehouse environment: - AWS_ACCESS_KEY_ID=${AWS_ACCESS_KEY_ID} - AWS_SECRET_ACCESS_KEY=${AWS_SECRET_ACCESS_KEY} - AWS_REGION=${AWS_REGION} ports: - 8888:8888 - 8080:8080 - 10000:10000 - 10001:10001 rest: image: tabulario/iceberg-rest:0.6.0 ports: - 8181:8181 environment: - AWS_ACCESS_KEY_ID=${AWS_ACCESS_KEY_ID} - AWS_SECRET_ACCESS_KEY=${AWS_SECRET_ACCESS_KEY} - AWS_REGION=${AWS_REGION} - CATALOG_WAREHOUSE=s3://bucket-name/ - CATALOG_IO__IMPL=org.apache.iceberg.aws.s3.S3FileIO ``` Use Spark-SQL to: - List databases: ```none spark-sql ()> show databases; testdb ``` - Show tables in database: ```none spark-sql ()> show tables in testdb; testtable ``` - Select data from table: ```none spark-sql ()> select * from testdb.testtable; ``` ## [](#use-with-aws-glue-data-catalog-and-aws-lake-formation)Use with AWS Glue Data Catalog and AWS Lake Formation The connector can be used with the AWS Glue Data Catalog and the AWS Lake Formation service. AWS Lake Formation only lets you use the role form of authentication. The connectors UI does not support Lake Formation-specific properties. Use the JSON editor instead. Sample configuration: ```json { ... "iceberg.catalog.client.assume-role.region": "the-region", "iceberg.catalog.client.assume-role.arn": "arn:aws:iam::account-number:role/role-name", "iceberg.catalog.glue.account-id": "NNN", "iceberg.catalog.catalog-impl": "org.apache.iceberg.aws.glue.GlueCatalog", "iceberg.catalog.client.assume-role.tags.LakeFormationAuthorizedCaller": "iceberg-connect", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog": "catalog_name", "iceberg.catalog.warehouse": "s3://bucket-name/my/data", "iceberg.catalog.s3.path-style-access": "true" } ``` ## [](#test-the-connection)Test the connection After the connector is created, execute SELECT query on the Iceberg table to verify data. It may take a couple of minutes for the records to be visible in Iceberg. Check connector state and logs for errors. ## [](#troubleshoot)Troubleshoot Iceberg connection settings are checked for validity during first data processing. The connector can be successfully created with incorrect configuration and fail only when there are messages in source topic to process. | Message | Action | | --- | --- | | NoSuchTableException: Table does not exist | Make sure Iceberg table exists and the connector iceberg.tables configuration contains correct table name in {namespace}.{table} format. | | UnknownHostException: incorrectcatalog: Name or service not known | Cannot connect to Iceberg catalog. Check if Iceberg catalog URI is correct and accessible. | | DataException: An error occurred converting record, topic: topicName, partition, 0, offset: 0 | The connector cannot read the message format. Ensure the connector mapping configuration and data format are correct. | | NullPointerException: Cannot invoke "java.lang.Long.longValue()" because "value" is null | The connector cannot read the message format. Ensure the connector mapping configuration and data format are correct. | ## [](#suggested-reading)Suggested reading - For details about the Iceberg Sink connector configuration properties, see [Iceberg-Kafka-Connect](https://github.com/tabular-io/iceberg-kafka-connect) - For details about the Iceberg Sink connector internals, see [Iceberg-Kafka-Connect documentation](https://github.com/tabular-io/iceberg-kafka-connect/tree/main/docs) --- # Page 328: Create a JDBC Sink Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-jdbc-sink-connector.md --- # Create a JDBC Sink Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a JDBC Sink Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-jdbc-sink-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-jdbc-sink-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-jdbc-sink-connector.adoc description: Use the Redpanda Cloud UI to create a JDBC Sink Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use a JDBC Sink connector to export structured data from Redpanda to a relational database. ## [](#prerequisites)Prerequisites Before you can create a JDBC Sink connector in the Redpanda Cloud, you must have a: - Relational database instance that is accessible from the JDBC Sink connector instance - Database user ## [](#limitations)Limitations The JDBC Sink connector has the following limitations: - Only `JSON` or `AVRO` formats can be used as a value converter. - Only the following databases are supported: - MySQL 5.7 and 8.0 - PostgreSQL 8.2 and higher using the version 3.0 of the PostgreSQL® protocol - SQLite - SQL Server - Microsoft SQL versions: Azure SQL Database, Azure Synapse Analytics, Azure SQL Managed Instance, SQL Server 2014, SQL Server 2016, SQL Server 2017, SQL Server 2019 ## [](#create-a-jdbc-sink-connector)Create a JDBC Sink connector To create the JDBC Sink connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Export to JDBC**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to export | topics | Comma-separated list of the cluster topics you want to replicate. | | Topics regex | topics.regex | Java regular expression of topics to replicate. For example: specify .* to replicate all available topics in the cluster. Applicable only when Use regular expressions is selected. | | JDBC URL | connection.url | The database connection JDBC URL. | | User | connection.user | Name of the database user to be used when connecting to the database. | | Password | connection.password | Password of the database user to be used when connecting to the database. | | Redpanda message key format | key.converter | Format of the key in the Redpanda topic. BYTES is the default. | | Redpanda message value format | value.converter | Format of the value in the Redpanda topic. JSON is the default. | | Auto-create | auto.create | When enabled, automatically creates the destination table (if it is missing) based on the record schema (issues a CREATE). The default is disabled. | | Max Tasks | tasks.max | Maximum number of tasks to use for this connector. The default is 1. Each task replicates exclusive set of partitions assigned to it. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-jdbc-sink-connector-configuration)Advanced JDBC Sink connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Include fields | fields.whitelist | List of comma-separated record value field names. If the value of this property is empty, the connector uses all fields from the record to migrate to a database. Otherwise, the connector uses only the record fields that are specified (in a comma-separated format). Note that Primary Key Fields is applied independently in the context of which fields form the primary key columns in the destination database, while this configuration is applicable for the other columns. | | Topics to tables mapping | topics.to.tables.mapping | Kafka topics to database tables mapping. Comma-separated list of topic to table mapping in the format: topic_name:table_name. If the destination table is found in the mapping, then it overrides the generated one defined in table.name.format. | | Table name format | table.name.format | A format string for the destination table name, which may contain ${topic} as a placeholder for the original topic name. For example, kafka_${topic} for the topic orders maps to the table name kafka_orders. The default is ${topic}. | | Table name normalize | table.name.normalize | Specifies whether or not to normalize destination table names for topics. When enabled, the alphanumeric characters (a-z, A-Z, 0-9) and remain as is, others (such as .) are replaced with . By default, is disabled. | | Quote SQL identifiers | sql.quote.identifiers | Specifies whether or not to delimit (in most databases, a quote with double quotation marks) identifiers (for example, table names and column names) in SQL statements. By default, enabled. | | Auto-evolve | auto.evolve | Whether to automatically add columns in the table schema when found to be missing relative to the record schema by issuing ALTER. | | Batch size | batch.size | Specifies how many records to attempt to batch together for insertion into the destination table, when possible. The default is 3000. | | DB time zone | db.timezone | Name of the JDBC timezone that should be used in the connector when querying with time-based criteria. Default is UTC. | | Insert mode | insert.mode | The insertion mode to use. The supported modes are:INSERT: standard SQL INSERT statementsMULTI: multi-row INSERT statementsUPSERT: use the appropriate upsert semantics for the target database if it is supported by the connector; for example, INSERT .. ON CONFLICT .. DO UPDATE SET ..UPDATE: use the appropriate update semantics for the target database if it is supported by the connector; for example, UPDATE. | | Primary key mode | pk.mode | The primary key mode to use. Supported modes are:NONE: no keys utilizedkafka: Kafka coordinates (the topic, partition, and offset) are used as the primary keyRECORD_KEY: fields from the record key are used, which may be a primitive or a structRECORD_VALUE: fields from the record value are used, which must be a struct. | | Primary key fields | pk.fields | Comma-separated list of primary key field names. The runtime interpretation of this configuration depends on the pk.mode. Supported modes are:none: ignored because no fields are used as primary key in this mode.kafka: must be a trio representing the Kafka coordinates (the topic, partition, and offset). Defaults to connect_topic,connect_partition,__connect_offset if empty.record_key: if empty, all fields from the key struct will be used, otherwise used to extract the desired fields. For primitive key, only a single field name must be configured.record_value: if empty, all fields from the value struct will be used, otherwise used to extract the desired fields. | | Maximum retries | max.retries | The maximum number of times to retry on errors before failing the task. The default is 10. | | Retry backoff (ms) | retry.backoff.ms | The time in milliseconds to wait before a retry attempt is made following an error. The default is 3000. | | Database dialect | dialect.name | The name of the database dialect that should be used for this connector. By default. the connector automatically determines the dialect based upon the JDBC connection URL. Use if you want to override that behavior and specify a specific dialect. | | Error tolerance | errors.tolerance | Error tolerance response during connector operation. Default value is none and signals that any error will result in an immediate connector task failure. Value of all changes the behavior to skip over problematic records. | | Dead letter queue topic name | errors.deadletterqueue.topic.name | The name of the topic to be used as the dead letter queue (DLQ) for messages that result in an error when processed by this sink connector, its transformations, or converters. The topic name is blank by default, which means that no messages are recorded in the DLQ. | | Dead letter queue topic replication factor | errors.deadletterqueue.topic .replication.factor | Replication factor used to create the dead letter queue topic when it doesn’t already exist. | | Enable error context headers | errors.deadletterqueue.context .headers.enable | When true, adds a header containing error context to the messages written to the dead letter queue. To avoid clashing with headers from the original record, all error context header keys, start with __connect.errors. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - Use the default `Redpanda message value format` = `JSON` (`org.apache.kafka.connect.json.JsonConverter`) property in your configuration. - Topics should contain data in JSON format with a defined JSON schema. For example: ```json { "schema": { "type": "struct", "fields": [ ] }, "payload": { } } ``` ## [](#test-the-connection)Test the connection After the connector is created, ensure that: - There are no errors in logs and in Redpanda Console. - Database tables contain data from Redpanda topics. ## [](#troubleshoot)Troubleshoot JDBC Sink connector issues are reported as failed tasks. Select **Show Logs** to view error details. | Message | Action | | --- | --- | | PSQLException: FATAL: database "invalid-database" does not exist | Make sure the JDBC URL specifies an existing database name. | | UnknownHostException: invalid-host | Make sure the JDBC URL specifies a valid database host name. | | PSQLException: Connection to postgres:1234 refused. Check that the hostname and port are correct and that the postmaster is accepting TCP/IP connections | Make sure the JDBC URL specifies a valid database host name and port, and that the port is accessible. | | PSQLException: FATAL: password authentication failed for user "postgres" | Verify that the User and Password are correct. | | ConnectException: topic_name.Value (STRUCT) type doesn’t have a mapping to the SQL database column type | The JDBC Sink connector is not compatible with the Debezium PostgreSQL Source connector. Kafka Connect JSON produced by the Debezium Connector is not compatible with what the JDBC Sink Connector is expecting. Try changing a topic name. The JDBC Source connector is compatible with the JDBC Sink connector, and can be used as an alternative for a Debezium PostgreSQL source connector. | --- # Page 329: Create a JDBC Source Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-jdbc-source-connector.md --- # Create a JDBC Source Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a JDBC Source Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-jdbc-source-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-jdbc-source-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-jdbc-source-connector.adoc description: Use the Redpanda Cloud UI to create a JDBC Source Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use a JDBC Source connector to import batches of rows from MySQL, PostgreSQL, SQLite, and SQL Server relational databases into Redpanda topics. ## [](#prerequisites)Prerequisites - Relational database instance that is accessible from the JDBC Source connector instance. - Database user has been created. ## [](#limitations)Limitations The JDBC Source connector has the following limitations: - Only `JSON` or `AVRO` formats can be used as a value converter. - Only the following databases are supported: - MySQL 5.7 and 8.0 - PostgreSQL 8.2 and higher using the version 3.0 of the PostgreSQL® protocol - SQLite - SQL Server - Microsoft SQL versions: Azure SQL Database, Azure Synapse Analytics, Azure SQL Managed Instance, SQL Server 2014, SQL Server 2016, SQL Server 2017, SQL Server 2019 ## [](#create-a-jdbc-source-connector)Create a JDBC Source connector To create the JDBC Source connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from JDBC**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topic prefix | topic.prefix | Prefix to prepend to table names to generate the name of the Kafka topic to which to publish data, or in the case of a custom query, the full name of the topic to publish to. | | JDBC URL | connection.url | The database connection JDBC URL. | | User | connection.user | Name of the database user to be used when connecting to the database. | | Password | connection.password | Password of the database user to be used when connecting to the database. | | Redpanda message value format | value.converter | Format of the value in the Redpanda topic. JSON is the default. | | Max Tasks | tasks.max | Maximum number of tasks to use for this connector. The default is 1. Each task replicates an exclusive set of partitions assigned to it. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-jdbc-source-connector-configuration)Advanced JDBC Source connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | JDBC connection attempts | connection.attempts | Maximum number of attempts to retrieve a valid JDBC connection. The default is 3. | | JDBC connection backoff (ms)) | connection.backoff.ms | Backoff time between connection attempts. The default is 10000. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. BYTES is the default. | | Kafka message headers format | header.converter | Format of the headers in the Kafka topic. The default is SIMPLE. | | Include tables | table.whitelist | List of tables to include when copying. If specified, you cannot specify the Exclude Tables property. | | Exclude tables | table.blacklist | List of tables to exclude when copying. If specified, you cannot specify the Include Tables property. | | Qualify table names | table.names.qualify | Specifies whether or not to use fully-qualified table names when querying the database. If disabled, queries are performed with unqualified table names. This property may be useful if the database has been configured with a search path that automatically directs unqualified queries to the correct table when there are multiple tables available with the same unqualified name. | | Catalog pattern | catalog.pattern | Catalog pattern used to fetch table metadata from the database. null (default) means that the catalog name is not to be used to narrow the search to fetch all table metadata, regardless of the catalog. `""`retrieves those without a catalog. | | Schema pattern | schema.pattern | Schema pattern used to fetch table metadata from the database: * "" retrieves those without a schema. * null (default) specifies that the schema name is not to be used to narrow the search, so that all table metadata is fetched, regardless of the schema. | | DB time zone | db.timezone | Name of the JDBC timezone that should be used in the connector when querying with time-based criteria. Default is UTC. | | Max rows per batch | batch.max.rows | Maximum number of rows to include in a single batch when polling for new data. You can use this property to limit the amount of data buffered internally in the connector. The default is 100. | | Incrementing column name | incrementing.column.name | The name of the strictly incrementing column to use to detect new rows. An empty value indicates the column should be autodetected by looking for an auto-incrementing column. This column cannot not be nullable. | | Incrementing column initial value | incrementing.initial | For the incrementing column, consider only the rows that have a value greater than this. Specify if you need to pick up rows with negative or zero value, or if you want to skip rows. The default is -1. To avoid excessive memory usage leading to a large data set, carefully select the initial value. | | Table loading mode | mode | The mode for updating a table each time it is polled. Options include:bulk: perform a bulk load of the entire table each time it is polled.incrementing: use a strictly incrementing column on each table to detect only new rows. Note that this does not detect modifications or deletions of existing rows.timestamp: use a timestamp (or timestamp-like) column to detect new and modified rows. Based on the assumption that the column is updated with each write, and that values are monotonically incrementing, but not necessarily unique.timestamp+incrementing: use two columns, a timestamp column that detects new and modified rows, and a strictly incrementing column, which provides a globally unique ID for updates so that each row can be assigned a unique stream offset. | | Map Numeric Values, Integral or Decimal, By Precision and Scale | numeric.mapping | Map NUMERIC values by precision and optionally scale to integral or decimal types:none (default): use if all NUMERIC columns are to be represented by Connect’s DECIMAL logical type. This may lead to serialization issues with Avro because Connect’s DECIMAL type is mapped to its binary representationbest_fit: use if NUMERIC columns should be cast to Connect’s INT8, INT16, INT32, INT64, or FLOAT64 based upon the column’s precision and scale. Is often preferred because it maps to the most appropriate primitive type.precision_only: use to map NUMERIC columns based only on the column’s precision (assuming that column’s scale is 0). | | Poll interval (ms) | poll.interval.ms | Frequency used to poll for new data in each table. The default is 5000. | | Query | query | Specifies the query to use to select new or updated rows. Use to join tables, select subsets of columns in a table, or to filter data. When specified, this connector will only copy data using this query, and whole-table copying will be disabled. Different query modes may still be used for incremental updates, but to properly construct the incremental query, it must be possible to append a WHERE clause to this query (that is, no WHERE clauses can be used). If you use a WHERE clause, it must handle incremental queries itself. | | Quote SQL identifiers | sql.quote.identifiers | Specifies whether or not to delimit (in most databases, a quote with double quotation marks) identifiers (for example, table names and column names) in SQL statements. | | Metadata change monitoring interval (ms) | table.poll.interval.ms | Frequency to poll for new or removed tables, which may result in updated task configurations to start polling for data in added tables, or stop polling for data in removed tables. The default is 60000. | | Table types | table.types | By default, the JDBC connector only detects tables with type TABLE from the source Database. This property allows a command separated list of table types to extract. Options include: TABLE (default) VIEW SYSTEM TABLE GLOBAL TEMPORARY LOCAL TEMPORARY ALIAS SYNONYM. In most cases, it is best to specify TABLE or VIEW. | | Timestamp column name | timestamp.column.name | Comma separated list of one or more timestamp columns to detect new or modified rows using the COALESCE SQL function. Rows whose first non-null timestamp value is greater than the largest previous timestamp value seen aare discovered with each poll. At least one column should not be nullable. | | Delay interval (ms) | timestamp.delay.interval.ms | The amount of time to wait after a row with a certain timestamp appears before including it in the result. You can add a delay to allow transactions with earlier timestamp to complete. The first execution fetches all available records (that is, starting at a timestamp greater than 0) until current time minus the delay. Every following execution will get data from the last time fetched until the current time, minus the delay. | | Initial timestamp (ms) since epoch | timestamp.initial.ms | The initial value of the timestamp when selecting records. Value can be negative. The records having a timestamp greater than the value are included in the result. To avoid excessive memory usage leading to a large data set, carefully select the initial timestamp. | | Validate non null | validate.non.null | By default, the JDBC connector validates that all incrementing and timestamp tables have NOT NULL set for the columns being used as their ID/timestamp. If the tables don’t, then the JDBC connector will fail to start. Setting to false disables these checks. | | Database dialect | dialect.name | The name of the database dialect that should be used for this connector. By default. the connector automatically determines the dialect based upon the JDBC connection URL. Use if you want to override that behavior and specify a specific dialect. | | Topic creation enabled | topic.creation.enable | Specifies whether or not to allow automatic creation of topics. Default is enabled. | | Topic creation partitions | topic.creation.default. partitions | Specifies the number of partitions for the created topics. The default is 1. | | Topic creation replication factor | topic.creation.default. replication.factor | Specifies the replication factor for the created topics. The default is -1. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - You can use Schema Registry as an alternative to the JSON schema. - Use `Kafka message value format` = `AVRO` (`io.confluent.connect.avro.AvroConverter`) to use Schema Registry with `AvroConverter`. Use the following properties to select the database data set to read from: - `Include tables` - `Exclude tables` - `Catalog pattern` - `Schema pattern` ## [](#test-the-connection)Test the connection After the connector is created, check to ensure that: - There are no errors in logs and in Redpanda Console. - Redpanda topics contain data from relational database tables. ## [](#troubleshoot)Troubleshoot Most JDBC Source connector issues are identified in the connector creation phase. Invalid `Include tables` are reported in logs. Select **Show Logs** to view error details. | Message | Action | | --- | --- | | PSQLException: FATAL: database "invalid-database" does not exist | Make sure the JDBC URL specifies an existing database name. | | PSQLException: The connection attempt failed. for configuration Couldn’t open connection / PSQLException: Connection to postgres:1234 refused. Check that the hostname and port are correct and that the postmaster is accepting TCP/IP connections | Make sure the JDBC URL specifies a valid database host name and port, and that the port is accessible. | | PSQLException: FATAL: password authentication failed for user "postgres" | Verify that the User and Password are correct. | | IllegalArgumentException: Number of groups must be positive. | Make sure Include tables contains a valid tables list.Include tables setting is case-sensitive, even though the underlying database isn’t. Revise Include tables = tablename to Include Tables: tableName.Postgres occasionally refuses a connection for the first time. Retry creating the connector. | --- # Page 330: Create a MirrorMaker2 Checkpoint Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-mmaker-checkpoint-connector.md --- # Create a MirrorMaker2 Checkpoint Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a MirrorMaker2 Checkpoint Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-mmaker-checkpoint-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-mmaker-checkpoint-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-mmaker-checkpoint-connector.adoc description: Use the Redpanda Cloud UI to create a MirrorMaker2 Checkpoint Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use the MirrorMaker2 Checkpoint connector to import consumer group offsets from other Kafka clusters. ## [](#prerequisites)Prerequisites - The external Kafka cluster is accessible. - A service account with read-only access to the external cluster is available. - The Kafka cluster topics connector is running for the same source cluster, with a matching configuration. ## [](#limitations)Limitations The MirrorMaker2 Checkpoint connector does not migrate consumer group offsets that are lower than the highest offsets synced by the MirrorMaker2 Source connector by the time the MirrorMaker2 Checkpoint connector is started. ## [](#create-a-mirrormaker2-checkpoint-connector)Create a MirrorMaker2 Checkpoint connector To create the MirrorMaker2 Checkpoint connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from Kafka cluster offsets**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to replicate | topics | Comma-separated topic names and regexes you want to replicate. | | Source cluster broker list | source.cluster.bootstrap.servers | A comma-separated list of host/port pairs to use for establishing the initial connection to the Kafka cluster. The client will make use of all servers regardless of which servers are specified here for bootstrapping. | | Source cluster security protocol | source.cluster.security.protocol | The protocol used to communicate with source brokers. The default is PLAINTEXT. | | Source cluster SASL mechanism | source.cluster.sasl.mechanism | SASL mechanism used for connections to source cluster. Default is PLAIN. | | Source cluster SASL username | source.cluster.sasl.username | SASL username used for connections to source cluster. | | Source cluster SASL password | source.cluster.sasl.password | SASL password used for connections to source cluster. | | Groups | groups | Consumer groups to replicate. Supports comma-separated group IDs and regexes. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-mirrormaker2-checkpoint-connector-configuration)Advanced MirrorMaker2 Checkpoint connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Source cluster SSL custom certificate | source.cluster.ssl.truststore.certificates | Trusted certificates in the PEM format. | | Source cluster SSL keystore key | source.cluster.ssl.keystore.key | Private key in the PEM format. | | Source cluster SSL keystore certificate chain | source.cluster.ssl.keystore.certificate.chain | Certificate chain in the PEM format. | | Topics exclude | topics.exclude | Excluded topics. Supports comma-separated topic names and regexes. | | Source cluster alias | source.cluster.alias | When using DefaultReplicationPolicy, topic names will be prefixed with it. | | Replication policy class | replication.policy.class | Class that defines the remote topic naming convention. Use IdentityReplicationPolicy to preserve topic names. DefaultReplicationPolicy prefixes the topic with the source cluster alias. | | Emit checkpoints interval seconds | emit.checkpoints.interval.seconds | Frequency of checkpoints. The default is 60. | | Sync group offsets enabled | sync.group.offsets.enabled | Specifies whether or not to periodically write the translated offsets to the __consumer_offsets topic in the target cluster, as long as no active consumers in that group are connected to the target cluster. | | Sync group offsets interval seconds | sync.group.offsets.interval.seconds | Frequency of consumer group offset sync. The default is 60. | | Refresh groups interval seconds | refresh.groups.interval.seconds | Frequency of group refreshes. The default is 600. | | Offset-Syncs topic location | offset-syncs.topic.location | The location (source or target) of the offset-syncs topic. The default is source. | | Checkpoints topic replication factor | checkpoints.topic.replication.factor | Replication factor for checkpoints topic. The default is -1. | ## [](#test-the-connection)Test the connection After the connector is created: - Ensure that there are no errors in logs and in Redpanda Console. - Wait for the Kafka cluster topics connector to catch up. Then check to confirm that the consumer groups are replicated. ## [](#use-the-connectors-api)Use the Connectors API When using the Connectors API, instead of specifying a value for `source.cluster.sasl.username` and `source.cluster.sasl.password`, you can specify a value for `source.cluster.sasl.jaas.config`. ## [](#troubleshoot)Troubleshoot Most MirrorMaker2 Checkpoint connector issues are reported as a failed task at the time of creation. Select **Show Logs** to view error details. | Message | Action | | --- | --- | | Connection to node -1 (/127.0.0.1:9092) could not be established. Broker may not be available. / LOGS: Timed out while checking for or creating topic 'mm2-offset-syncs.target.internal'. This could indicate a connectivity issue / TimeoutException: Timed out waiting for a node assignment | Make sure broker URLs are correct and that the source cluster security protocol is correct. | | SaslAuthenticationException: SASL authentication failed: security: Invalid credentials | Check to confirm that the username and password specified are correct. | | java.lang.IllegalArgumentException: No serviceName defined in either JAAS or Kafka config | Check to confirm that the username and password specified are correct. | | Client SASL mechanism 'PLAIN' not enabled in the server, enabled mechanisms are [SCRAM-SHA-256, SCRAM-SHA-512] | Check to confirm that the respective Source cluster SASL mechanism is correct. | | SaslAuthenticationException: SASL authentication failed: security: Invalid credentials | Make sure the respective Source cluster SASL mechanism is correct (for example, SCRAM-SHA-256 instead of SCRAM-SHA-512). | | terminated during authentication. This may happen due to any of the following reasons: (1) Authentication failed due to invalid credentials with brokers older than 1.0.0, (2) Firewall blocking Kafka TLS traffic (eg it may only allow HTTPS traffic), (3) Transient network issue | Enable the SSL using Source cluster security protocol (specify SSL or SASL_SSL). | --- # Page 331: Create a MirrorMaker2 Heartbeat Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-mmaker-heartbeat-connector.md --- # Create a MirrorMaker2 Heartbeat Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a MirrorMaker2 Heartbeat Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-mmaker-heartbeat-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-mmaker-heartbeat-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-mmaker-heartbeat-connector.adoc description: Use the Redpanda Cloud UI to create a MirrorMaker2 Heartbeat Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use a MirrorMaker2 Heartbeat connector to generate heartbeat messages to a local cluster’s `heartbeat` topic. There are no prerequisites or limitations associated with this connector. ## [](#create-a-mirrormaker2-heartbeat-connector)Create a MirrorMaker2 Heartbeat connector To create the MirrorMaker2 Heartbeat connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from Heartbeat**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Emit heartbeats interval seconds | emit.heartbeats.interval.seconds | Frequency of heartbeats. The default is 1. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-mirrormaker2-heartbeat-connector-configuration)Advanced MirrorMaker2 Heartbeat connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Source cluster alias | source.cluster.alias | Used to generate the heartbeat topic key. The default is source. | | Target cluster alias | target.cluster.alias | Used to generate the heartbeat topic key. The default is target. | | Heartbeats topic replication factor | heartbeats.topic.replication.factor | Replication factor for heartbeats topic. The default is -1. | ## [](#test-the-connection)Test the connection After the connector is created, check to ensure that: - There are no errors in logs and in Redpanda Console. - Check to confirm the `heartbeat` topic has heartbeat messages. --- # Page 332: Create a MirrorMaker2 Source Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-mmaker-source-connector.md --- # Create a MirrorMaker2 Source Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a MirrorMaker2 Source Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-mmaker-source-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-mmaker-source-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-mmaker-source-connector.adoc description: Use the Redpanda Cloud UI to create a MirrorMaker2 Source Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use a MirrorMaker2 Source connector to import messages from another Kafka cluster. You can also use it to: - Replicate messages from an external Kafka or Redpanda cluster. - Create topics on the local cluster, with a configuration matching external topics. - Replicate topic access-control lists (ACLs). ## [](#prerequisites)Prerequisites - The external Kafka cluster must be accessible. - A service account with full access to the external cluster must be available. You can also use a service account with read-only ACLs when the `offset-syncs` topic location is set to `target`. You must have describe and/or describe-configs ACLs for the connector to read topic configurations on the source cluster and create the topics on the target cluster, unless you create the topics yourself. ## [](#limitations)Limitations - ACLs are copied, but service accounts are not created. - Only topic ACLs are copied (group ACLs are not). - Only ACLs for topics matching the connector configuration are copied (write ACLs are not copied). - All permissions ACLs are downgraded to read-only. ## [](#create-a-mirrormaker2-source-connector)Create a MirrorMaker2 Source connector To create the MirrorMaker2 Source connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from Kafka cluster topics**. 3. On the **Create Connector** form page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Regexes of topics to import | topics | Comma-separated topic names and regexes you want to replicate. | | Source cluster broker list | source.cluster.bootstrap.servers | A comma-separated list of host/port pairs to use for establishing the initial connection to the Kafka cluster. The client will make use of all servers regardless of which servers are specified here for bootstrapping. This list only impacts the initial hosts used to discover the full set of servers, and should be in the form host1:port1,host2:port2,.... Because these servers are only used for the initial connection to discover the full cluster membership (which may change dynamically), it need not contain the full set of servers (you may want more than one, though, in case a server is down). | | Source cluster security protocol | source.cluster.security.protocol | The protocol to use to communicate with source brokers. Default is PLAINTEXT. | | Source cluster SASL mechanism | source.cluster.sasl.mechanism | SASL mechanism used for connections to source cluster. Default is PLAIN. | | Source cluster SASL username | source.cluster.sasl.username | SASL username used for connections to source cluster. | | Source cluster SASL password | source.cluster.sasl.password | SASL password used for connections to source cluster. | | Sync topic configs enabled | sync.topic.configs.enabled | Specifies whether to periodically configure remote topics to match their corresponding upstream topics. | | Sync topic ACLs enabled | sync.topic.acls.enabled | Specifies whether or not to periodically configure remote topic ACLs to match their corresponding upstream topics. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. > 📝 **NOTE** > > Offsets are not guaranteed to match between the source and target. For example, if data-retention deletes occur on the source topic and the earliest offset is `#5000`, then when that event is created on the target topic the offset for that event will be `#0`. > > Events written on the target topic use the timestamp that was set on the source event. For example, if the source event has a timestamp `2023-05-22 17:00`, then this would also be the timestamp on the target event. ### [](#advanced-mirrormaker2-source-connector-configuration)Advanced MirrorMaker2 Source connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Source cluster SSL custom certificate | source.cluster.ssl.truststore.certificates | Trusted certificates in the PEM format. | | Source cluster SSL keystore key | source.cluster.ssl.keystore.key | Private key in the PEM format. | | Source cluster SSL keystore certificate chain | source.cluster.ssl.keystore.certificate.chain | Certificate chain in the PEM format. | | Sync topic configs interval seconds | sync.topic.configs.interval.seconds | Frequency of topic config sync. | | Sync topic ACLs interval seconds | sync.topic.acls.interval.seconds | Frequency of topic ACL sync. | | Topics exclude | topics.exclude | Excluded topics. Supports comma-separated topic names and regexes. | | Source cluster alias | source.cluster.alias | When using DefaultReplicationPolicy, topic names will be prefixed with it. | | Replication policy class | replication.policy.class | Class that defines the remote topic naming convention. Use IdentityReplicationPolicy to preserve topic names. DefaultReplicationPolicy prefixes the topic with the source cluster alias. | | Replication factor | replication.factor | Replication factor for newly created remote topics. Set -1 for cluster default. | | Refresh topics interval seconds | refresh.topics.interval.seconds | Frequency of topic refresh. | | Offset-Syncs topic location | offset-syncs.topic.location | The location (source or target) of the offset-syncs topic. The default is source. | | Offset-Syncs topic replication factor | offset-syncs.topic.replication.factor | Replication factor for offset-syncs topic. The default is -1. | | Config properties exclude | config.properties.exclude | Topic config properties that should not be replicated. Supports comma-separated property names and regexes. | | Compression type | producer.override.compression.type | The compression type for all data generated by the producer. The default is none (no compression). | | Max size of a request | producer.override.max.request.size | The maximum size of a request in bytes. The default is 1048576. | | Auto offset reset | consumer.auto.offset.reset | What to do when there is no initial offset in Kafka, or if the current offset does not exist any more on the server (for example, because that data has been deleted). 'earliest' - automatically reset the offset to the earliest offset. 'latest' - automatically reset the offset to the latest offset. 'none' - throw exception to the consumer if no previous offset is found for the consumer’s group. | | Offset lag max | offset.lag.max | How out-of-sync a remote partition can be before it is resynced. This setting impacts the MirrorMaker2 Checkpoint connector as it is the maximum lag for syncing consumer groups. The default is 100 records. | ## [](#map-data)Map data The value converter does not require any schema; it copies data as bytes. ## [](#test-the-connection)Test the connection After the connector is created: - Ensure that there are no errors in logs and in Redpanda Console. - Confirm that Redpanda topics are being replicated. You should see messages coming into the topics. ## [](#use-the-connectors-api)Use the Connectors API When using the Connectors API, instead of specifying a value for `source.cluster.sasl.username` and `source.cluster.sasl.password`, you can specify a value for `source.cluster.sasl.jaas.config`. ## [](#troubleshoot)Troubleshoot Most MirrorMaker2 Source connector issues are reported as a failed task at the time of creation. Select **Show Logs** to view error details. | Message | Action | | --- | --- | | Connection to node -1 (/127.0.0.1:9092) could not be established. Broker may not be available. / LOGS: Timed out while checking for or creating topic 'mm2-offset-syncs.target.internal'. This could indicate a connectivity issue / TimeoutException: Timed out waiting for a node assignment | Make sure broker URLs are correct and that the security.protocol is correct. | | SaslAuthenticationException: SASL authentication failed: security: Invalid credentials | Confirm that the username and password specified are correct. | | Terminated during authentication. This may happen due to any of the following reasons: (1) Authentication failed due to invalid credentials with brokers older than 1.0.0, (2) Firewall blocking Kafka TLS traffic (eg it may only allow HTTPS traffic), (3) Transient network issue | Error indicates that the SSL should be enabled using Source cluster security protocol (use SSL or SASL_SSL). | | RecordTooLargeException: The message is N bytes (…​) | Use producer.override.max.request.size property to change max request size. | | RecordTooLargeException: The request included (…​) | The target server is not able to receive messages because it is too large in size. Disabled compression can be a root cause. Consider enabling compression: "Compression type": "snappy", | | Scheduler for MirrorSourceConnector caught exception in scheduled task: syncing topic ACLs | MirrorMaker2 requires an authorizer to be configured by the broker side, but it is not. Change the Sync topic ACLs enabled MirrorMaker2 property to false (default is true) to disable ACL syncing. | | TopicAuthorizationException: Topic authorization failed | Confirm the service account for the source cluster contains describe and/or describe-configs ACLs. | | OffsetOutOfRangeException Fetch position FetchPosition{offset=0, …​ ] | If the 0 offset for your topic does not exist in the source cluster, set Auto offset reset to either earliest or latest. | --- # Page 333: Create a MongoDB Sink Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-mongodb-sink-connector.md --- # Create a MongoDB Sink Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a MongoDB Sink Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-mongodb-sink-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-mongodb-sink-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-mongodb-sink-connector.adoc description: Use the Redpanda Cloud UI to create a MongoDB Sink Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. The MongoDB Sink managed connector exports Redpanda structured data to a MongoDB database. ## [](#prerequisites)Prerequisites - Valid credentials with the `readWrite` role to access the MongoDB database. For more granular access, you need to allow `insert`, `remove` and `update` actions for specific databases or collections. ## [](#limitations)Limitations If you want to use the MongoDB sink connector with the `MongoDB` CDC handler for data sourced from MongoDB (using the MongoDB source connector), you must select `STRING` or `BYTES` as the value converter for both the source and sink connectors. ## [](#create-a-mongodb-sink-connector)Create a MongoDB Sink connector To create a MongoDB Sink connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Export to MongoDB Sink**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to export | topics | A comma-separated list of the cluster topics you want to export to MongoDB. | | Topics regex | topics.regex | Java regular expression of topics to replicate. For example: specify .* to replicate all available topics in the cluster. Applicable only when Use regular expressions is selected. | | MongoDB Connection URL | connection.url | The MongoDB connection URI string to connect to your MongoDB instance or cluster. For example, mongodb://locahost/. | | MongoDB username | connection.username | A valid MongoDB user. | | MongoDB password | connection.password | The password for the account associated with the MongoDB user. | | MongoDB database name | database | The name of an existing MongoDB database to store output files in. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. Default is STRING. | | Kafka message value format | value.converter | Format of the value in the Redpanda topic. Default is STRING. | | Default MongoDB collection name | collection | (Optional). Single sink collection name to write to. If following multiple topics, then this will be the default collection to which they are mapped. | | Max Tasks | tasks.max | Maximum number of tasks to use for this connector. The default is 1. Each task replicates exclusive set of partitions assigned to it. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-mongodb-sink-connector-configuration)Advanced MongoDB Sink connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | CDC handler | change.data.capture.handler | The CDC (change data capture) handler to use for processing. The MongoDB handler requires plain JSON or BSON format. The default is NONE. | | Key projection type | key.projection.type | The type of key projection to use: either AllowList or BlockList. | | Key projection list | key.projection.list | A comma-separated list of field names for key projection. | | Value projection type | value.projection.type | Only use with Value projection list. The type of value projection to use: AllowList or BlockList. The default is NONE. | | Value projection list | value.projection.list | A comma-separated list of field names for value projection. | | Field renamer mapping | field.renamer.mapping | An inline JSON array with objects describing field name mappings. For example: [{"oldName":"key.fieldA","newName":"field1"},{"oldName":"value.xyz","newName":"abc"}]. | | Field used for time | timeseries.timefield | Name of the top level field used for time. Inserted documents must specify this field, and it must be of the BSON datetime type. | | Field describing the series | timeseries.metafield | The name of the top-level field that contains metadata in each time series document. The metadata in the specified field should be data that is used to label a unique series of documents. The metadata should rarely, if ever, change. This field is used to group related data and may be of any BSON type, except for array. The metadata field may not be the same as the timeField or _id. | | Convert the field to a BSON datetime type | timeseries.timefield.auto.convert | Converts the timeseries field to a BSON datetime type. If the value is a numeric value it will use the milliseconds from epoch. Any fractional parts are discarded. If the value is a STRING it will use the timeseries.timefield.auto.convert.date.format property to parse the date. | | DateTimeFormatter pattern for the date | timeseries.timefield.auto.convert .date.format | The DateTimeFormatter pattern to use when converting string dates. Defaults to support ISO style date times. A string is expected to contain both the date and time. If the string only contains date information, then the time since epoch is taken from the start of that day. If a string representation does not contain a timezone offset, then the extracted date and time is interpreted as UTC. | | Data expiry time in seconds | timeseries.expire.after.seconds | The amount of time in seconds that the data will be kept in MongoDB before being automatically deleted. | | Data expiry time | timeseries.granularity | The expected interval between subsequent measurements for a time series. Possible values are "seconds", "minutes" or "hours". | | Error tolerance | errors.tolerance | Error tolerance response during connector operation. Default value is none and signals that any error will result in an immediate connector task failure. Value of all changes the behavior to skip over problematic records. | | Dead letter queue topic name | errors.deadletterqueue.topic.name | The name of the topic to be used as the dead letter queue (DLQ) for messages that result in an error when processed by this sink connector, its transformations, or converters. The topic name is blank by default, which means that no messages are recorded in the DLQ. | | Dead letter queue topic replication factor | errors.deadletterqueue.topic .replication.factor | Replication factor used to create the dead letter queue topic when it doesn’t already exist. | | Enable error context headers | errors.deadletterqueue.context .headers.enable | When true, adds a header containing error context to the messages written to the dead letter queue. To avoid clashing with headers from the original record, all error context header keys, start with __connect.errors. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - `JSON` (`org.apache.kafka.connect.json.JsonConverter`) when your messages are structured JSON. Select `Message JSON contains schema`, with the `schema` and `payload` fields. - `AVRO` (`io.confluent.connect.avro.AvroConverter`) when your messages contain AVRO-encoded messages, with schema stored in the Schema Registry. - `STRING` (`org.apache.kafka.connect.storage.StringConverter`) when your messages contain plaintext JSON. - `BYTES` (`org.apache.kafka.connect.converters.ByteArrayConverter`) when your messages contain BSON. ## [](#test-the-connection)Test the connection After the connector is created, verify that your new collections apper in your MongoDB database: show collections ## [](#use-the-connectors-api)Use the Connectors API When using the Connectors API, instead of specifying a value for `connection.url`, `connection.username`, and `connection.password`, you can specify a value for `connection.uri` in the form `mongodb+srv://username:password@cluster0.xxx.mongodb.net`. ## [](#troubleshoot)Troubleshoot Issues are reported using a failed task error message. Select **Show Logs** to view error details. | Message | Action | | --- | --- | | Invalid value wrong_uri for configuration connection.uri: The connection string is invalid. Connection strings must start with either 'mongodb://' or 'mongodb+srv:// | Check to make sure the Connection URI is a valid MongoDB URL. | | Unable to connect to the server. | Check to ensure that the Connection URI is valid and that the MongoDB server accepts connections. | | Invalid user permissions authentication failed. Exception authenticating MongoCredential{mechanism=SCRAM-SHA-1, userName='user', source='admin', password=, mechanismProperties=}. | Check to ensure that you specified valid username and password credentials. | | DataException: Could not convert key into a BsonDocument. | Make sure your message keys are valid JSONs or skip configuration for fields that require valid JSON keys. | | DataException: Error: operationType field doc is missing. | Make sure the input record format is correct (produced by a MongoDB source connector if you use MongoDB CDC handler). | | DataException: Value document is missing or CDC operation is not a string | Make sure the input record format is correct (produced by a Debezium source connector if you use Debezium CDC handler). | | JsonParseException: Unrecognized token 'text': was expecting (JSON String, Number, Array, Object or token 'null', 'true' or 'false') | Make sure the input record format is JSON. | | Unexpected documentKey field type, expecting a document but found BsonString…​: {…​} | Make sure the source data is in the plain JSON or BSON format (value converter STRING or BYTES). | ## [](#suggested-reading)Suggested reading - [MongoDB Kafka Sink Connector](https://www.mongodb.com/docs/kafka-connector/current/sink-connector/) --- # Page 334: Create a MongoDB Source Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-mongodb-source-connector.md --- # Create a MongoDB Source Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a MongoDB Source Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-mongodb-source-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-mongodb-source-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-mongodb-source-connector.adoc description: Use the Redpanda Cloud UI to create a MongoDB Source Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. The MongoDB Source managed connector imports collections from a MongoDB database into Redpanda topics. ## [](#prerequisites)Prerequisites - Valid credentials with the `read` role to access the MongoDB database. For more granular access, you need to allow `find` and `changeStream` actions for specific databases or collections. ## [](#create-a-mongodb-source-connector)Create a MongoDB Source connector To create a MongoDB Source connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from MongoDB**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topic prefix | topic.prefix | Prefix to prepend to database and collection names to generate the name of the Kafka topic to which to publish data. Used by the DefaultTopicMapper. | | MongoDB Connection URL | connection.url | The MongoDB connection URL string as supported by the official drivers. For example, mongodb://locahost/. | | MongoDB username | connection.username | A valid MongoDB user. | | MongoDB password | connection.password | The password for the account associated with the MongoDB user. | | Database to watch | database | The MongoDb database from which the connector imports data into Redpanda topics. The connector monitors changes in this database. Leave the field empty to watch all databases. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. Default is STRING. Use AVRO or JSON for schematic output, STRING for plain JSON, or BYTES for BSON. | | Kafka message value format | value.converter | Format of the value in the Redpanda topic. Default is STRING. Use AVRO or JSON for schematic output, STRING for plain JSON, or BYTES for BSON. | | Collection to watch | collection | The collection in the MongoDB database to watch. If not set, then all collections are watched. | | Start up behavior when there is no source offset available | startup.mode | Specifies how the connector should start up when there is no source offset available. Resuming a change stream requires a resume token, which the connector stores as reads from the source offset. If no source offset is available, the connector may either ignore all or some existing source data, or may at first copy all existing source data and then continue with processing new data. Possible values are:latest (default): The connector creates a new change stream, processes change events from it and stores resume tokens from them, thus ignoring all existing source data.timestamp: actuates startup.mode.timestamp.* properties. If no such properties are configured, then timestamp is equivalent to latest.copy_existing: actuates startup.mode.copy.existing.* properties. The connector creates a new change stream and stores its resume token, copies all existing data from all the collections being used as the source, then processes new data starting from the stored resume token. Note that reads of all the data during the copy and subsequent change stream events may produce duplicated events. During the copy, clients can make changes to the source data, which may be represented both by the copying process and the change stream. However, as the change stream events are idempotent, it’s possible to apply them multiple times with the same effect as if they were applied once. Renaming a collection during the copying process is not supported. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-mongodb-source-connector-configuration)Advanced MongoDB Source connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Enable Infer Schemas for the value | output.schema.infer.value | Specifies whether or not to infer the schema for the value. Each Document is processed in isolation, which may lead to multiple schema definitions for the data. Only enable when Kafka message value format is set to AVRO or JSON. | | startAtOperationTime | startup.mode.timestamp .start.at.operation.time | Actuated only if startup.mode = timestamp specifies the starting point for the change stream. Must be either an integer number of seconds because the Epoch is in the decimal format (for example: 30), or an instant in the ISO-8601 format with one second precision (for example: 1970-01-01T00:00:30Z), or a BSON timestamp in the canonical extended JSON (v2) format (for example: {"$timestamp": {"t": 30, "i": 0}}). You can specify 0 to start at the beginning of the oplog. Requires MongoDB 4.0 or above. For more detail, see the $changeStream definition. | | Copy existing namespace regex | startup.mode.copy.existing .namespace.regex | Use a regular expression to define which existing namespaces data should be copied from. A namespace is the database name and collection, separated by a period (for example, database.collection). Example: The following regular expression only includes collections starting with a in the demo database: demo\.a.*. | | Copy existing initial pipeline | startup.mode.copy.existing .pipeline | An inline JSON array with objects describing the pipeline operations to run when copying existing data. Specifying this property can improve the use of indexes by the copying manager and make copying more efficient. Use this property if there is any filtering of collection data in the pipeline configuration to speed up the copying process. For example: [{"$match": {"closed": "false"}}]. | | Pipeline to apply to the change stream | pipeline | An inline JSON array with objects describing the pipeline operations to run. For example: [{"$match": {"operationType": "insert"}}, {"$addFields": {"Kafka": "Rules!"}}]. | | fullDocument | change.stream.full.document | Specifies what to return for update operations when using a change stream. When set to updateLookup, the change stream for partial updates will include both a delta describing the changes to the document, and a copy of the entire document that was changed _ at some point_ after the change occurred. See db.collection.watch for more detail. | | fullDocumentBeforeChange | change.stream.full.document .before.change | Specifies the pre-image configuration when creating a change stream. The pre-image is not available in source records published while copying existing data as a result of enabling copy.existing. The pre-image configuration has no effect on copying. Requires MongoDB 6.0 or above. For details, see possible values. | | Publish only the fullDocument | publish.full.document.only | When enabled, only publishes the actual changed document (rather than the full change stream document). Automatically sets change.stream.full.document=updateLookup so updated documents will be included. | | Send a null value on a delete event | publish.full.document.only .tombstone.on.delete | When enabled, requires publish.full.document.only=true. Default is false (disabled). | | Error tolerance | mongo.errors.tolerance | Error tolerance response during connector operation. Default value is none and signals that any error will result in an immediate connector task failure. Value of all changes the behavior to skip over problematic records. | | Heartbeat interval milliseconds | heartbeat.interval.ms | The length of time it takes when sending heartbeat messages to record the post-batch resume token when no source records have been published. Improves the resumability of the connector for low volume namespaces. Specify 0 to disable. | | heartbeat topic name | heartbeat.topic.name | The name of the topic to publish heartbeats to. Defaults to __mongodb_heartbeats. | | Offset partition name | offset.partition.name | Use to specify a custom offset partition name. If blank, the default partition name based on the connection details is used. | | Topic creation enabled | topic.creation.enable | Specifies whether or not to allow automatic creation of topics. Default is true. | | Topic creation partitions | topic.creation.default. partitions | Specifies the number of partitions for the created topics. The default is 1. | | Topic creation replication factor | topic.creation.default. replication.factor | Specifies the replication factor for the created topics. The default is -1. | ## [](#map-data)Map data - `AVRO` (`io.confluent.connect.avro.AvroConverter`) or `JSON` (`org.apache.kafka.connect.json.JsonConverter`) for output with a preset schema. Additionally, you can set `Enable Infer Schemas` for the value. Each document will be processed in isolation, which may lead to multiple schema definitions for the data. - `STRING` (`org.apache.kafka.connect.storage.StringConverter`) when your messages contain plaintext JSON. - `BYTES` (`org.apache.kafka.connect.converters.ByteArrayConverter`) when your messages contain BSON. After the connector is created, check to ensure that: - There are no errors in logs and in Redpanda Console. - Redpanda topics contain data from relational database tables. ## [](#use-the-connectors-api)Use the Connectors API When using the Connectors API, instead of specifying a value for `connection.url`, `connection.username`, and `connection.password`, you can specify a value for `connection.uri` in the form `mongodb+srv://username:password@cluster0.xxx.mongodb.net`. ## [](#troubleshoot)Troubleshoot Most MongoDB Source connector issues are identified in the connector creation phase. Invalid Include Tables are reported in logs. Select **Show Logs** to view error details. | Message | Action | | --- | --- | | Invalid value wrong_uri for configuration connection.uri: The connection string is invalid. Connection strings must start with either 'mongodb://' or 'mongodb+srv:// | Check to make sure the MongoDB Connection URL is a valid MongoDB URL. | | Unable to connect to the server. | Check to ensure that the MongoDB Connection URL is valid and that the MongoDB server accepts connections. | | Invalid user permissions authentication failed. Exception authenticating MongoCredential{mechanism=SCRAM-SHA-1, userName='user', source='admin', password=, mechanismProperties=}. | Check to ensure that you specified valid username and password credentials. | | MongoCommandException: Command failed with error 8000 (AtlasError): 'user is not allowed to do action [find] on [db1.characters]' on server ac-nboibsg-shard-00-01.4hagsz0.mongodb.net:27017. The full response is {"ok": 0, "errmsg": "user is not allowed to do action [find] on [db1.characters]", "code": 8000, "codeName": "AtlasError"} | Check the permissions of the MongoDB user. Also confirm that the MongoDB server accepts connections. | | Command failed with error 286 (ChangeStreamHistoryLost): 'PlanExecutor error during aggregation :: caused by :: Resume of change stream was not possible, as the resume point may no longer be in the oplog | See Troubleshoot invalid resume token | ## [](#suggested-reading)Suggested reading - [MongoDB Kafka Source Connector](https://www.mongodb.com/docs/kafka-connector/current/source-connector/) --- # Page 335: Create a MySQL (Debezium) Source Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-mysql-source-connector.md --- # Create a MySQL (Debezium) Source Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a MySQL (Debezium) Source Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-mysql-source-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-mysql-source-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-mysql-source-connector.adoc description: Use the Redpanda Cloud UI to create a MySQL (Debezium) Source Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use a MySQL (Debezium) Source connector to import a stream of changes from MySQL, AmazonRDS, and Amazon Aurora. ## [](#prerequisites)Prerequisites - A MySQL database that is accessible from the connector instance. - A MySQL user exists. This database user for the Debezium connector must have LOCK TABLES privileges. For details, see [MySQL Creating a user](https://debezium.io/documentation/reference/stable/connectors/mysql.html#mysql-creating-user). - A [binlog must be enabled](https://debezium.io/documentation/reference/stable/connectors/mysql.html#enable-mysql-binlog) for the source MySQL cluster. ## [](#limitations)Limitations - Only `JSON`, `CloudEvents` or `AVRO` formats can be used as a a Kafka message key and value format. - The MySQL (Debezium) Source connector can work with only a single task at a time. ## [](#create-a-mysql-debezium-source-connector)Create a MySQL (Debezium) Source connector To create the MySQL (Debezium) Source connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from MySQL (Debezium)**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topic prefix | topic.prefix | A topic prefix that identifies and provides a namespace for the particular database server/cluster that is capturing changes. The topic prefix should be unique across all other connectors because it is used as a prefix for all Kafka topic names that receive events emitted by this connector. Only alphanumeric characters, hyphens, dots, and underscores are accepted. | | Hostname | database.hostname | A resolvable hostname or IP address of the MySQL database server. | | Port | database.port | Integer port number of the MySQL database server. | | User | database.user | Name of the MySQL user to be used when connecting to the MySQL database. | | Password | database.password | The password of the MySQL database user who will be connecting to the MySQL database. | | SSL mode | database.ssl.mode | Specifies whether to use an encrypted connection to the MySQL server. Select disable to use an unencrypted connection. Select 'preferred' to use an encrypted connection if the server supports secure connections. If the server does not support secure connections, falls back to an unencrypted connection. Select require to use a secure, or encrypted connection. If a secure connection cannot be established when required is selected, then the connector fails. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. | | Message key JSON contains schema | key.converter.schemas.enable | Enable to specify that the message key contains schema in the schema field. | | Kafka message value format | value.converter | Format of the value in the Redpanda topic. | | Message value JSON contains schema | value.converter.schemas.enable | Enable to specify that the message value contains schema in the schema field. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ## [](#map-data)Map data Use `Include databases`, `Include tables`, and `Include columns` to define data mapping. Alternatively, use `Exclude databases`, `Exclude tables`, and `Exclude columns`. Following is an example table in `db` database: ```sql CREATE TABLE IF NOT EXISTS Persons ( Id int PRIMARY KEY, FirstName varchar(255), LastName varchar(255) ); ``` The table has one record: ```sql INSERT INTO Persons (FirstName, LastName) VALUES (1, 'Winnie', 'the Pooh'); ``` The connector configuration for the table: ```bash column.include.list = db\\.Persons\\.(Id|FirstName|LastName) table.include.list = db\\.Persons database.include.list = db topic.prefix = frommysql ``` The connector configuration will create the Redpanda topic `frommysql.db.Persons`. For `Kafka message value format` = `JSON` (`org.apache.kafka.connect.json.JsonConverter`), the connector produces JSON messages with a schema like the following: ```json { "payload": { "schema": { // schema definition }, "payload": { "before": null, "after": { "Id": 1, "FirstName": "Winnie", "LastName": "the Pooh" }, ... } }, "encoding": "json", "schemaId": 0 } ``` For `Kafka message value format` = `AVRO` (`io.confluent.connect.avro.AvroConverter`), the connector creates a Schema Registry `frommysql.db.Persons-value` record and produces messages like the following: ```js { "payload": { "before": null, "after": { "mysql.db.Persons.Value": { "Id": 1, "FirstName": { "string": "Winnie" }, "LastName": { "string": "the Pooh" } } }, ... }, "encoding": "avro", "schemaId": 2 } ``` For `Kafka message value format` = `CloudEvents` (`io.debezium.converters.CloudEventsConverter`), the connector uses `JSON` or `AVRO` data serializer. - For `JSON` data serializer, enable `Message value CloudEvents JSON contains schema` to include JSON schema in message - For `AVRO` data serializer, connector creates schema in Schema Registry and produces messages in CloudEvents data format. ## [](#test-the-connection)Test the connection After the connector is created: - Check the connector status and confirm that there are no errors in logs and in Redpanda Console. - Review the Redpanda topic to confirm that it contains the expected data. ## [](#troubleshoot)Troubleshoot If the connector configuration is invalid, an error appears upon clicking **Finish**. If the connector fails, check the error message or select **Show Logs** to view error details. - **Topics not created by the connector** Create the topic manually or let the connector create it by setting (use desired number of partitions and replication factor): Topic creation enabled: true Topic creation partitions: 1 Topic creation replication factor: -1 Or in JSON: ```json "topic.creation.enable": true, "topic.creation.default.partitions": "1", "topic.creation.default.replication.factor": "-1" ``` - **Connector requires binlog file 'mysql-bin-changelog.257116', but MySQL only has mysql-bin-changelog.257123** Task threw an uncaught and unrecoverable exception. Task is being killed and will not recover until manually restarted" Connector requires binlog file 'mysql-bin-changelog.257116', but MySQL only has mysql-bin-changelog.257123, mysql-bin-changelog.257124, mysql-bin-changelog.257125 The connector needs a binlog file that was already purged. Change the `Snapshot mode` property from the default to `when_needed`. Additional errors and corrective actions follow. | Message | Action | | --- | --- | | Unable to connect: Public Key Retrieval is not allowed | Set Allow public key retrieval property to true. | | Unable to connect: Communications link failure | Confirm that Hostname and Port are correct. | | Access denied for user | Confirm that User and Password credentials are valid. | | Caused by: io.confluent.kafka.schemaregistry.client.rest.exceptions.RestClientException: Invalid schema Invalid namespace: from-mysql.db.Persons; error code: 422 | The Schema Registry namespace is incorrect. Consider changing the Topic prefix value, remove unallowed characters. | ## [](#suggested-reading)Suggested reading - [Debezium connector for MySQL](https://debezium.io/documentation/reference/stable/connectors/mysql.html) --- # Page 336: Create a PostgreSQL (Debezium) Source Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-postgresql-connector.md --- # Create a PostgreSQL (Debezium) Source Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a PostgreSQL (Debezium) Source Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-postgresql-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-postgresql-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-postgresql-connector.adoc description: Use the Redpanda Cloud UI to create a PostgreSQL (Debezium) Source Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use a PostgreSQL (Debezium) Source connector to import updates to Redpanda from PostgreSQL. ## [](#prerequisites)Prerequisites Before you can create a PostgreSQL (Debezium) Source connector in the Redpanda Cloud, you must: - [Make the PostgreSQL (Debezium) database accessible](https://debezium.io/documentation/reference/stable/connectors/postgresql.html#postgresql-security) from connectors instance. - [Create a PostgreSQL (Debezium) user](https://debezium.io/documentation/reference/stable/connectors/postgresql.html#postgresql-permissions) with the necessary permissions. ## [](#limitations)Limitations The PostgreSQL (Debezium) Source connector has the following limitations: - Only `JSON`, `CloudEvents` or `AVRO` formats can be used for a Kafka message key and value format. - PostgreSQL (Debezium) connector can work with only a single task at a time. ## [](#create-a-postgresql-debezium-source-connector)Create a PostgreSQL (Debezium) Source connector To create the PostgreSQL (Debezium) Source connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from PostgreSQL (Debezium)**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topic prefix | topic.prefix | A topic prefix that identifies and provides a namespace for the particular database server/cluster that is capturing changes. The topic prefix should be unique across all other connectors because it is used as a prefix for all Kafka topic names that receive events emitted by this connector. Only alphanumeric characters, hyphens, dots, and underscores are accepted. | | Hostname | database.hostname | A resolvable hostname or IP address of the PostgreSQL database server. | | Port | database.port | Integer port number of the PostgreSQL database server. | | User | database.user | Name of the PostgreSQL user to be used when connecting to the PostgreSQL database. | | Password | database.password | The password of the PostgreSQL database user who will be connecting to the PostgreSQL database. | | Database | database.dbname | The name of the database from which the connector will import changes. | | SSL mode | database.sslmode | Specifies whether to use an encrypted connection to the PostgreSQL server. Select disable to use an unencrypted connection. Select require to use a secure, or encrypted connection. If a secure connection cannot be established when required is selected, then the connector fails. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. | | Message key JSON contains schema | key.converter.schemas.enable | Enable to specify that the message key contains schema in the schema field. | | Kafka message value format | value.converter | Format of the value in the Redpanda topic. | | Message value JSON contains schema | value.converter.schemas.enable | Enable to specify that the message value contains schema in the schema field. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - Use `Include Schemas`, `Include Tables` and `Include Columns` properties to define lists of columns, tables, and schemas to read from. Alternatively, use `Exclude Schemas`, `Exclude Tables`, and `Exclude Columns` to define lists of columns, tables, and schemas to exclude from sources list. - Use only `JSON` (`org.apache.kafka.connect.json.JsonConverter`), `AVRO` (`io.confluent.connect.avro.AvroConverter`) and `CloudEvents` (`io.debezium.converters.CloudEventsConverter`) formats for the Kafka message key and value format. ## [](#test-the-connection)Test the connection After the connector is created: 1. Open Redpanda Console, click the **Topics** tab and select a topic. Check to check to confirm that it contains data migrated from PostgreSQL. Alternatively, use the `rpk consume` to check the topic. 2. Click the **Connectors** tab to confirm no issues have been reported for the connector. ## [](#troubleshoot)Troubleshoot If the connector configuration is invalid, an error appears upon clicking **Finish**. Select **Show Logs** to view error details. Additional errors and corrective actions follow. | Message | Action | | --- | --- | | Missing tables or topics | The Debezium connector replicates tables one by one. Wait for other tables to be replicated. If the database is quite large, then replication takes longer to complete. | | non-existing-db | Make sure the provided database name in Database is correct, and that the database exists. | | The connection attempt failed / Connection to postgres:9999 refused | Check to make sure that hostname and port are correct. | | Password authentication failed for user | Make sure that the User and Password credentials are valid. | | The Plugin name value is invalid | Make sure that Plugin contains a valid value, either decoderbufs or pgoutput. | | Postgres server wal_level property is replica | Specify wal_level as logical for your database. | | RecordTooLargeException: The message is 1050766 bytes when serialized, which is larger than 1048576, the value of the max.request.size configuration. | Increase the max request size to unblock the connector and allow large messages to pass: "producer.override.max.request.size": "209715200". The connector may be reaching memory limits and failing if the amount of data to pass or your messages are too large. | ## [](#suggested-reading)Suggested reading - [Debezium connector for PostgreSQL](https://debezium.io/documentation/reference/stable/connectors/postgresql.html) --- # Page 337: Create an S3 Sink Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-s3-sink-connector.md --- # Create an S3 Sink Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create an S3 Sink Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-s3-sink-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-s3-sink-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-s3-sink-connector.adoc description: Use the Redpanda Cloud UI to create an AWS S3 Sink Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. The Amazon S3 Sink connector exports Apache Kafka messages to files in AWS S3 buckets. ## [](#prerequisites)Prerequisites Before you can create an AWS S3 sink connector in the Redpanda Cloud, you must complete these tasks: 1. [Create an AWS account](https://docs.aws.amazon.com/accounts/latest/reference/manage-acct-creating.html). 2. [Create an S3 bucket](https://docs.aws.amazon.com/AmazonS3/latest/userguide/creating-bucket.html) that you will send data to. 3. [Create an IAM user](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_users_create.html) that will be used to connect to the S3 service. 4. [Attach the following policy](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_users_change-permissions.html) to the user, replacing `bucket-name` with the name you specified in step 2. ```js { "Version": "2012-10-17", "Statement": [ { "Principal": "*", "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject", "s3:AbortMultipartUpload", "s3:ListMultipartUploadParts", "s3:ListBucketMultipartUploads" ], "Resource": [ "arn:aws:s3:::bucket-name/*", "arn:aws:s3:::bucket-name" ] } ] } ``` 5. [Create access keys](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_credentials_access-keys.html) for the user created in step 3. 6. Copy the access key ID and the secret access key. You will need them to configure the connector. ## [](#limitations)Limitations - You can use only the `STRING` and `BYTES` input formats for `CSV` output format. - You can use only the `PARQUET` format when your messages contain schema. ## [](#create-an-aws-s3-sink-connector)Create an AWS S3 Sink connector To create the AWS S3 Sink connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Export to S3**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to export | topics | Comma-separated list of the cluster topics whose records will be exported to the S3 bucket. | | Topics regex | topics.regex | Java regular expression of topics to replicate. For example: specify .* to replicate all available topics in the cluster. Applicable only when Use regular expressions is selected. | | AWS access key ID | aws.access.key.id | Enter the AWS access key ID. | | AWS secret access key | aws.secret.access.key | Enter the AWS secret access key. | | AWS S3 bucket name | aws.s3.bucket.name | Specify the name of the AWS S3 bucket to which the connector is to send data. | | AWS S3 region | aws.s3.region | Select the region for the S3 bucket used for storing the records. The default us-east-1. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. The default is BYTES. | | Kafka message value format | value.converter | Format of the value in the Redpanda topic. The default is BYTES. | | S3 file format | format.output.type | Format of the files created in S3: CSV (the default), AVRO, JSON, JSONL, or PARQUET. You can use the CSV format output only with BYTES and STRING. | | Avro codec | avro.codec | The Avro compression codec to be used for Avro output files. Available values: null (the default), deflate, snappy, and bzip2. | | Max Tasks | tasks.max | Maximum number of tasks to use for this connector. The default is 1. Each task replicates exclusive set of partitions assigned to it. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-aws-s3-sink-connector-configuration)Advanced AWS S3 Sink connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | File name template | file.name.template | The template for file names on S3. Supports {{ variable }} placeholders for substituting variables. Supported placeholders are:topicpartitionstart_offset (the offset of the first record in the file)timestamp:unit=yyyy|MM|dd|HH (the timestamp of the record)key (when used, other placeholders are not substituted) | | File name prefix | file.name.prefix | The prefix to be added to the name of each file put in S3. | | Output fields | format.output.fields | Fields to place into output files. Supported values are: 'key', 'value', 'offset', 'timestamp', and 'headers'. | | Value field encoding | format.output.fields.value.encoding | The type of encoding to be used for the value field. Supported values are: 'none' and 'base64'. | | Envelope for primitives | format.output.envelope | Specifies whether or not to enable additional JSON object wrapping of the actual value. | | Output file compression | file.compression.type | The compression type to be used for files put into S3. Supported values are: 'none' (default), 'gzip', 'snappy', and 'zstd'. | | Max records per file | file.max.records | The maximum number of records to put in a single file. Must be a non-negative number. 0 is interpreted as "unlimited", which is the default. In this case files are only flushed after file.flush.interval.ms. | | File flush interval milliseconds | file.flush.interval.ms | The time interval to periodically flush files and commit offsets. Value specified must be a non-negative number. Default is 60 seconds. 0 indicates that it is disabled. In this case, files are only flushed after reaching file.max.records record size. | | AWS S3 bucket check | aws.s3.bucket.check | If set to true (default), the connector will attempt to put a test file to the S3 bucket to validate access. | | AWS S3 part size bytes | s3.part.size | The part size in S3 multi-part uploads in bytes. Maximum is 2147483647 (2GB) and default is 5242880 (5MB). | | S3 retry backoff | aws.s3.backoff.delay.ms | S3 default base sleep time (in milliseconds) for non-throttled exceptions. Default is 100. | | S3 maximum back-off | aws.s3.backoff.max.delay.ms | S3 maximum back-off time (in milliseconds) before retrying a request. Default is 20000. | | S3 max retries | aws.s3.backoff.max.retries | Maximum retry limit (if the value is greater than 30, there can be integer overflow issues during delay calculation). Default is 3. | | Error tolerance | errors.tolerance | Error tolerance response during connector operation. Default value is none and signals that any error will result in an immediate connector task failure. Value of all changes the behavior to skip over problematic records. | | Dead letter queue topic name | errors.deadletterqueue.topic.name | The name of the topic to be used as the dead letter queue (DLQ) for messages that result in an error when processed by this sink connector, its transformations, or converters. The topic name is blank by default, which means that no messages are recorded in the DLQ. | | Dead letter queue topic replication factor | errors.deadletterqueue.topic .replication.factor | Replication factor used to create the dead letter queue topic when it doesn’t already exist. | | Enable error context headers | errors.deadletterqueue.context .headers.enable | When true, adds a header containing error context to the messages written to the dead letter queue. To avoid clashing with headers from the original record, all error context header keys, start with __connect.errors. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - `JSON` (`org.apache.kafka.connect.json.JsonConverter`) when your messages are JSON-encoded. Select `Message JSON contains schema`, with the `schema` and `payload` fields. - `AVRO` (`io.confluent.connect.avro.AvroConverter`) when your messages contain AVRO-encoded messages, with schema stored in the Schema Registry. - `STRING` (`org.apache.kafka.connect.storage.StringConverter`) when your messages contain textual data. - `BYTES` (`org.apache.kafka.connect.converters.ByteArrayConverter`) when your messages contain arbitrary data. You can also select the output data format for your S3 files as follows: - `CSV` to produce data in the `CSV` format. For `CSV` only, you can set `STRING` and `BYTES` input formats. - `JSON` to produce data in the `JSON` format as an array of record objects. - `JSONL` to produce data in the `JSON` format, each message as a separate JSON, one per line. - `PARQUET` to produce data in the `PARQUET` format when your messages contain schema. - `AVRO` to produce data in the `AVRO` format when your messages contain schema. ## [](#test-the-connection)Test the connection After the connector is created, test the connection by writing to one of your topics, then checking the contents of the S3 bucket in the AWS management console. Files should appear after the file flush interval (default is 60 seconds). ## [](#troubleshoot)Troubleshoot If there are any connection issues, an error message is returned. Depending on the `AWS S3 bucket check` property value, the error results in a failed connector (`AWS S3 bucket check = true`) or a failed task (`AWS S3 bucket check = false`). Select **Show Logs** to view error details. Additional errors and corrective actions follow. | Message | Action | | --- | --- | | The AWS Access Key Id you provided does not exist in our records | AWS access key ID is invalid. Check to confirm that a valid existing AWS access key is specified. | | The authorization header is malformed; the region us-east-1 is wrong; expecting us-east-2 | The selected region (AWS S3 region) of the AWS bucket is incorrect. Check to confirm that you have specified the region in which the bucket was created. | | The specified bucket does not exist | Create the bucket specified in the AWS S3 bucket name property, or provide the correct name of the existing bucket. | | No files in the S3 bucket | Be sure to wait until the connector completes the first file flush (default 60 seconds). Verify that the topics specified are correct. Then verify that the topics contain messages to be pushed to S3. | --- # Page 338: Create a Snowflake Sink Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-snowflake-connector.md --- # Create a Snowflake Sink Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a Snowflake Sink Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-snowflake-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-snowflake-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-snowflake-connector.adoc description: Use the Redpanda Cloud UI to create a Snowflake Sink Connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use the Snowflake Sink connector to ingest and store Redpanda structured data into a Snowflake database for analytics and decision-making. ## [](#prerequisites)Prerequisites Before you can create a Snowflake Sink connector in the Redpanda Cloud, you must: 1. [Create a role](https://docs.snowflake.com/en/user-guide/kafka-connector-install#creating-a-role-to-use-the-kafka-connector) for use by Kafka Connect. 2. [Create a key pair](https://docs.snowflake.com/en/user-guide/key-pair-auth#configuring-key-pair-authentication) for authentication. 3. [Create a database](https://docs.snowflake.com/en/user-guide/getting-started-tutorial-create-objects#creating-a-database) to hold the data you intend to stream from Redpanda Cloud messages. ## [](#limitations)Limitations Refer to the [Snowflake Kafka Connector Limitations](https://docs.snowflake.com/en/user-guide/kafka-connector-overview#kafka-connector-limitations) documentation for details. ## [](#create-a-snowflake-sink-connector)Create a Snowflake Sink connector To create a Snowflake Sink connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Export to Snowflake**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topics to export | topics | A comma-separated list of the cluster topics you want to export to Snowflake. | | Topics regex | topics.regex | Java regular expression of topics to replicate. For example: specify .* to replicate all available topics in the cluster. Applicable only when Use regular expressions is selected. | | Snowflake URL name | snowflake.url.name | The Snowflake URL to be used for the connection. | | Snowflake database name | snowflake.database.name | The Snowflake database name to be used for the exported data. | | Snowflake user name | snowflake.user.name | The name of the user who created the key pair. | | Snowflake private key | snowflake.private.key | The private key name for the Snowflake user. | | Snowflake private key passphrase | snowflake.private.key.passphrase | (Optional) If created and encrypted, the passphrase of the private key. | | Snowflake role name | snowflake.role.name | The name of the role created in Prerequisites. | | Kafka message value format | value.converter | The format of the value in the Redpanda topic. The default is SNOWFLAKE_JSON. | | Max Tasks | tasks.max | Maximum number of tasks to use for this connector. The default is 1. Each task replicates exclusive set of partitions assigned to it. | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ### [](#advanced-snowflake-sink-connector-configuration)Advanced Snowflake Sink connector configuration In most instances, the preceding basic configuration properties are sufficient. If you require additional property settings, then specify any of the following _optional_ advanced connector configuration properties by selecting **Show advanced options** on the **Create Connector** page: | Property name | Property key | Description | | --- | --- | --- | | Snowflake schema name | snowflake.schema.name | The Snowflake database schema name. The default is PUBLIC. | | Snowflake ingestion method | snowflake.ingestion.method | The default, SNOWPIPE, allows for structured data, while SNOWPIPE_STREAMING is lower latency option. | | Snowflake topic2table map | snowflake.topic2table.map | (Optional) Map of topics to tables. Format is comma-separated tuples. For example, :,:. | | Buffer count records | buffer.count.records | Number of records buffered in memory per partition before triggering Snowflake ingestion. Default is 10000. | | Buffer flush time | buffer.flush.time | The time in seconds to flush cached data. Default is 120. | | Buffer size bytes | buffer.size.bytes | Cumulative size of records buffered in memory per partition before triggering Snowflake ingestion. Default is 5000000. | | Error tolerance | errors.tolerance | Error tolerance response during connector operation. Default value is none and signals that any error will result in an immediate connector task failure. Value of all changes the behavior to skip over problematic records. | | Dead letter queue topic name | errors.deadletterqueue.topic.name | The name of the topic to be used as the dead letter queue (DLQ) for messages that result in an error when processed by this sink connector, its transformations, or converters. The topic name is blank by default, which means that no messages are recorded in the DLQ. | | Dead letter queue topic replication factor | errors.deadletterqueue.topic .replication.factor | Replication factor used to create the dead letter queue topic when it doesn’t already exist. | | Enable error context headers | errors.deadletterqueue.context .headers.enable | When true, adds a header containing error context to the messages written to the dead letter queue. To avoid clashing with headers from the original record, all error context header keys, start with __connect.errors. | ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - `JSON` formatted records should use `SNOWFLAKE_JSON` (`com.snowflake.kafka.connector.records.SnowflakeJsonConverter`). - `AVRO` formatted records that use Kafka’s Schema Registry Service should use `SNOWFLAKE_AVRO` (`com.snowflake.kafka.connector.records.SnowflakeAvroConverter`). - `AVRO` formatted records that contain the schema (and therefore do not need Kafka’s Schema Registry Service) should use `SNOWFLAKE_AVRO_WITHOUT_SCHEMA_REGISTRY` (`com.snowflake.kafka.connector.records.SnowflakeAvroConverterWithoutSchemaRegistry`). - Plain text formatted records should use `STRING` (`org.apache.kafka.connect.storage.StringConverter`). ## [](#test-the-connection)Test the connection After the connector is created, verify in your Snowflake worksheet that your table is populated: SELECT \* FROM TEST.PUBLIC.TABLE\_NAME; It may take a couple of minutes for the records to be visible in Snowflake. ## [](#troubleshoot)Troubleshoot After submitting the connector for creation in Redpanda Console, the Snowflake Sink connector attempts to authenticate to the Snowflake database to validate the configuration. This validation must be successful before the connector is created. It can take up 10 seconds or more to respond. If the connector fails, check the error message or select **Show Logs** to view error details. Additional errors and corrective actions follow. | Message | Action | | --- | --- | | snowflake.url.name is not a valid snowflake url | Check to make sure Snowflake URL name contains a valid Snowflake URL. | | snowflake.user.name: Cannot connect to Snowflake | Check to make sure Snowflake user name contains a valid Snowflake user. | | snowflake.private.key must be a valid PEM RSA private key / java.lang.IllegalArgumentException: Last encoded character (before the padding, if any) is a valid base 64 alphabet but not a possible value. Expect the discarded bits to be zero. | Snowflake private key is invalid. Provide a valid key. | | snowflake.database.name+ database does not exist | Specify a valid database name in snowflake.database.name. | | Object does not exist, or operation cannot be performed | Snowflake error that can have several causes: an invalid role is being used, there is no existing Snowflake table, or an incorrect schema name is specified. Verify that the connector configuration and Snowflake settings are valid. | | Config:value.converter has provided value:com.snowflake.kafka.connector.records.SnowflakeJsonConverter. If ingestionMethod is:snowpipe_streaming, Snowflake Custom Converters are not allowed. | Use STRING for the Kafka message value format. | ## [](#suggested-reading)Suggested reading - For more about limitations, see [Kafka Connector Limitations](https://docs.snowflake.com/en/user-guide/kafka-connector-overview#kafka-connector-limitations) - For testing the connection, see [Using Worksheets for Queries / DML / DDL](https://docs.snowflake.com/en/user-guide/ui-worksheet) - For details about all Snowflake Sink connector properties, see [Kafka Configuration Properties](https://docs.snowflake.com/en/user-guide/kafka-connector-install#required-properties) --- # Page 339: Create a SQL Server (Debezium) Source Connector **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-sqlserver-connector.md --- # Create a SQL Server (Debezium) Source Connector > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a SQL Server (Debezium) Source Connector latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/create-sqlserver-connector page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/create-sqlserver-connector.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/create-sqlserver-connector.adoc description: Use the Redpanda Cloud UI to create a SQL Server (Debezium) Source Connector. page-git-created-date: "2024-10-03" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can use an SQL Server (Debezium) Source connector to import updates to Redpanda from SQL Server. ## [](#prerequisites)Prerequisites Before you can create an SQL Server (Debezium) Source connector in the Redpanda Cloud, you must: - Make the SQL Server (Debezium) database accessible from the connector instance. - Create a SQL Server (Debezium) user with the necessary permissions. ## [](#limitations)Limitations The SQL Server (Debezium) Source connector has the following limitations: - Only `JSON`, `CloudEvents` or `AVRO` formats can be used for a Kafka message key and value format. - SQL Server (Debezium) connector can work with only a single task at a time per database name. ## [](#create-an-sql-server-debezium-source-connector)Create an SQL Server (Debezium) Source connector To create the SQL Server (Debezium) Source connector: 1. In Redpanda Cloud, click **Connectors** in the navigation menu, and then click **Create Connector**. 2. Select **Import from SQL Server (Debezium)**. 3. On the **Create Connector** page, specify the following required connector configuration options: | Property name | Property key | Description | | --- | --- | --- | | Topic prefix | topic.prefix | A topic prefix that identifies and provides a namespace for the particular database server/cluster that is capturing changes. The topic prefix should be unique across all other connectors because it is used as a prefix for all Kafka topic names that receive events emitted by this connector. Only alphanumeric characters, hyphens, dots, and underscores are accepted. | | Hostname | database.hostname | A resolvable hostname or IP address of the SQL Server database server. | | Port | database.port | Integer port number of the SQL Server database server. | | User | database.user | Name of the SQL Server user to be used when connecting to the SQL Server database. | | Password | database.password | The password of the SQL Server database user who will be connecting to the SQL Server database. | | Database instance | database.instance | Specifies the instance name of the SQL Server named instance. If both database.port and database.instance are specified, database.instance is ignored. | | Databases | database.names | The comma-separated list of the SQL Server database names from which to stream the changes. | | Kafka message key format | key.converter | Format of the key in the Redpanda topic. | | Message key JSON contains schema | key.converter.schemas.enable | Enable to specify that the message key contains schema in the schema field. | | Kafka message value format | value.converter | Format of the value in the Redpanda topic. | | Message value JSON contains schema | value.converter.schemas.enable | Enable to specify that the message value contains schema in the schema field. | | Max tasks | tasks.max | The maximum number of tasks that the connector can use to capture data from the database instance. If the Databases list contains more than one element, you can increase the value of this property to a number less than or equal to the number of elements in the list. Default: 1 | | Connector name | name | Globally-unique name to use for this connector. | 4. Click **Next**. Review the connector properties specified, then click **Create**. ## [](#map-data)Map data Use the appropriate key or value converter (input data format) for your data as follows: - Use the `Include Schemas`, `Include Tables`, and `Include Columns` properties to define lists of columns, tables, and schemas to read from. Alternatively, use `Exclude Schemas`, `Exclude Tables`, and `Exclude Columns` to define lists of columns, tables, and schemas to exclude from sources list. - Use only `JSON` (`org.apache.kafka.connect.json.JsonConverter`), `AVRO` (`io.confluent.connect.avro.AvroConverter`), and `CloudEvents` (`io.debezium.converters.CloudEventsConverter`) formats for the Kafka message key and value format. ## [](#test-the-connection)Test the connection After the connector is created: 1. Open Redpanda Console, click the **Topics** tab, and select a topic. Check to confirm that it contains data migrated from SQL Server. Alternatively, run `rpk consume` to check the topic. 2. Click the **Connectors** tab to confirm that no issues have been reported for the connector. ## [](#suggested-reading)Suggested reading - [Debezium connector for SQL Server](https://debezium.io/documentation/reference/stable/connectors/sqlserver.html) --- # Page 340: Disable Kafka Connect **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc.md --- # Disable Kafka Connect > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Disable Kafka Connect latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/disable-kc page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/disable-kc.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/disable-kc.adoc description: Learn how to disable Kafka Connect using the Cloud API. page-git-created-date: "2025-08-07" page-git-modified-date: "2026-05-26" --- Kafka Connect is disabled by default on new clusters. If you previously enabled Kafka Connect on a cluster and want to disable it, you can use the [Cloud API](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview). > 📝 **NOTE** > > Redpanda Support does not manage or monitor Kafka Connect, but Support can enable the feature for your account. ## [](#verify-kafka-connect-is-enabled)Verify Kafka Connect is enabled If Kafka Connect is enabled on your cluster, you will see it configured on the **Connect** page in the Redpanda Cloud UI. You can also verify with the Cloud API: ```bash curl -sX GET "https://api.redpanda.com/v1/clusters/{cluster.id}" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -H 'accept: application/json' | jq -r '.cluster.kafka_connect' ``` Replace `{cluster.id}` with your actual cluster ID. You can find the cluster ID in the Redpanda Cloud UI. Look in the **Details** section of the cluster overview. If Kafka Connect is enabled, the response will show: ```bash "enabled": true ``` ## [](#prerequisites)Prerequisites - You have the cluster ID of a cluster that has Kafka Connect enabled. - You have a valid bearer token for the Cloud API. For details, see [Authenticate to the API](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication). > ❗ **IMPORTANT** > > Make sure to stop any active connectors gracefully before disabling Kafka Connect to avoid data loss or incomplete processing. ## [](#disable-kafka-connect)Disable Kafka Connect After you are authenticated to the Cloud API, make a [`PATCH /v1/clusters/{cluster.id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) request, replacing `{cluster.id}` with your actual cluster ID. ```bash curl -X PATCH "https://api.redpanda.com/v1/clusters/{cluster.id}" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -H "Content-Type: application/json" \ -d '{"kafka_connect":{"enabled":false}}' ``` The `PATCH` request returns the ID of a long-running operation. You can check the status of the operation by polling the [`GET /operations/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-operationservice_getoperation) endpoint: ```bash curl -X GET "https://api.redpanda.com/v1/operations/\{operation.id\}" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -H "Content-Type: application/json" ``` When the operation is complete, the status will show `"state": "STATE_COMPLETED"`. You can verify that Kafka Connect has been disabled by running the verification command from the previous section. The response should show: ```bash "enabled": false ``` --- # Page 341: Monitor Kafka Connect **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/monitor-connectors.md --- # Monitor Kafka Connect > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Monitor Kafka Connect latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/monitor-connectors page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/monitor-connectors.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/monitor-connectors.adoc description: Use metrics to monitor the health of Kafka Connect. page-git-created-date: "2024-06-06" page-git-modified-date: "2026-05-26" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. You can monitor the health of Kafka Connect with metrics that Redpanda exports through a Prometheus HTTPS endpoint. You can use Grafana to visualize the metrics and set up alerts. The most important metrics to be monitored by alerts are: - connector failed tasks - connector lag / connector lag rate ## [](#view-connector-logs)View connector logs Connector logs are written to the system topic `__redpanda.connectors_logs`. You can view logs in Redpanda Cloud on the Topics page for your cluster, or you can download logs with `rpk`. For example: ```bash # Last 100 messages (most recent) rpk topic consume __redpanda.connectors_logs -o -100 -n 100 # Last 10 minutes rpk topic consume __redpanda.connectors_logs -o @-10m:end # Stream new logs only (like tail -f) rpk topic consume __redpanda.connectors_logs -o end # Filter by connector name rpk topic consume __redpanda.connectors_logs -o @-10m:end -O json \ | jq -r 'select(.message | test(""; "i"))' ``` > 📝 **NOTE** > > Access to system topics may be restricted by organization/project roles. Log retention follows cluster/system-topic policies and messages may expire. ## [](#limitations)Limitations The connectors dashboard renders metrics that are exported by managed connectors. However, when a connector does not create a task (for example, an empty topic list), the dashboard will not show metrics for that connector. --- # Page 342: Sizing Connectors **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/sizing-connectors.md --- # Sizing Connectors > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Sizing Connectors latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/sizing-connectors page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/sizing-connectors.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/sizing-connectors.adoc description: How to choose number of tasks to set for a connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. ## [](#connector-tasks)Connector tasks When you set up a connector, its main responsibility is to validate the configuration and spawn _connector tasks_, which perform the work. Setting up multiple tasks for a connector allows for parallelization of the work, resulting in higher throughputs. Before setting up connector tasks, consider the following: - For source connectors, the ability to add tasks to achieve higher throughput depends on the connector implementation and configuration. For many connectors, only a single connector task is allowed (for example, Debezium allows a single task only). When Redpanda Cloud does not offer an option to set the number of tasks, the source connector runs only one task. - For sink connectors, parallelism is achieved by evenly distributing configured topic partitions for the connector amongst connector tasks. The number of partitions must be equal to or greater than the number of tasks. ## [](#single-task-throughput)Single task throughput Connector throughput depends on many factors, including converters used, compression, message size, and the performance of external systems. As a rule of thumb, expect a single connector task to provide 1-2 MB/s of throughput. ## [](#specify-number-of-connector-tasks-for-a-sink-connector)Specify number of connector tasks for a sink connector It can be a challenge to determine the number of connector tasks to use for a given workload, so you must experiment to find the right number. Start with low number of connector tasks and wait a couple of minutes to view performance. Keep increasing the number of tasks until satisfactory throughput is achieved. Keep in mind that the underlying infrastructure must scale to provide room for additional connector tasks. Waiting roughly 10 minutes after each change should provide sufficient time for the system to scale up. --- # Page 343: Single Message Transforms **URL**: https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/transforms.md --- # Single Message Transforms > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Single Message Transforms latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: managed-connectors/transforms page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: managed-connectors/transforms.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/managed-connectors/transforms.adoc description: Single Message Transforms (SMTs) let you modify the data and its characteristics as it passes through a connector. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-05" --- > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. Single Message Transforms (SMTs) help you modify data and its characteristics as it passes through a connector, without needing additional stream processors. Prior to using an SMT with production data, test the configuration on a smaller subset of data to verify the behavior of the SMT. ## [](#cast)Cast Cast SMT lets you change the data type of fields in a Redpanda message, updating the schema if one is present. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.Cast$Key`) or value (`org.apache.kafka.connect.transforms.Cast$Value`). ### [](#configuration)Configuration | Property key | Description | | --- | --- | | spec | Comma-separated list of field names and the type to which they should be cast; for example: my-field1:int32,my-field2:string. Allowed types are: `int8, int16, int32, int64, float32, float64, boolean, and string. | ### [](#example)Example "transforms": "Cast", "transforms.Cast.type": "org.apache.kafka.connect.transforms.Cast$Value", "transforms.Cast.spec": "price:float64" Before: {"price": 1234, "product\_id": "9987"} After: {"price": 1234.0,"product\_id": "9987"} ## [](#dropheaders)DropHeaders DropHeaders SMT removes one or more headers from each record. ### [](#configuration-2)Configuration | Property key | Description | | --- | --- | | headers | Comma-separated list of header names to drop. | ### [](#example-2)Example Sample configuration: "transforms": "DropHeader", "transforms.DropHeader.type": "org.apache.kafka.connect.transforms.DropHeaders", "transforms.DropHeader.headers": "source-id,conv-id" ## [](#eventrouter-debezium)EventRouter (Debezium) The outbox pattern is a way to safely and reliably exchange data between multiple (micro) services. An outbox pattern implementation avoids inconsistencies between a service’s internal state (as typically persisted in its database) and state in events consumed by services that need the same data. To implement the outbox pattern in a Debezium application, configure a Debezium connector to: - Capture changes in an outbox table - Apply the Debezium outbox EventRouter Single Message Transformation > 📝 **NOTE** > > EventRouter SMT is available for managed Debezium connectors only. ### [](#configuration-3)Configuration | Property key | Description | | --- | --- | | route.by.field | Specifies the name of a column in the outbox table. The default behavior is that the value in this column becomes a part of the name of the topic to which the connector emits the outbox messages. | | route.topic.replacement | Specifies the name of the topic to which the connector emits outbox messages. The default topic name is outbox.event. followed by the aggregatetype column value in the outbox table record. | | table.expand.json.payload | Specifies whether the JSON expansion of a String payload should be done. If no content is found, or if there’s a parsing error, the content is kept "as is". | | fields.additional.placement | Specifies one or more outbox table columns to add to outbox message headers or envelopes. Specify a comma-separated list of pairs. In each pair, specify the name of a column and whether you want the value to be in the header or the envelope. | | table.field.event.key | Specifies the outbox table column that contains the event key. When this column contains a value, the SMT uses that value as the key in the emitted outbox message. This is important for maintaining the correct order in Kafka partitions. | ### [](#example-3)Example Sample JSON configuration: "transforms": "outbox", "transforms.outbox.route.by.field": "type", "transforms.outbox.route.topic.replacement": "my-topic.${routedByValue}", "transforms.outbox.table.expand.json.payload": "true", "transforms.outbox.table.field.event.key": "aggregate\_id", "transforms.outbox.table.fields.additional.placement": "before:envelope", "transforms.outbox.type": "io.debezium.transforms.outbox.EventRouter" ### [](#suggested-reading)Suggested reading - [Debezium Outbox Event Router SMT](https://debezium.io/documentation/reference/stable/transformations/outbox-event-router.html) ## [](#extractfield)ExtractField ExtractField SMT pulls the specified field from a Struct when a schema is present, or a Map for schemaless data. Any null values are passed through unmodified. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.ExtractField$Key`) or value (`org.apache.kafka.connect.transforms.ExtractField$Value`). ### [](#configuration-4)Configuration | Property key | Description | | --- | --- | | field | Field name to extract. | ### [](#example-4)Example Sample configuration: "transforms": "ExtractField", "transforms.ExtractField.type": "org.apache.kafka.connect.transforms.ExtractField$Value", "transforms.ExtractField.field": "product\_id" Before: ```json {"product_id":9987,"price":1234} ``` After: ```json {"value":9987} ``` ## [](#filter)Filter Filter SMT drops all records, filtering them from subsequent transformations in the chain. This is intended to be used conditionally to filter out records matching (or not matching) a particular predicate. ### [](#configuration-5)Configuration | Property key | Description | | --- | --- | | predicate | Name of predicate filtering records. | ### [](#example-5)Example Sample configuration: "transforms": "Filter", "transforms.Filter.type": "org.apache.kafka.connect.transforms.Filter", "transforms.Filter.predicate": "IsMyTopic", "predicates": "IsMyTopic", "predicates.IsMyTopic.type": "org.apache.kafka.connect.transforms.predicates.TopicNameMatches", "predicates.IsMyTopic.pattern": "my-topic" ### [](#predicates)Predicates Managed connectors support the following predicates: #### [](#topicnamematches)TopicNameMatches `org.apache.kafka.connect.transforms.predicates.TopicNameMatches` - A predicate that is true for records with a topic name that matches the configured regular expression. | Property key | Description | | --- | --- | | pattern | A Java regular expression for matching against the name of a record’s topic. | #### [](#hasheaderkey)HasHeaderKey `org.apache.kafka.connect.transforms.predicates.HasHeaderKey` - A predicate that is true for records with at least one header with the configured name. | Property key | Description | | --- | --- | | name | The header name. | #### [](#recordistombstone)RecordIsTombstone `org.apache.kafka.connect.transforms.predicates.RecordIsTombstone` - A predicate that is true for records that are tombstones (that is, they have null values). ## [](#flatten)Flatten Flatten SMT flattens a nested data structure, generating names for each field by concatenating the field names at each level with a configurable delimiter character. Applies to Struct when a schema is present, or a Map for schemaless data. Array fields and their contents are not modified. The default delimiter is `.`. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.Flatten$Key`) or value (`org.apache.kafka.connect.transforms.Flatten$Value`). ### [](#configuration-6)Configuration | Property key | Description | | --- | --- | | delimiter | Delimiter to insert between field names from the input record when generating field names for the output record. | ### [](#example-6)Example "transforms": "flatten", "transforms.flatten.type": "org.apache.kafka.connect.transforms.Flatten$Value", "transforms.flatten.delimiter": "." Before: ```json { "user": { "id": 10, "name": { "first": "Red", "last": "Panda" } } } ``` After: ```json { "user.id": 10, "user.name.first": "Red", "user.name.last": "Panda" } ``` ## [](#headerfrom)HeaderFrom HeaderFrom SMT moves or copies fields in the key or value of a record into that record’s headers. Corresponding elements of `fields` and `headers` together identify a field and the header it should be moved or copied to. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.HeaderFrom$Key`) or value (`org.apache.kafka.connect.transforms.HeaderFrom$Value`). ### [](#configuration-7)Configuration | Property key | Description | | --- | --- | | fields | Comma-separated list of field names in the record whose values are to be copied or moved to headers. | | headers | Comma-separated list of header names, in the same order as the field names listed in the fields configuration property. | | operation | Either move if the fields are to be moved to the headers (removed from the key/value), or copy if the fields are to be copied to the headers (retained in the key/value). | ### [](#example-7)Example "transforms": "HeaderFrom", "transforms.HeaderFrom.type": "org.apache.kafka.connect.transforms.HeaderFrom$Value", "transforms.HeaderFrom.fields": "id,last\_login\_ts", "transforms.HeaderFrom.headers": "user\_id,timestamp", "transforms.HeaderFrom.operation": "move" Before: - Record value: { "id": 11, "name": "Harry Wilson", "last\_login\_ts": 1715242380 } - Record header: { "conv\_id": "uier923" } After: - Record value: { "name": "Harry Wilson" } - Record header: { "conv\_id": "uier923", "user\_id": 11, "timestamp": 1715242380 } ## [](#hoistfield)HoistField HoistField SMT wraps data using the specified field name in a Struct when schema present, or a Map in the case of schemaless data. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.HoistField$Key`) or value (`org.apache.kafka.connect.transforms.HoistField$Value`). ### [](#configuration-8)Configuration | Property key | Description | | --- | --- | | field | Field name for the single field that will be created in the resulting Struct or Map. | ### [](#example-8)Example "transforms": "HoistField", "transforms.HoistField.type": "org.apache.kafka.connect.transforms.HoistField$Value", "transforms.HoistField.field": "name" Message: ```none Red Panda ``` After: ```none {"name":"Red"} {"name":"Panda"} ``` ## [](#insertfield)InsertField InsertField SMT inserts field(s) using attributes from the record metadata or a configured static value. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.InsertField$Key`) or value (`org.apache.kafka.connect.transforms.InsertField$Value`). ### [](#configuration-9)Configuration | Property key | Description | | --- | --- | | offset.field | Field name for Redpanda offset. | | partition.field | Field name for Redpanda partition. | | static.field | Field name for static data field. | | static.value | The static field value. | | timestamp.field | Field name for record timestamp. | | topic.field | Field name for Redpanda topic. | ### [](#example-9)Example Sample configuration: "transforms": "InsertField", "transforms.InsertField.type": "org.apache.kafka.connect.transforms.InsertField$Value", "transforms.InsertField.static.field": "cluster\_id", "transforms.InsertField.static.value": "19423" Before: ```json {"product_id":9987,"price":1234} ``` After: ```json {"price":1234,"cluster_id":"19423","product_id":9987} ``` ## [](#maskfield)MaskField MaskField SMT replaces the contents of fields in a record. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.MaskField$Key`) or value (`org.apache.kafka.connect.transforms.MaskField$Value`). ### [](#configuration-10)Configuration | Property key | Description | | --- | --- | | fields | Comma-separated list of fields to mask. | | replacement | Custom value replacement used to mask field values. | ### [](#example-10)Example "transforms": "MaskField", "transforms.MaskField.type": "org.apache.kafka.connect.transforms.MaskField$Value", "transforms.MaskField.fields": "metadata", "transforms.MaskField.replacement": "\*\*\*" Before: {"product\_id":9987,"price":1234,"metadata":"test"} After: {"metadata":"\*\*\*","price":1234,"product\_id":9987} ## [](#regexrouter)RegexRouter RegexRouter SMT updates the record topic using the configured regular expression and replacement string. Under the hood, the regex is compiled to a `java.util.regex.Pattern`. If the pattern matches the input topic, `java.util.regex.Matcher#replaceFirst()` is used with the replacement string to obtain the new topic. ### [](#configuration-11)Configuration | Property key | Description | | --- | --- | | regex | Regular expression to use for matching. | | replacement | Replacement string. | ### [](#example-11)Example This configuration snippet shows how to add the prefix `prefix_` to the beginning of a topic. "transforms": "AppendPrefix", "transforms.AppendPrefix.type": "org.apache.kafka.connect.transforms.RegexRouter", "transforms.AppendPrefix.regex": ".\*", "transforms.AppendPrefix.replacement": "prefix\_$0" Before: `topic-name` After: `prefix_topic-name` ## [](#replacefield)ReplaceField ReplaceField SMT filters or renames fields in a Redpanda record. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.ReplaceField$Key`) or value (`org.apache.kafka.connect.transforms.ReplaceField$Value`). ### [](#configuration-12)Configuration | Property key | Description | | --- | --- | | exclude | Fields to exclude. This takes precedence over the fields to include. | | include | Fields to include. If specified, only these fields are used. | | renames | List of comma-separated pairs. For example: foo:bar,abc:xyz | ### [](#example-12)Example Sample configuration: "transforms": "ReplaceField", "transforms.ReplaceField.type": "org.apache.kafka.connect.transforms.ReplaceField$Value", "transforms.ReplaceField.renames": "product\_id:item\_number" Before: ```json {"product_id":9987,"price":1234} ``` After: ```json {"item_number":9987,"price":1234} ``` ## [](#replacetimestamp-redpanda)ReplaceTimestamp (Redpanda) ReplaceTimestamp (Redpanda) SMT is designed to support using a record key/value field as a record timestamp, which then can be used to partition data with an S3 connector. Use the concrete transformation type designed for the record key (`com.redpanda.connectors.transforms.ReplaceTimestamp$Key`) or value (`com.redpanda.connectors.transforms.ReplaceTimestamp$Value`). > 📝 **NOTE** > > ReplaceTimestamp is available for Sink connector only. ### [](#configuration-13)Configuration | Property key | Description | | --- | --- | | field | Specifies the name of a field to be used as a source of timestamp. | ### [](#example-13)Example To use `my-timestamp` field as a source of the timestamp for the record, update a connector config with: "transforms": "ReplaceTimestamp", "transforms.ReplaceTimestamp.type": "com.redpanda.connectors.transforms.ReplaceTimestamp$Value", "transforms.ReplaceTimestamp.field": "my-timestamp" for messages in a format: { "name": "my-name", ... "my-timestamp": 1707928150868, ... } The SMT needs structured data to be able to extract the field from it, which means either a Map in the case of schemaless data, or a Struct when a schema is present. The timestamp value should be of a numeric type (epoch millis), or a Java Date object (which is the case when using `"connect.name":"org.apache.kafka.connect.data.Timestamp"` in schema). ## [](#schemaregistryreplicator-redpanda)SchemaRegistryReplicator (Redpanda) SchemaRegistryReplicator (Redpanda) SMT is a transform to replicate schemas. > 📝 **NOTE** > > SchemaRegistryReplicator SMT is designed to be used with the MirrorMaker2 connector only. To use it, remove the `_schema` topic from the topic exclude list. ### [](#example-14)Example Sample configuration: "transforms": "schema-replicator", "transforms.schema-replicator.type": "com.redpanda.connectors.transforms.SchemaRegistryReplicator" ## [](#setschemametadata)SetSchemaMetadata SetSchemaMetadata SMT sets the schema name, version, or both on the record’s key (`org.apache.kafka.connect.transforms.SetSchemaMetadata$Key`) or value (`org.apache.kafka.connect.transforms.SetSchemaMetadata$Value`) schema. ### [](#configuration-14)Configuration | Property key | Description | | --- | --- | | schema.name | Schema name to set. | | schema.version | Schema version to set. | ### [](#example-15)Example Sample configuration: "transforms": "SetSchemaMetadata", "transforms.SetSchemaMetadata.type": "org.apache.kafka.connect.transforms.SetSchemaMetadata$Value", "transforms.SetSchemaMetadata.schema.name": "transaction-value" "transforms.SetSchemaMetadata.schema.version": "3" ## [](#timestampconverter)TimestampConverter TimestampConverter SMT converts timestamps between different formats, such as Unix epoch, strings, and Connect Date/Timestamp types. It applies to individual fields or to the entire value. Use the concrete transformation type designed for the record key (`org.apache.kafka.connect.transforms.TimestampConverter$Key`) or value (`org.apache.kafka.connect.transforms.TimestampConverter$Value`). ### [](#configuration-15)Configuration | Property key | Description | | --- | --- | | field | The field containing the timestamp, or empty if the entire value is a timestamp. Default: "". | | target.type | The desired timestamp representation: string, unix, Date, Time, or Timestamp. | | format | A SimpleDateFormat-compatible format for the timestamp. Used to generate the output when target.type=string or used to parse the input if the input is a string. Default: "". | | unix.precision | The desired Unix precision for the timestamp: seconds, milliseconds, microseconds, or nanoseconds. Used to generate the output when type=unix or used to parse the input if the input is a Long. Note: This SMT causes precision loss during conversions from, and to, values with sub-millisecond components. Default: milliseconds. | ### [](#example-16)Example Sample configuration: "transforms": "TimestampConverter", "transforms.TimestampConverter.type": "org.apache.kafka.connect.transforms.TimestampConverter$Value", "transforms.TimestampConverter.field": "last\_login\_date", "transforms.TimestampConverter.format": "yyyy-MM-dd", "transforms.TimestampConverter.target.type": "string" Before: `1702041416` After: `2023-12-08` ## [](#timestamprouter)TimestampRouter TimestampRouter SMT updates the record’s topic field as a function of the original topic value and the record timestamp. This is mainly useful for sink connectors, because the topic field is often used to determine the equivalent entity name in the destination system (for example, a database table or search index name). > 📝 **NOTE** > > TimestampRouter SMT should be used with sink connectors only. ### [](#configuration-16)Configuration | Property key | Description | | --- | --- | | topic.format | Format string that can contain ${topic} and ${timestamp} as placeholders for the topic and timestamp, respectively. | | timestamp.format | Format string for the timestamp that is compatible with java.text.SimpleDateFormat. | ### [](#example-17)Example Sample configuration: "transforms": "router", "transforms.router.type": "org.apache.kafka.connect.transforms.TimestampRouter", "transforms.router.topic.format": "${topic}\_${timestamp}", "transforms.router.timestamp.format": "YYYY-MM-dd" ## [](#valuetokey)ValueToKey ValueToKey SMT replaces the record key with a new key formed from a subset of fields in the record value. ### [](#configuration-17)Configuration | Property key | Description | | --- | --- | | fields | Comma-separated list of field names on the record value to extract as the record key. | ### [](#example-18)Example Sample configuration: "transforms": "valueToKey", "transforms.valueToKey.type": "org.apache.kafka.connect.transforms.ValueToKey", "transforms.valueToKey.fields": "txn-id" ## [](#error-handling)Error handling By default, `Error tolerance` is set to `NONE`, so SMTs fail for any exception (notably, data parsing or data processing errors). To avoid the connector crashing for data issues, set `Error tolerance` to `ALL`, and specify `Dead Letter Queue Topic Name` as a place where failed messages are redirected. --- # Page 344: Produce Data **URL**: https://docs.redpanda.com/cloud-data-platform/develop/produce-data.md --- # Produce Data > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Produce Data latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: produce-data/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: produce-data/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/produce-data/index.adoc description: Learn how to configure producers and idempotent producers. page-git-created-date: "2024-07-25" page-git-modified-date: "2024-08-01" --- - [Configure Producers](configure-producers/) Learn about configuration options for producers, including write caching and acknowledgment settings. - [Idempotent Producers](idempotent-producers/) Idempotent producers assign a unique ID to every write request, guaranteeing that each message is recorded only once in the order in which it was sent. - [Configure Leader Pinning](leader-pinning/) Learn about Leader Pinning and how to configure a preferred partition leader location based on cloud availability zones or regions. --- # Page 345: Configure Producers **URL**: https://docs.redpanda.com/cloud-data-platform/develop/produce-data/configure-producers.md --- # Configure Producers > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Producers latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: produce-data/configure-producers page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: produce-data/configure-producers.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/produce-data/configure-producers.adoc description: Learn about configuration options for producers, including write caching and acknowledgment settings. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- Producers are client applications that write data to Redpanda in the form of events. Producers communicate with Redpanda through the Kafka API. When a producer publishes a message to a Redpanda cluster, it sends it to a specific partition. Every event consists of a key and value. When selecting which partition to produce to, if the key is blank, then the producer publishes in a round-robin fashion between the topic’s partitions. If a key is provided, then the partition hashes the key using the murmur2 algorithm and modulates across the number of partitions. ## [](#producer-acknowledgment-settings)Producer acknowledgment settings The `acks` property sets the number of acknowledgments the producer requires the leader to have received before considering a request complete. This controls the durability of records that are sent. Redpanda guarantees data safety with fsync, which means flushing to disk. - With `acks=all`, every write is fsynced by default. - With `write.caching` enabled at the topic level, Redpanda fsyncs to disk according to `flush.ms` and `flush.bytes`, whichever is reached first. ### [](#acks0)`acks=0` The producer doesn’t wait for acknowledgments from the leader and doesn’t retry sending messages. This increases throughput and lowers latency of the system at the expense of durability and data loss. This option allows a producer to immediately consider a message acknowledged when it is sent to the Redpanda broker. This means that a producer does not have to wait for any response from the Redpanda broker. This is the least safe option, because a leader-broker crash can cause data loss if the data has not yet replicated to the other brokers in the replica set. However, this setting is useful when you want to optimize for the highest throughput and are willing to risk some data loss. Because of the lack of guarantees, this setting is the most network bandwidth-efficient. This is helpful for use cases like IoT/sensor data collection, where updates are periodic or stateless and you can afford some degree of data loss, but you want to gather as much data as possible in a given time interval. ### [](#acks1)`acks=1` The producer waits for an acknowledgment from the leader, but it doesn’t wait for the leader to get acknowledgments from followers. This setting doesn’t prioritize throughput, latency, or durability. Instead, `acks=1` attempts to provide a balance between all of them. Replication is not guaranteed with this setting because it happens in the background, after the leader broker sends an acknowledgment to the producer. This setting could result in data loss if the leader broker crashes before any followers manage to replicate the message or if a majority of replicas go down at the same time before fsyncing the message to the disk. ### [](#acksall)`acks=all` The producer receives an acknowledgment after the majority of (implicitly, all) replicas acknowledge the message. Redpanda guarantees data safety by fsyncing every message to disk before acknowledgement back to clients. This increases durability at the expense of lower throughput and increased latency. Sometimes referred to as `acks = -1`, this option instructs the broker that replication is considered complete when the message has been replicated (and fsynced) to the majority of the brokers responsible for the partition in the cluster. As soon as the fsync call is complete, the message is considered acknowledged and is made visible to readers. > 📝 **NOTE** > > This property has an important distinction compared to Kafka’s behavior. In Kafka, a message is considered acknowledged without the requirement that it has been fsynced. Messages that have not been fsynced to disk may be lost in the event of a broker crash. So when using `acks=all`, the Redpanda default configuration is more resilient than Kafka’s. You can also consider using write caching, which is a relaxed mode of `acks=all` that acknowledges a message as soon as it is received and acknowledged on a majority of brokers, without waiting for it to fsync to disk. This provides lower latency while still ensuring that a majority of brokers acknowledge the write. ### [](#retries)`retries` This property controls the number of times a message is re-sent to the broker if the broker fails to acknowledge it. This is essentially the same as if the client application resends the erroneous message after receiving an error response. The default value of `retries` in most client libraries is 0. This means that if the send fails, the message is not re-sent at all. If you increase this to a higher value, check the `max.in.flight.requests.per.connection` value as well, because leaving that property at its default value can potentially cause ordering issues in the target topic where the messages arrive. This occurs if two batches are sent to a single partition and the first fails and is retired, but the second succeeds so the records in the second batch may appear first. ### [](#max-in-flight-requests-per-connection)`max.in.flight.requests.per.connection` This property controls how many unacknowledged messages can be sent to the broker simultaneously at any given time. The default value is 5 in most client libraries. If you set this to 1, then the producer does not send any more messages until the previous one is either acknowledged or an error happens, which can prompt a retry. If you set this to a value higher than 1, then the producer sends more messages at the same time, which can help increase throughput but adds a risk of message reordering if retries are enabled. When you configure the producer to be [idempotent](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/idempotent-producers/), up to five requests can be guaranteed to be in flight with the order preserved. ### [](#enable-idempotence)`enable.idempotence` To enable idempotence, set `enable.idempotence` to `true` (the default) in your Redpanda configuration. When idempotence is enabled, the producer ensures that exactly one copy of every message is written to the broker. When set to `false`, the producer retries sending a message for any reason (such as transient errors like brokers not being available or not enough replicas exception), and it can lead to duplicates. In most client libraries `enable.idempotence` is set to true by default. Internally, this is implemented using a special identifier that is assigned to every producer (the producer ID or PID). This ID, along with a sequence number, is included in every message sent to the broker. The broker checks if the PID/sequence number combination is larger than the previous one and, if not, it discards the message. To guarantee true idempotent behavior, you must also set `acks=all` to ensure that all brokers record messages in order, even in the event of node failures. In this configuration, both the producer and the broker prefer safety and durability over throughput. Idempotence is only guaranteed within a session. A session starts after the producer is instantiated and a connection is established between the client and the Redpanda broker. When the connection is closed, the session ends. If your application code retries a request, the producer client assigns a new ID to that request, which may lead to duplicate messages. ## [](#message-batching)Message batching Batching is an efficient way to save on both network bandwidth and disk size, because messages can be compressed easier. When a producer prepares to send messages to a broker, it first fills up a buffer. When this buffer is full, the producer compresses (if instructed to do so) and sends out this batch of messages to the broker. The number of batches that can be sent in a single request to the broker is limited by the `max.request.size` property. The number of requests that can simultaneously be in this sending state is controlled by the `max.in.flight.requests.per.connection` value, which defaults to 5 in most client libraries. Tune the batching configuration with the following properties: ### [](#buffer-memory)`buffer.memory` This property controls the total amount of memory available to the producer for buffering. If messages are sent faster than they can be delivered to the broker, the producer application may run out of memory, which causes it to either block subsequent send calls or throw an exception. The `max.block.ms` property controls the amount of time the producer blocks before throwing an exception if it cannot immediately send messages to the broker. ### [](#batch-size)`batch.size` This property controls the maximum size of coupled messages that can be batched together in one request. The producer automatically puts messages being sent to the same partition into one batch. This configuration property is given in bytes, as opposed to the number of messages. When the producer is gathering messages to assign to a batch, at some point it hits this byte-size limit, which triggers it to send the batch to the broker. However, the producer does not necessarily wait (for as much time as set using `linger.ms`) until the batch is full. Sometimes, it can even send single-message batches. This means that setting the batch size too large is not necessarily undesirable, because it won’t cause throttling when sending messages; rather, it only causes increased memory usage. Conversely, setting the batch size too small can cause the producer to send batches of messages faster, which can cause network overhead, meaning a reduced throughput. The default value is usually 16384, but you can set this as low as 0, which turns off batching entirely. ### [](#linger-ms)`linger.ms` This property controls the maximum amount of time the producer waits before sending out a batch of messages, if it is not already full. This means you can somewhat force the producer to make sure that batches are filled as efficiently as possible. If you’re willing to tolerate some latency, setting this value to a number larger than the default of `0` causes the producer to send fewer, more efficient batches of messages. If you set the value to `0`, there is still a high chance messages arrive around the same time to be batched together. ## [](#common-producer-configurations)Common producer configurations ### [](#compression-type)`compression.type` This property controls how the producer should compress a batch of messages before sending it to the broker. The default is `none`, which means the batch of messages is not compressed at all. Compression occurs on full batches, so you can improve batching throughput by setting this property to use one of the available compression algorithms (along with increasing batch size). The available options are: `zstd`, `lz4`, `gzip`, and `snappy`. ### [](#serializers)Serializers Serializers are responsible for converting a message to a byte array. You can influence the speed/memory efficiency of your streaming setup by choosing one of the built-in serializers or writing a custom one. The performance consequences of using serializers is not typically significant. For example, if you opt for the JSON serializer, you have more data to transport with each message because every record contains its schema in a verbose format, which impacts your compression speeds and network throughput. Alternatively, going with AVRO or Protobuf allows you to only define the schema in one place, while also enabling features like schema evolution. ## [](#broker-timestamps)Broker timestamps Redpanda employs a unique strategy to help ensure the accuracy of retention operations. In this strategy, closed segments are only eligible for deletion when the age of all messages in the segment exceeds a configured threshold. However, when a producer sends a message to a topic, the timestamp set by the producer may not accurately reflect the time the message reaches the broker. To address this time skew, each time a producer sends a message to a topic, Redpanda records the broker’s system date and time in the `broker_timestamp` property of the message. This property helps maintain accurate retention policies, even when the message’s creation timestamp deviates from the broker’s time. > 📝 **NOTE** > > Clock synchronization should be monitored by the server owner, as Redpanda does not monitor clock synchronization. While Redpanda does not rely on clocks for correctness, if you are using `LogAppendTime` (server timestamp set by Redpanda), server clocks may affect the time your application sees. ## [](#producer-optimization-strategies)Producer optimization strategies You can optimize for speed (throughput and latency) or safety (durability and availability) by adjusting properties. Finding the optimal configuration depends on your use case. There are many configuration options within Redpanda. The configuration options mentioned here work best when combined with other broker and consumer configuration options. See also: - [Consumer Offsets](https://docs.redpanda.com/cloud-data-platform/develop/consume-data/consumer-offsets/) ### [](#optimize-for-speed)Optimize for speed To get data into Redpanda as quickly as possible, you can maximize latency and throughput in a variety of ways: - Experiment with [acks](#producer-acknowledgment-settings) settings. The quicker a producer receives a reply from the broker that the message has been committed, the sooner it can send the next message, which generally results in higher throughput. Hence, if you set `acks=1`, then the leader broker does not need to wait for replication to occur, and it can reply as soon as it finishes committing the message. This can result in less durability overall. - Enable [write caching](<#Write caching>), which acknowledges a message as soon as it is received and acknowledged on a majority of brokers, without waiting for it to fsync to disk. This provides lower latency while still ensuring that a majority of brokers acknowledge the write. - Experiment with other component’s properties, like the topic partition size. - Explore how the producer batches messages. Increasing the value of `batch.size` and `linger.ms` can increase throughput by making the producer add more messages into one batch before sending it to the broker and waiting until the batches can properly fill up. This approach negatively impacts latency though. By contrast, if you set `linger.ms` to `0` and `batch.size` to `1`, you can achieve lower latency, but sacrifice throughput. ### [](#optimize-for-safety)Optimize for safety For applications where you must guarantee that there are no lost messages, duplicates, or service downtime, you can use higher durability `acks` settings. If you set `acks=all`, then the producer waits for a majority of replicas to acknowledge the message before it can send the next message, resulting in lower latency, because there is more communication required between brokers. This approach can guarantee higher durability because the message is replicated to all brokers. You can also increase durability by increasing the number of retries the broker can make in case messages are not delivered successfully. The trade-off is that duplicates may enter the system and potentially alter the ordering of messages. --- # Page 346: Idempotent Producers **URL**: https://docs.redpanda.com/cloud-data-platform/develop/produce-data/idempotent-producers.md --- # Idempotent Producers > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Idempotent Producers latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: produce-data/idempotent-producers page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: produce-data/idempotent-producers.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/produce-data/idempotent-producers.adoc description: Idempotent producers assign a unique ID to every write request, guaranteeing that each message is recorded only once in the order in which it was sent. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- When a producer writes messages to a topic, each message should be recorded only once in the order in which it was sent. However, network issues such as a connection failure can result in a timeout, which prevents a write request from succeeding. In such cases, the client retries the write request until one of these events occurs: - The client receives an acknowledgment from the broker that the write was successful. - The retry limit is reached. - The message delivery timeout limit is reached. Since there is no way to tell if the initial write request succeeded before the disruption, a retry can result in a duplicate message. A retry can also cause subsequent messages to be written out of order. Idempotent producers prevent this problem by assigning a unique ID to every write request. The request ID consists of the producer ID and a sequence number. The sequence number identifies the order in which each write request was sent. If a retry results in a duplicate message, Redpanda detects and rejects the duplicate message and maintains the original order of the messages. If new write requests continue while a previous request is being retried, the new requests are stored in the client’s memory in the order in which they were sent. The client must also retry these requests once the previous request is successful. ## [](#enable-idempotence-for-producers)Enable idempotence for producers To make producers idempotent, the `enable.idempotence` property must be set to `true` in your producer configuration, as well as in the Redpanda cluster configuration, where it is set to `true` by default. Some Kafka clients have `enable.idempotence` set to `false` by default. In this case, set the property to `true` by following the instructions for your particular client. Idempotence is guaranteed within a session. A session starts once a producer is created and a connection is established between the client and the Kafka broker. > 📝 **NOTE** > > Idempotent producers retry unsuccessful write requests automatically. If you manually retry a write request, the client will assign a new ID to that request, which may lead to duplicate messages. --- # Page 347: Configure Leader Pinning **URL**: https://docs.redpanda.com/cloud-data-platform/develop/produce-data/leader-pinning.md --- # Configure Leader Pinning > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Leader Pinning latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: produce-data/leader-pinning page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: produce-data/leader-pinning.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/produce-data/leader-pinning.adoc description: Learn about Leader Pinning and how to configure a preferred partition leader location based on cloud availability zones or regions. learning-objective-1: Configure preferred partition leader placement using rack labels learning-objective-2: Configure ordered rack preference for priority-based leader failover learning-objective-3: Identify conditions where Leader Pinning cannot place leaders in preferred racks page-git-created-date: "2024-12-04" page-git-modified-date: "2026-05-26" --- Produce requests that write data to Redpanda topics are routed through the topic partition leader, which syncs messages across its follower replicas. For a Redpanda cluster deployed across multiple availability zones (AZs), Leader Pinning ensures that a topic’s partition leaders are geographically closer to clients, which helps decrease networking costs and guarantees lower latency. If consumers are located in the same preferred region or AZ for Leader Pinning, and you have not set up [follower fetching](https://docs.redpanda.com/cloud-data-platform/develop/consume-data/follower-fetching/), Leader Pinning can also help reduce networking costs on consume requests. After reading this page, you will be able to: - Configure preferred partition leader placement using rack labels - Configure ordered rack preference for priority-based leader failover - Identify conditions where Leader Pinning cannot place leaders in preferred racks ## [](#set-leader-rack-preferences)Set leader rack preferences Configure Leader Pinning if you have Redpanda deployed in a multi-AZ or multi-region cluster and your ingress is concentrated in a particular AZ or region. Use the topic configuration property `redpanda.leaders.preference` to configure Leader Pinning for individual topics. The property accepts the following string values: - `none`: Disable Leader Pinning for the topic. - `racks:[,,…​]`: Specify the preferred location (rack) of all topic partition leaders. The list can contain one or more racks, and you can list the racks in any order. Spaces in the list are ignored, for example: `racks:rack1,rack2` and `racks: rack1, rack2` are equivalent. You cannot specify empty racks, for example: `racks: rack1,,rack2`. If you specify multiple racks, Redpanda tries to distribute the partition leader locations equally across brokers in these racks. - `ordered_racks:[,,…​]`: Supported in Redpanda v26.1 or later. Specify the preferred racks in priority order. Redpanda places leaders in the first listed rack when available, failing over to each subsequent rack when higher-priority racks are unavailable. If all listed racks are unavailable, leaders fall back to any other available brokers. Brokers with no rack assignment are treated as lowest priority. To find the rack identifiers of all brokers, run: ```bash rpk cluster info ``` Expected output ```bash CLUSTER ======= redpanda.be267958-279d-49cd-ae86-98fc7ed2de48 BROKERS ======= ID HOST PORT RACK 0* 54.70.51.189 9092 us-west-2a 1 35.93.178.18 9092 us-west-2b 2 35.91.121.126 9092 us-west-2c ``` To set the topic property: ```bash rpk topic alter-config --set redpanda.leaders.preference=ordered_racks:, ``` If there is more than one broker in the preferred AZ (or AZs), Leader Pinning distributes partition leaders uniformly across brokers in the AZ. ## [](#limitations)Limitations Leader Pinning controls which replica is elected as leader, and does not move replicas to different brokers. If all of a topic’s replicas are on brokers in non-preferred racks, no replica exists in the preferred racks to elect as leader, and Redpanda may elect a non-preferred leader indefinitely. For example, consider a cluster deployed across four racks (A, B, C, D) with Leader Pinning configured as `ordered_racks:A,B,C,D`. With a replication factor of 3, rack awareness can only place replicas in three of the four racks. If the highest-priority rack (A) does not receive a replica, no replica exists there to elect as leader, and Redpanda may elect a non-preferred leader indefinitely. To prevent this scenario, ensure the topic’s replication factor at least equals the total number of racks in the cluster, so every rack, including the highest-priority rack, receives a replica. ## [](#leader-pinning-failover-across-availability-zones)Leader Pinning failover across availability zones If there are three AZs: A, B, and C, and A becomes unavailable, the failover behavior with `racks` is as follows: - The topic with `A` as the preferred leader AZ will have its partition leaders uniformly distributed across B and C. - The topic with `A,B` as the preferred leader AZs will have its partition leaders in B. - The topic with `B` as the preferred leader AZ will have its partition leaders in B as well. ### [](#failover-with-ordered-rack-preference)Failover with ordered rack preference With `ordered_racks`, the failover order follows the configured priority list. Leaders move to the next available rack in the list when higher-priority racks become unavailable. For a topic configured with `ordered_racks:A,B,C`: - The topic with `A` as the first-priority rack will have its partition leaders in A. - If A becomes unavailable, leaders move to B. - If A and B become unavailable, leaders move to C. - If A, B, and C all become unavailable, leaders fall back to any available brokers. If a higher-priority rack recovers and the topic’s replication factor ensures that rack receives a replica, Redpanda automatically moves leaders back to the highest available preferred rack. ## [](#suggested-reading)Suggested reading - For latency-tolerant, high-throughput workloads where cross-AZ networking charges are a major cost driver, also consider [Cloud Topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/cloud-topics/) - [Follower Fetching](https://docs.redpanda.com/cloud-data-platform/develop/consume-data/follower-fetching/) --- # Page 348: Topics **URL**: https://docs.redpanda.com/cloud-data-platform/develop/topics.md --- # Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: topics/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: topics/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/topics/index.adoc description: Overview of standard topics in Redpanda Cloud. page-git-created-date: "2026-03-31" page-git-modified-date: "2026-03-31" --- - [Topics Overview](create-topic/) Learn how to create a topic for a Redpanda Cloud cluster. - [Manage Topics](config-topics/) Learn how to create topics, update topic configurations, and delete topics or records. - [Manage Cloud Topics](cloud-topics/) Cloud Topics are Redpanda topics that enable users to trade off latency for lower costs. --- # Page 349: Manage Cloud Topics **URL**: https://docs.redpanda.com/cloud-data-platform/develop/topics/cloud-topics.md --- # Manage Cloud Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Cloud Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: topics/cloud-topics page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: topics/cloud-topics.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/topics/cloud-topics.adoc description: Cloud Topics are Redpanda topics that enable users to trade off latency for lower costs. page-git-created-date: "2026-03-31" page-git-modified-date: "2026-05-29" --- Starting in v26.1, Redpanda provides [Cloud Topics](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#cloud-topic) to support multi-modal streaming workloads in the most cost-effective way possible: as a per-topic configuration running mixed latency workloads. While standard Redpanda [topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/config-topics/) that use local storage or Tiered Storage are ideal for latency-sensitive workloads (for example, for audit logs or analytics), Cloud Topics are optimized for latency-tolerant, high-throughput workloads where cross-AZ networking charges are a major consideration that can become the dominant cost driver at high throughput. These workloads can include observability streams, offline analytics, AI/ML model training data feeds, or development environments that have flexible latency requirements. Instead of replicating every byte across expensive network links, Cloud Topics leverage durable, inexpensive cloud storage (S3, ADLS, GCS, MinIO) as the primary mechanism to both replicate data and serve it to consumers. This eliminates over 90% of the cost of replicating data over network links in multi-AZ clusters. The end-to-end latency experienced when using Cloud Topics can range from 500 ms to as high as a few seconds with different object stores. Lower latencies may be achievable in certain environments, but Cloud Topics is optimized for throughput rather than low latency or tightly constrained tail latency. This latency profile is often acceptable for many streaming workloads, and can unlock new streaming use cases that previously were not cost effective. With Cloud Topics, data from the client is not acknowledged until it is uploaded to object storage. This maintains durability in the face of infrastructure failures, but results in an increase in both produce latency and end-to-end latency, driven by both batching of produced data and the inherent latency of the underlying object store. You should generally expect end-to-end latencies of 1-2 seconds with public cloud stores. After reading this page, you will be able to: - Describe the latency and cost trade-offs of Cloud Topics compared to standard Redpanda topics - Create a Cloud Topic using rpk on a cluster that has cloud storage enabled - Identify Cloud Topics limitations and configurations that reduce cross-AZ networking costs ## [](#prerequisites)Prerequisites - [Install rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) v26.1 or later. ## [](#limitations)Limitations - In Redpanda versions earlier than v26.2, shadow links do not support Cloud Topics. - Once created, a Cloud Topic cannot be converted back to a standard Redpanda topic that uses local storage or Tiered Storage v1. Conversely, existing topics created as local or Tiered Storage v1 topics cannot be converted to Cloud Topics. Starting in Redpanda v26.2, Cloud Topics can be converted to and from Tiered Storage v2 topics. ## [](#create-cloud-topics)Create Cloud Topics Cloud Topics don’t require a separate cluster property to enable them. When cloud storage is enabled for your cluster, you can create Cloud Topics directly. To create a Cloud Topic, set the topic property `redpanda.storage.mode` to `cloud`: ```bash rpk topic create -c redpanda.storage.mode=cloud ``` ```console TOPIC STATUS audit.analytics.may2025 OK ``` You can make a topic a Cloud Topic only at topic creation time. In addition to replication, cross-AZ ingress (producer) and egress (consumer) traffic can also contribute substantially to cloud networking costs. When running multi-AZ clusters in general, Redpanda strongly recommends using [Follower Fetching](https://docs.redpanda.com/cloud-data-platform/develop/consume-data/follower-fetching/), which allows consumers to avoid crossing network zones. When possible, you can use [leader pinning](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/leader-pinning/), which positions a topic’s partition leader close to the producers, providing a similar benefit for ingress traffic. These features can add additional savings to the replication cost savings of Cloud Topics. For client-side tuning guidance, see [Configure producers for Cloud Topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/configure-producers-for-cloud-topics/). --- # Page 350: Manage Topics **URL**: https://docs.redpanda.com/cloud-data-platform/develop/topics/config-topics.md --- # Manage Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: topics/config-topics page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: topics/config-topics.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/topics/config-topics.adoc description: Learn how to create topics, update topic configurations, and delete topics or records. page-git-created-date: "2026-03-31" page-git-modified-date: "2026-05-26" --- Topics provide a way to organize events in a data streaming platform. ## [](#create-a-topic)Create a topic Creating a topic can be as simple as specifying a name for your topic on the command line. For example, to create a topic named `xyz`, run: ```bash rpk topic create xyz ``` This command creates a topic named `xyz` with one partition and three replicas, because these are the default values set in the cluster configuration file. Replicas are copies of partitions that are distributed across different brokers, so if one broker goes down, other brokers still have a copy of the data. Redpanda Cloud supports 40,000 topics per cluster. ### [](#choose-the-number-of-partitions)Choose the number of partitions A partition acts as a log file where topic data is written. Dividing topics into partitions allows producers to write messages in parallel and consumers to read messages in parallel. The higher the number of partitions, the greater the throughput. > 💡 **TIP** > > As a general rule, select a number of partitions that corresponds to the maximum number of consumers in any consumer group that will consume the data. For example, suppose you plan to create a consumer group with 10 consumers. To create topic `xyz` with 10 partitions, run: ```bash rpk topic create xyz -p 10 ``` ## [](#update-topic-configurations)Update topic configurations After you create a topic, you can update the topic property settings for all new data written to it. For example, you can add partitions or change the cleanup policy. ### [](#add-partitions)Add partitions You can assign a certain number of partitions when you create a topic, and add partitions later. For example, suppose you add brokers to your cluster, and you want to take advantage of the additional processing power. To increase the number of partitions for existing topics, run: ```bash rpk topic add-partitions [TOPICS...] --num [#] ``` Note that `--num <#>` is the number of partitions to _add_, not the total number of partitions. > 📝 **NOTE** > > If a topic already has messages and you add partitions, the existing messages won’t be redistributed to the new partitions. If you require messages to be redistributed, then you must create a new topic with the new partition count, then stream the messages from the old topic to the new topic so they are appropriately distributed according to the new partition hashing. ### [](#reduce-the-number-of-partitions)Reduce the number of partitions You cannot reduce the number of partitions on an existing topic. The Kafka API does not support it: a record’s partition is chosen when the record is produced, and records that have already been written stay in the partition where they landed, so removing a partition would orphan its data. Redpanda assigns a record that has a key to a partition by hashing the key, and leaves a record without a key to the producer’s partitioner, which usually spreads such records across all available partitions. To move a topic’s data to fewer partitions, copy it to a new topic and switch your clients over. Before you start, check that your applications can tolerate the following: - **Duplicates**: Redpanda Connect delivers records at least once, so a restart or a retry during the copy can write the same record to the new topic twice. A lag of zero shows only how far the copy has committed, not that the new topic is free of duplicates. Either make your consumers idempotent, or deduplicate on a record ID after the copy. - **Ordering**: records keep their keys, so all records for a key still land on one partition and keep their order relative to each other, but the global order of records across partitions is not preserved. - **Retention**: the copy starts at the oldest record that is still retained. Records that retention or compaction has already removed cannot be copied, and both keep running during the copy, so complete the copy well within the topic’s retention period. - **Consumer offsets**: consumer group offsets are stored per topic, so the offsets your consumers committed on the original topic do not carry over. Each consumer starts from the beginning of the new topic and replays what the copy wrote, unless you set its offsets explicitly with `rpk group seek`. - **Write downtime**: producers must stop writing to the original topic before you switch clients over, so plan a window in which the topic accepts no writes. The following procedure reduces a topic named `orders` from three partitions to one. 1. Check the configuration of the original topic so that you can recreate it. Note the replication factor, and every row whose `SOURCE` is `DYNAMIC_TOPIC_CONFIG`, which is an override you must set on the new topic: ```bash rpk topic describe orders ``` Example output (abbreviated) ```bash SUMMARY ======= NAME orders PARTITIONS 3 REPLICAS 1 CONFIGS ======= KEY VALUE SOURCE cleanup.policy delete DEFAULT_CONFIG retention.bytes -1 DEFAULT_CONFIG retention.local.target.ms 86400000 DEFAULT_CONFIG retention.ms 604800000 DYNAMIC_TOPIC_CONFIG segment.bytes 134217728 DEFAULT_CONFIG ``` Here, only `retention.ms` is an override. If the topic has Tiered Storage settings, a custom cleanup policy, or other overrides, carry all of them over: a new topic created without them silently falls back to the cluster defaults. 2. Create the new topic with the target number of partitions, the replication factor of the original topic, and each override from the previous step: ```bash rpk topic create orders-reduced --partitions 1 --replicas 1 --topic-config retention.ms=604800000 ``` Example output ```bash TOPIC STATUS orders-reduced OK ``` 3. Copy the data with a [Redpanda Connect](https://docs.redpanda.com/connect/get-started/about/) pipeline. This configuration reads all records that are still available in the original topic and preserves record keys, so records for the same key land on the same partition of the new topic: `reduce-partitions.yaml` ```yaml input: redpanda: seed_brokers: [""] topics: ["orders"] consumer_group: orders-to-orders-reduced start_offset: earliest output: redpanda: seed_brokers: [""] topic: orders-reduced key: ${! @kafka_key } ``` Give the consumer group a name that is unique to this copy, such as `-to-`. `start_offset: earliest` applies only when the group has no committed offset, so a group name that has been used before resumes from where it left off and skips records. ```bash rpk connect run reduce-partitions.yaml ``` Example output ```bash level=info msg="Launching a Redpanda Connect instance, use CTRL+C to close" level=info msg="Output type redpanda is now active" level=info msg="Input type redpanda is now active" ``` 4. Stop the producers that write to the original topic. Leave the pipeline running so that it copies the last records they wrote. 5. Wait for the copy to drain. It is complete when the consumer group reports a `LAG` of `0` for every partition of the original topic: ```bash rpk group describe orders-to-orders-reduced ``` Example output ```bash GROUP orders-to-orders-reduced COORDINATOR-NODE 0 COORDINATOR-PARTITION __consumer_offsets/0 STATE Stable BALANCER cooperative-sticky MEMBERS 1 TOTAL-LAG 0 TOPIC PARTITION CURRENT-OFFSET LOG-START-OFFSET LOG-END-OFFSET LAG MEMBER-ID CLIENT-ID HOST orders 0 3 0 3 0 redpanda-connect-15d7a80f-590f-4cde-bc16-4854fa2754 redpanda-connect 10.0.0.1 orders 1 3 0 3 0 redpanda-connect-15d7a80f-590f-4cde-bc16-4854fa2754 redpanda-connect 10.0.0.1 orders 2 3 0 3 0 redpanda-connect-15d7a80f-590f-4cde-bc16-4854fa2754 redpanda-connect 10.0.0.1 ``` 6. Compare the record counts of the two topics. For each topic, the number of available records is the sum of `HIGH-WATERMARK` minus `LOG-START-OFFSET` across its partitions. A higher count on the new topic means the copy wrote duplicates: ```bash rpk topic describe orders -p rpk topic describe orders-reduced -p ``` Example output ```bash PARTITION LEADER EPOCH REPLICAS LOG-START-OFFSET HIGH-WATERMARK 0 0 1 [0] 0 3 1 0 1 [0] 0 3 2 0 1 [0] 0 3 PARTITION LEADER EPOCH REPLICAS LOG-START-OFFSET HIGH-WATERMARK 0 0 1 [0] 0 9 ``` Nine records across the three original partitions, and the same nine on the single partition of the new topic. 7. Point your producers and consumers at the new topic. Consumers start from the beginning of the new topic unless you set their offsets with `rpk group seek`. 8. Stop the pipeline with Ctrl+C. 9. When you no longer need the original topic, delete it to reclaim storage. See [Delete a topic](#delete-a-topic). > ⚠️ **CAUTION** > > Do not delete the original topic until the new topic holds the data you expect and your consumers are running against it. Deleting a topic deletes its data. ### [](#change-the-cleanup-policy)Change the cleanup policy The cleanup policy determines how to clean up the partition log files when they reach a certain size: - `delete` deletes data based on age or log size. Topics retain all records until then. - `compact` compacts the data by only keeping the latest values for each KEY. - `compact,delete` combines both methods. Unlike compacted topics, which keep only the most recent message for a given key, topics configured with a `delete` cleanup policy provide a running history of all changes for those topics. > ⚠️ **WARNING** > > All topic properties take effect immediately after being set. Do not modify properties on internal Redpanda topics (such as `__consumer_offsets`, `_schemas`, or other system topics) as this can cause cluster instability. For example, to change a topic’s policy to `compact`, run: ```bash rpk topic alter-config [TOPICS…] —-set cleanup.policy=compact ``` ### [](#configure-write-caching)Configure write caching Write caching is a relaxed mode of [`acks=all`](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/configure-producers/#acksall) that provides better performance at the expense of durability. It acknowledges a message as soon as it is received and acknowledged on a majority of brokers, without waiting for it to be written to disk. This provides lower latency while still ensuring that a majority of brokers acknowledge the write. Write caching applies to user topics. It does not apply to transactions or consumer offsets: data written in the context of a transaction and consumer offset commits is always written to disk and fsynced before being acknowledged to the client. Only enable write caching on workloads that can tolerate some data loss in the case of multiple, simultaneous broker failures. Leaving write caching disabled safeguards your data against complete data center or availability zone failures. #### [](#configure-at-topic-level)Configure at topic level To override the cluster-level setting at the topic level, set the topic-level property `write.caching`: `rpk topic alter-config my_topic --set write.caching=true` With `write.caching` enabled at the topic level, Redpanda fsyncs to disk according to `flush.ms` and `flush.bytes`, whichever is reached first. ### [](#remove-a-configuration-setting)Remove a configuration setting You can remove a configuration that overrides the default setting, and the setting will use the default value again. For example, suppose you altered the cleanup policy to use `compact` instead of the default, `delete`. Now you want to return the policy setting to the default. To remove the configuration setting `cleanup.policy=compact`, run `rpk topic alter-config` with the `--delete` flag: ```bash rpk topic alter-config [TOPICS...] --delete cleanup.policy ``` ## [](#list-topic-configuration-settings)List topic configuration settings To display all the configuration settings for a topic, run: ```bash rpk topic describe -c ``` The `-c` flag limits the command output to just the topic configurations. This command is useful for checking the default configuration settings before you make any changes and for verifying changes after you make them. The following command output displays after running `rpk topic describe test-topic`, where `test-topic` was created with default settings: ```bash rpk topic describe test_topic SUMMARY ======= NAME test_topic PARTITIONS 1 REPLICAS 3 CONFIGS ======= KEY VALUE SOURCE cleanup.policy delete DYNAMIC_TOPIC_CONFIG compression.type producer DEFAULT_CONFIG max.message.bytes 20971520 DEFAULT_CONFIG message.timestamp.type CreateTime DEFAULT_CONFIG redpanda.datapolicy function_name: script_name: DEFAULT_CONFIG redpanda.remote.delete true DEFAULT_CONFIG redpanda.remote.read false DEFAULT_CONFIG redpanda.remote.write false DEFAULT_CONFIG retention.bytes -1 DEFAULT_CONFIG retention.local.target.bytes -1 DEFAULT_CONFIG retention.local.target.ms 86400000 DEFAULT_CONFIG retention.ms 604800000 DEFAULT_CONFIG ``` ## [](#delete-a-topic)Delete a topic To delete a topic, run: ```bash rpk topic delete ``` When a topic is deleted, its underlying data is deleted, too. To delete multiple topics at a time, provide a space-separated list. For example, to delete two topics named `topic1` and `topic2`, run: ```bash rpk topic delete topic1 topic2 ``` You can also use the `-r` flag to specify one or more regular expressions; then, any topic names that match the pattern you specify are deleted. For example, to delete topics with names that start with “f” and end with “r”, run: ```bash rpk topic delete -r '^f.*' '.*r$' ``` Note that the first regular expression must start with the `^` symbol, and the last expression must end with the `$` symbol. This requirement helps prevent accidental deletions. ## [](#delete-records-from-a-topic)Delete records from a topic Redpanda allows you to delete data from the beginning of a partition up to a specific offset (a monotonically increasing sequence number for records in a partition). Deleting records frees up disk space, which is especially helpful if your producers are pushing more data than anticipated in your retention plan. Delete records when you know that all consumers have read up to that given offset, and the data is no longer needed. There are different ways to delete records from a topic, including using the [`rpk topic trim-prefix`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-topic/rpk-topic-trim-prefix/) command, using the `DeleteRecords` Kafka API with Kafka clients, or using Redpanda Cloud. > 📝 **NOTE** > > - To delete records, `cleanup.policy` must be set to `delete` or `compact,delete`. > > - Object storage is deleted asynchronously. After messages are deleted, the partition’s start offset will have advanced, but garbage collection of deleted segments may not be complete. > > - Similar to Kafka, after deleting records, local storage and object storage may still contain data for deleted offsets. (Redpanda does not truncate segments. Instead, it bumps the start offset, then it attempts to delete as many whole segments as possible.) Data before the new start offset is not visible to clients but could be read by someone with access to the local disk of a Redpanda node. > ⚠️ **WARNING** > > When you delete records from a topic with a timestamp, Redpanda advances the partition start offset to the first record whose timestamp is after the threshold. If record timestamps are not in order with respect to offsets, this may result in unintended deletion of data. Before using a timestamp, verify that timestamps increase in the same order as offsets in the topic to avoid accidental data loss. For example: > > ```bash > rpk topic consume -n 50 --format '%o %d{go[2006-01-02T15:04:05Z07:00]} %k %v' > ``` ## [](#next-steps)Next steps [Configure Producers](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/configure-producers/) --- # Page 351: Configure Producers for Cloud Topics **URL**: https://docs.redpanda.com/cloud-data-platform/develop/topics/configure-producers-for-cloud-topics.md --- # Configure Producers for Cloud Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Producers for Cloud Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: topics/configure-producers-for-cloud-topics page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: topics/configure-producers-for-cloud-topics.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/topics/configure-producers-for-cloud-topics.adoc description: Learn about producer configuration considerations for Cloud Topics. page-git-created-date: "2026-05-20" page-git-modified-date: "2026-05-26" --- This page describes how to tune the client producer for Cloud Topics (note that general producer configuration guidance still applies). See [Configure Producers](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/configure-producers/). With idempotency enabled, Kafka’s protocol allows only five in-flight requests at a time on a single broker connection. This imposes a tight limit on how much data can be in flight at a single time. Over high-latency links, this is already a problem for standard topics. Cloud Topics use a 250ms batching interval at the broker to reduce cloud storage costs, which effectively makes every producer connection a high-latency link. Compared to standard topics, this changes producer tuning priorities. Configuring batch size, linger time, and request size correctly is critical for achieving good single-producer throughput. This section covers the key settings and example values for the most common client libraries. Calculate the maximum throughput for a single producer as follows: ```text broker_count * max_in_flight * request_size / latency_seconds ``` Throughput formula definitions: - `broker_count`: The number of brokers in the cluster. - `max_in_flight`: The maximum number of in-flight requests on a single connection. With an idempotent producer, this is usually five. - `latency_seconds`: The latency of a single produce request. Assuming normal operation, this can be equated to the 250ms as earlier. In practice, it is typically lower, as the timer starts with the first byte arriving. - `request_size`: The size of a single produce request. This is the primary tuning lever because the other factors are effectively constants. Another way to increase throughput in the system is to increase producer count, as doing so naturally increases available parallelism in the system as more connections are opened to the brokers. ## [](#producer-settings)Producer settings Most Kafka client libraries offer an assortment of tunables, with many properties impacting request size. The most important ones are described here, along with an example for the Kafka Java client, along with caveats from other libraries. ### [](#batches-and-requests)Batches and requests A Kafka produce request consists of one or more batches. A batch contains multiple records. Records mostly correspond to application-level messages as passed to the Kafka client API. At the broker level, the unit that matters is the batch, as everything happens at the batch level. Hence, creating the largest possible batches is the most important factor for performance in general, and equally applies for Cloud Topics. A request can contain multiple batches, this can sometimes alleviate the need for massive batches. However, this requires that a producer is producing to multiple partitions and that there are enough partitions per broker to fill the request with batches. As explained in [Java client](#java-client), there are also exceptions to this in some client libraries, so you should not blindly rely on them. ### [](#java-client)Java client For the Java client, Redpanda Data recommends the following as minimal settings: | Setting | Recommended value | | --- | --- | | linger.ms | 10ms | | batch.size | 131072 | | max.request.size | 1048576 (default) | As shown in the preceding formula, with a basic three-broker setup and enough partitions, you can achieve a maximum throughput of ~65 MB/s per producer. If more throughput is needed on a single producer, increase batch and max request size. Setting `linger.ms` to a higher value is recommended, but it’s less critical for Cloud Topics because as soon as there are five requests in flight, messages are force-batched even after crossing the `linger.ms` threshold. ### [](#librdkafka)librdkafka librdkafka is a commonly-used Kafka C library. However, it’s also the backing library for many other Kafka clients like confluent-kafka-python. librdkafka only allows a single batch in a produce request, unlike most other Kafka client libraries. This significantly cuts down how much data is packed into a single request. Thus, it’s important to increase `batch.size` to even higher values, as it is effectively the limiting factor. librdkafka has a default of 1MB, which typically allows for decent throughput. ### [](#idempotency-considerations)idempotency considerations Disabling idempotency allows you to avoid the in-flight limitation and vastly increases the number of concurrent in-flight requests and throughput. However, running without idempotency can result in duplicate and out-of-order messages, which for most applications is a problem, and is not recommended. In cases where message requirements are fairly lax, it can be a viable alternative. --- # Page 352: Topics Overview **URL**: https://docs.redpanda.com/cloud-data-platform/develop/topics/create-topic.md --- # Topics Overview > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Topics Overview latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: topics/create-topic page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: topics/create-topic.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/topics/create-topic.adoc description: Learn how to create a topic for a Redpanda Cloud cluster. page-git-created-date: "2026-03-31" page-git-modified-date: "2026-07-15" --- Topics provide a way to organize events. After creating a cluster, you can create a topic in it. Each cluster can have up to 40,000 topics. Topic properties are populated from information stored in the broker. Redpanda features, such as Tiered Storage, are enabled and configured by default in Redpanda Cloud. You can optionally overwrite some settings. > ⚠️ **WARNING** > > Modifying the properties of topics that are created and managed by Redpanda applications can cause unexpected errors. This may lead to connector and cluster failures. | Property | Description | | --- | --- | | Partitions | The number of partitions for the topic. | | Replication factor | The number of partition replicas for the topic.Redpanda Cloud requires a minimum of 3 topic replicas. If a topic is created with a replication factor of 1, Redpanda resets the replication factor to 3. | | Cleanup policy | The policy that determines how to clean up old log segments.The default is delete. | | Retention time | The maximum length of time to keep messages in a topic.The default is 7 days. | | Retention size | The maximum size of each partition. If a partition reaches this size and more messages are added, the oldest messages are deleted.The default is infinite. | | Message size | The maximum size of a message or batch for a newly-created topic.The default is 20 MiB for BYOC and Dedicated clusters, and 8 MiB for Serverless clusters. You can increase this value up to 32 MiB for BYOC and Dedicated clusters, and 20 MiB for Serverless clusters, with the max.message.bytes topic property. | | Segment size | The maximum size of a log segment. When a segment reaches this size, Redpanda closes it and starts a new one.Redpanda Cloud sets the segment size for your cluster automatically. You can override it for a topic with the segment.bytes property. | ## [](#next-steps)Next steps - [Manage Topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/config-topics/) - [Manage Cloud Topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/cloud-topics/) --- # Page 353: Transactions **URL**: https://docs.redpanda.com/cloud-data-platform/develop/transactions.md --- # Transactions > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Transactions latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: transactions page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: transactions.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/develop/pages/transactions.adoc description: Learn how to use transactions; for example, you can fetch messages starting from the last consumed offset and transactionally process them one by one, updating the last consumed offset and producing events at the same time. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- Redpanda supports Apache Kafka®-compatible transaction semantics and APIs. For example, you can fetch messages starting from the last consumed offset and transactionally process them one by one, updating the last consumed offset and producing events at the same time. A transaction can span partitions from different topics, and a topic can be deleted while there are active transactions on one or more of its partitions. In-flight transactions can detect deletion events, remove the deleted partitions (and related messages) from the transaction scope, and commit changes to the remaining partitions. If a producer is sending multiple messages to the same or different partitions, and network connectivity or broker failure cause the transaction to fail, then it’s guaranteed that either all messages are written to the partitions or none. This is important for applications that require strict guarantees, like financial services transactions. Transactions guarantee both exactly-once semantics (EOS) and atomicity: - EOS helps developers avoid the anomalies of at-most-once processing (with potential lost events) and at-least-once processing (with potential duplicated events). Redpanda supports EOS when transactions are used in combination with [idempotent producers](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/idempotent-producers/). - Atomicity additionally commits a set of messages across partitions as a unit: either all messages are committed or none. Encapsulated data received or sent across multiple topics in a single operation can only succeed or fail globally. ## [](#use-transactions)Use transactions By default, the `[enable_transactions](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#enable_transactions)` cluster configuration property is set to true. However, in the following use cases, clients must explicitly use the Transactions API to perform operations within a transaction: - [Atomic (all or nothing) publishing of multiple messages](#atomic-publishing-of-multiple-messages) - [Exactly-once stream processing](#exactly-once-stream-processing) When you use transactions, you must set the [`transactional.id`](https://kafka.apache.org/documentation/#producerconfigs_transactional.id) property in the producer configuration. This property uniquely identifies the producer and enables reliable semantics across multiple producer sessions. It ensures that all transactions issued by a given producer are completed before any new transactions are started. ### [](#atomic-publishing-of-multiple-messages)Atomic publishing of multiple messages A banking IT system with an event-sourcing microservice architecture illustrates why transactions are necessary. In this system, each bank branch is implemented as an independent microservice that manages its own distinct set of accounts. Every branch maintains its own transaction history, stored as a Redpanda partition. When a branch starts, it replays the transaction history to reconstruct its current state. Financial transactions such as money transfers require the following guarantees: - A sender can’t withdraw more than the account withdrawal limit. - A recipient receives exactly the same amount sent. - A transaction is fast and is run at most once. - If a transaction fails, the system rolls back to the initial state. - Without withdrawals and deposits, the amount of money in the system remains constant with any history of money transfers. These requirements are easy to satisfy when the sender and the recipient of a financial transaction are hosted by the same branch. The operation doesn’t leave the consistency domain, and all checks and locks can be performed within a single service (ledger). Things get more complex with cross-branch financial transactions, because they involve several ledgers, and the operations should be performed atomically (all or nothing). The default approach (saga pattern) breaks a transaction into a sequence of reversible idempotent steps; however, this violates the isolation principle and adds complexity, making the application responsible for orchestrating the steps. Redpanda natively supports transactions, so it’s possible to atomically update several ledgers at the same time. For example: Show multi-ledger transaction example: ```java Properties props = new Properties(); props.put(ProducerConfig.BOOTSTRAP_SERVERS_CONFIG, "..."); props.put(ProducerConfig.ACKS_CONFIG, "all"); props.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true); props.put(ProducerConfig.TRANSACTIONAL_ID_CONFIG, "app-id"); Producer producer = null; while (true) { // waiting for somebody to initiate a financial transaction var sender_branch = ...; var sender_account = ...; var recipient_branch = ...; var recipient_account = ...; var amount = 42; if (producer == null) { try { producer = new KafkaProducer<>(props); producer.initTransactions(); } catch (Exception e1) { // TIP: log error for further analysis try { if (producer != null) { producer.close(); } } catch(Exception e2) { } producer = null; // TIP: notify the initiator of a transaction about the failure continue; } } producer.beginTransaction(); try { var f1 = producer.send(new ProducerRecord("ledger", sender_branch, sender_account, "" + (-amount))); var f2 = producer.send(new ProducerRecord("ledger", recipient_branch, recipient_account, "" + amount)); f1.get(); f2.get(); } catch (Exception e1) { // TIP: log error for further analysis try { producer.abortTransaction(); } catch (Exception e2) { // TIP: log error for further analysis try { producer.close(); } catch (Exception e3) { } producer = null; } // TIP: notify the initiator of a transaction about the failure continue; } try { producer.commitTransaction(); } catch (Exception e1) { try { producer.close(); } catch (Exception e3) {} producer = null; // TIP: notify the initiator of a transaction about the failure continue; } // TIP: notify the initiator of a transaction about the success } ``` When a transaction fails before a `commitTransaction` attempt completes, you can assume that it is not executed. When a transaction fails after a `commitTransaction` attempt completes, the true transaction status is unknown. Redpanda only guarantees that there isn’t a partial result: either the transaction is committed and complete, or it is fully rolled back. ### [](#exactly-once-stream-processing)Exactly-once stream processing Redpanda is commonly used as a pipe connecting different applications and storage systems. An application could use an OLTP database and then rely on change data capture to deliver the changes to a data warehouse. Redpanda transactions let you use streams as a smart pipe in your applications, building complex atomic operations that transform, aggregate, or otherwise process data transiting between external applications and storage systems. For example, here is the regular pipe flow: Postgresql -> topic -> warehouse Here is the smart pipe flow, with a transformation in `topic(1) -> topic(2)`: Postgresql -> topic(1) transform topic(2) -> warehouse The transformation reads a record from `topic(1)`, processes it, and writes it to `topic(2)`. Without transactions, an intermittent error can cause a message to be lost or processed several times. With transactions, Redpanda guarantees exactly-once semantics. For example: Show exactly-once processing example: ```java var source = "source-topic"; var target = "target-topic"; Properties pprops = new Properties(); pprops.put(ProducerConfig.BOOTSTRAP_SERVERS_CONFIG, "..."); pprops.put(ProducerConfig.ACKS_CONFIG, "all"); pprops.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true); pprops.put(ProducerConfig.TRANSACTIONAL_ID_CONFIG, UUID.randomUUID().toString()); Properties cprops = new Properties(); cprops.put(ConsumerConfig.BOOTSTRAP_SERVERS_CONFIG, "..."); cprops.put(ConsumerConfig.ENABLE_AUTO_COMMIT_CONFIG, false); cprops.put(ConsumerConfig.GROUP_ID_CONFIG, "app-id"); cprops.put(ConsumerConfig.AUTO_OFFSET_RESET_CONFIG, "earliest"); cprops.put(ConsumerConfig.ISOLATION_LEVEL_CONFIG, "read_committed"); Consumer consumer = null; Producer producer = null; boolean should_reset = false; while (true) { if (should_reset) { should_reset = false; if (consumer != null) { try { consumer.close(); } catch(Exception e) {} consumer = null; } if (producer != null) { try { producer.close(); } catch (Exception e2) {} producer = null; } } try { if (consumer == null) { consumer = new KafkaConsumer<>(cprops); consumer.subscribe(Collections.singleton(source)); } } catch (Exception e1) { // TIP: log error for further analysis should_reset = true; continue; } try { if (producer == null) { producer = new KafkaProducer<>(pprops); producer.initTransactions(); } } catch (Exception e1) { // TIP: log error for further analysis should_reset = true; continue; } ConsumerRecords records = null; try { records = consumer.poll(Duration.ofMillis(10000)); } catch (Exception e1) { // TIP: log error for further analysis should_reset = true; continue; } var it = records.iterator(); while (it.hasNext()) { var record = it.next(); // transformation var old_value = record.value(); var new_value = old_value.toUpperCase(); try { producer.beginTransaction(); producer.send(new ProducerRecord(target, record.key(), new_value)); var offsets = new HashMap(); offsets.put(new TopicPartition(source, record.partition()), new OffsetAndMetadata(record.offset() + 1)); producer.sendOffsetsToTransaction(offsets, consumer.groupMetadata()); } catch (Exception e1) { // TIP: log error for further analysis try { producer.abortTransaction(); } catch (Exception e2) { } should_reset = true; break; } try { producer.commitTransaction(); } catch (Exception e1) { // TIP: log error for further analysis should_reset = true; break; } } } ``` #### [](#exactly-once-processing-configuration-requirements)Exactly-once processing configuration requirements Redpanda’s default configuration supports exactly-once processing. To preserve this capability, ensure the following settings are maintained: - `enable_idempotence = true` - `enable_transactions = true` - `transaction_coordinator_delete_retention_ms` is greater than or equal to `transactional_id_expiration_ms` ## [](#best-practices)Best practices To help avoid common pitfalls and optimize performance, consider the following when configuring transactional workloads in Redpanda: ### [](#tune-producer-id-limits)Tune producer ID limits For production environments with heavy producer usage, configure both [`max_concurrent_producer_ids`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#max_concurrent_producer_ids) and [`transactional_id_expiration_ms`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#transactional_id_expiration_ms) to prevent out-of-memory (OOM) crashes. Setting limits on producer IDs helps manage memory usage in high-throughput environments, particularly when using transactions or idempotent producers. If you have\`kafka\_connections\_max\` configured, you can determine an appropriate value for `max_concurrent_producer_ids` based on your connection patterns. - Lower bound: `kafka_connections_max` / `number_of_shards`, assuming each producer connects to only one shard. - Upper bound: `topic_partitions_per_shard` \* `kafka_connections_max`, assuming producers connect to all shards. If `kafka_connections_max` is not configured, estimate the value for `max_concurrent_producer_ids` based on your application patterns. A conservative approach is to start with 1000-5000 per shard, then monitor and adjust as needed. Applications with many partitions per producer typically require higher values, such as 10000 or more per shard. Tune `transactional_id_expiration_ms` based on your application’s transaction patterns. Calculate this value by taking your longest expected transaction time and adding a safety buffer. For example, if transactions typically run for 30 minutes, consider setting this to 2-4 hours. Short-lived transactions can use values between 1-4 hours, while batch processing applications should match their batch interval plus buffer time. Interactive applications may benefit from shorter values to free up memory faster. Client applications should minimize producer ID churn. Reuse producer instances when possible, instead of creating new ones for each operation. Avoid using random transactional IDs, as some Flink configurations do, because this creates excessive producer ID churn. Instead, use consistent transactional IDs that can be resumed across application restarts. ### [](#configure-transaction-timeouts-and-limits)Configure transaction timeouts and limits - If a consumer is configured to use the read\_committed isolation level, it can only process successfully committed transactions. As a result, an ongoing transaction with a large timeout that becomes stuck could prevent the consumer from processing other committed transactions. To avoid this, don’t set the transaction timeout client setting (`transaction.timeout.ms` in the Kafka Java client implementation) to a value that is too high. The longer the timeout, the longer consumers may be blocked. ## [](#handle-transaction-failures)Handle transaction failures Different transactions require different approaches to handling failures within the application. Consider the approaches to failed or timed-out transactions in the provided use cases: - Publishing of multiple messages: The request came from outside the system, and it is the application’s responsibility to discover the true status of a timed-out transaction. (This example doesn’t use consumer groups to distribute partitions between consumers.) - Exactly-once streaming (consume-transform-loop): This is a closed system. Upon re-initialization of the consumer and producer, the system automatically discovers the moment it was interrupted and continues from that place. Additionally, this automatically scales by the number of partitions. Run another instance of the application, and it starts processing its share of partitions in the source topic. ## [](#transactions-with-compacted-segments)Transactions with compacted segments Transactions are supported on topics with compaction configured. The compaction process removes aborted transaction data from the log. The resulting compacted segment contains only committed data batches (and potentially harmless gaps in the offsets due to skipped batches). ## [](#suggested-reading)Suggested reading - [Kafka-compatible fast distributed transactions](https://redpanda.com/blog/fast-transactions) --- # Page 354: Get Started **URL**: https://docs.redpanda.com/cloud-data-platform/get-started.md --- # Get Started > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Get Started latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/index.adoc description: Get Started index page. page-git-created-date: "2024-06-06" page-git-modified-date: "2024-06-07" --- - [What’s New in Redpanda Cloud](whats-new-cloud/) Summary of new features in Redpanda Cloud. - [Redpanda Cloud Overview](cloud-overview/) Learn about Redpanda Cloud deployment options including BYOC, Dedicated, and Serverless clusters. - [BYOC Architecture](byoc-arch/) Learn about the control plane - data plane architecture in BYOC. --- # Page 355: How Redpanda Works **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/architecture.md --- # How Redpanda Works > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: How Redpanda Works latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: architecture page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: architecture.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/architecture.adoc description: Learn specifics about Redpanda architecture. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- At its core, Redpanda is a fault-tolerant transaction log for storing event streams. Producers and consumers interact with Redpanda using the Kafka API. To achieve high scalability, producers and consumers are fully decoupled. Redpanda provides strong guarantees to producers that events are stored durably within the system, and consumers can subscribe to Redpanda and read the events asynchronously. Redpanda achieves this decoupling by organizing events into topics. Topics represent a logical grouping of events that are written to the same log. A topic can have multiple producers writing events to it and multiple consumers reading events from it. This page provides details about how Redpanda works. For a high-level overview, see [Introduction to Redpanda](https://docs.redpanda.com/cloud-data-platform/get-started/intro-to-events/). ## [](#tiered-storage)Tiered Storage Redpanda Tiered Storage is a multi-tiered object storage solution that provides the ability to offload log segments to object storage in near real time. Tiered Storage can be combined with local storage to provide long-term data retention and disaster recovery on a per-topic basis. Consumers that read from more recent offsets continue to read from local storage, and consumers that read from historical offsets read from object storage, all with the same API. Consumers can read and reread events from any point within the maximum retention period, whether the events reside on local or object storage. As data in object storage grows, the metadata for it grows. To support efficient long-term data retention, Redpanda splits the metadata in object storage, maintaining metadata of only recently-updated segments in memory or local disk, while safely archiving the remaining metadata in object storage and caching it locally on disk. Archived metadata is then loaded only when historical data is accessed. This allows Tiered Storage to handle partitions of virtually any size or retention length. ## [](#partitions)Partitions To scale topics, Redpanda shards them into one or more partitions that are distributed across the nodes in a cluster. This allows for concurrent writing and reading from multiple nodes. When producers write to a topic, they route events to one of the topic’s partitions. Events with the same key (like a stock ticker) are always routed to the same partition, and Redpanda guarantees the order of events at the partition level. Consumers read events from a partition in the order that they were written. If a key is not specified, then events are sent to all topic partitions in a round-robin fashion. ## [](#raft-consensus-algorithm)Raft consensus algorithm Redpanda provides strong guarantees for data safety and fault tolerance. Events written to a topic partition are appended to a log file on disk. They can be replicated to other nodes in the cluster and appended to their copies of the log file on disk to prevent data loss in the event of failure. The [Raft consensus algorithm](https://raft.github.io/) is used for data replication. Every topic partition forms a Raft group consisting of a single elected leader and zero or more followers (as specified by the topic’s replication factor). A Raft group can tolerate ƒ failures given 2ƒ+1 nodes. For example, in a cluster with five nodes and a topic with a replication factor of five, the topic remains fully operational if two nodes fail. Raft is a majority vote algorithm. For a leader to acknowledge that an event has been committed to a partition, a majority of its replicas must have written that event to their copy of the log. When a majority (quorum) of responses have been received, the leader can make the event available to consumers and acknowledge receipt of the event when `acks=all (-1)`. [Producer acknowledgement settings](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/configure-producers/#producer-acknowledgement-settings) define how producers and leaders communicate their status while transferring data. As long as the leader and a majority of the replicas are stable, Redpanda can tolerate disturbances in a minority of the replicas. If [gray failures](https://blog.acolyer.org/2017/06/15/gray-failure-the-achilles-heel-of-cloud-scale-systems/) cause a minority of replicas to respond slower than normal, then the leader does not have to wait for their responses to progress, and any additional latency is not passed on to the clients. The result is that Redpanda is less sensitive to faults and can deliver predictable performance. ## [](#partition-leadership-elections)Partition leadership elections [Raft](https://raft.github.io/) uses a heartbeat mechanism to maintain leader authority and to trigger leader elections. The partition leader sends a heartbeat to all followers every 150 milliseconds to assert its leadership in the current term (an election cycle). If a follower does not receive a heartbeat within the election timeout, it triggers an election to choose a new partition leader. Configure the election timeout with the [`election_timeout_ms`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#election_timeout_ms) cluster property (default: 1500 milliseconds). The follower increments its term and votes for itself to be the leader for that term. It then sends a vote request to the other nodes and waits for one of the following scenarios: - It receives a majority of votes and becomes the leader. Raft guarantees that at most one candidate can be elected the leader for a given term. - Another follower establishes itself as the leader. While waiting for votes, the candidate may receive communication from another node in the group claiming to be the leader. The candidate only accepts the claim if its term is greater than or equal to the candidate’s term; otherwise, the communication is rejected and the candidate continues to wait for votes. - No leader is elected over a period of time. If multiple followers timeout and become election candidates at the same time, it’s possible that no candidate gets a majority of votes. When this happens, each candidate increments its term and triggers a new election round. Raft uses a random timeout between 150-300 milliseconds to ensure that split votes are rare and resolved quickly. As long as there is a timing inequality between heartbeat time, election timeout, and mean time between node failures (MTBF), then Raft can elect and maintain a steady leader and make progress. A leader can maintain its position as long as one of the ten heartbeat messages it sends to all of its followers every 1.5 seconds is received; otherwise, a new leader is elected. If a follower triggers an election, but the incumbent leader subsequently springs back to life and starts sending data again, then it’s too late. As part of the election process, the follower (now an election candidate) incremented the term and rejects requests from the previous term, essentially forcing a leadership change. If a cluster is experiencing wider network infrastructure problems that result in latencies above the heartbeat timeout, then back-to-back election rounds can be triggered. During this period, unstable Raft groups may not be able to form a quorum. This results in partitions rejecting writes, but data previously written to disk is not lost. Redpanda has a Raft-priority implementation that allows the system to settle quickly after network outages. ## [](#controller-partition-and-snapshots)Controller partition and snapshots Redpanda stores metadata update commands (such as creating and deleting topics or users) in a system partition called the controller partition. A new snapshot is created after each controller command is added, or, with rapid updates, after a set period of time (default is 60 seconds). Controller snapshots save the current cluster metadata state to disk, so startup is fast. For example, with a partition that has moved several times, a snapshot can restore the latest state without replaying every move command. Each broker has a snapshot file stored in the controller log directory, such as `/var/lib/redpanda/data/redpanda/controller/0_0/snapshot`. The controller partition is replicated by a Raft group that includes all cluster brokers, and the controller snapshot is the Raft snapshot for this group. Snapshots are hydrated when a broker joins the cluster or restarts. Snapshots are enabled by default for all clusters, both new and upgraded. ## [](#optimized-platform-performance)Optimized platform performance Redpanda is designed to exploit advances in modern hardware, from the network down to the disks. Network bandwidth has increased considerably, especially in object storage, and spinning disks have been replaced by SSD devices that deliver better I/O performance. CPUs are faster too, but this is largely due to the increased core counts as opposed to the increase in single-core speeds. Redpanda has tuners that detect your hardware configuration to automatically optimize itself. Examples of platform and kernel features that Redpanda uses to optimize its performance: - Direct Memory Access (DMA) for disk I/O - Sparse file system support with XFS - Distribution of interrupt request (IRQ) processing between CPU cores - Isolated processes with control groups (cgroups) - Disabled CPU power-saving modes - Upfront memory allocation, partitioned and pinned to CPU cores ## [](#tpc)Thread-per-core model Redpanda implements a thread-per-core programming model through its use of the [Seastar](https://seastar.io/) library. This allows Redpanda to pin each of its application threads to a CPU core to avoid context switching and blocking. It combines this with message passing to asynchronously communicate between the pinned threads. With this, Redpanda avoids the overhead of context switching and expensive locking operations to improve processing performance and efficiency. From a sizing perspective, Redpanda’s ability to efficiently use all available hardware enables it to scale up to get the most out of your infrastructure, before you’re forced to scale out to meet the demands of your workload. Redpanda delivers better performance with a smaller footprint, resulting in reduced operational costs and complexity. --- # Page 356: BYOC Architecture **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch.md --- # BYOC Architecture > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: BYOC Architecture latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: byoc-arch page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: byoc-arch.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/byoc-arch.adoc description: Learn about the control plane - data plane architecture in BYOC. page-git-created-date: "2025-04-01" page-git-modified-date: "2026-04-07" --- With Bring Your Own Cloud (BYOC) clusters, you deploy Redpanda in your own cloud (AWS, Azure, or GCP), and all data is contained in your own environment. This provides an additional layer of security and isolation. Redpanda handles provisioning, operations, and maintenance of the underlying infrastructure, including Kubernetes. ## [](#control-plane-data-plane)Control plane - data plane For high availability, Redpanda Cloud uses the following control plane - data plane architecture: ![Control plane and data plane](https://docs.redpanda.com/cloud-data-platform/shared/_images/control_d_plane.png) - **Control plane**: This is a Redpanda Cloud managed service that manages provisioning, operations, and maintenance of clusters with Kubernetes under the hood, including Kubernetes version upgrades and infrastructure maintenance. The control plane enforces rules in the data plane. You can use [RBAC](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/) or [GBAC](https://docs.redpanda.com/cloud-data-platform/security/authorization/gbac/gbac/) in the control plane to manage access to organization-level resources like clusters, resource groups, and networks. - **Data plane**: This is where your cluster lives. The term _data plane_ is sometimes used interchangeably with _cluster_. The data plane is where you manage topics, consumer groups, connectors, and schemas. You can use [RBAC](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac_dp/) or [GBAC](https://docs.redpanda.com/cloud-data-platform/security/authorization/gbac/gbac_dp/) in the data plane to configure cluster-level permissions for provisioned users at scale. IAM permissions allow the Redpanda Cloud agent to access the cloud provider API to create and manage cluster resources. The permissions follow the principle of least privilege, limiting access to only what is necessary. Clusters are configured and maintained in the control plane, but they remain available even if the network connection to the control plane is lost. > 💡 **TIP** > > In the Redpanda Cloud UI, you can identify which plane you’re in by the side navigation: > > - **Control Plane:** Visible after login at the organization level. Here you can select, create, and delete clusters, networks, and resource groups. > > - **Data Plane:** Visible after selecting a specific cluster. Here you can work with topics, consumer groups, connectors, and schemas. ## [](#byoc-setup)BYOC setup In a BYOC architecture, you deploy the data plane in your own VPC. All network connections into the data plane take place through either a public endpoint, or for private clusters, through Redpanda Cloud network connections such as VPC peering, AWS PrivateLink, Azure Private Link, or GCP Private Service Connect. Customer data never leaves the data plane. A BYOC cluster is initially set up from the control plane. This is a two-step process performed by `rpk cloud byoc apply`: 1. You bootstrap a virtual machine (VM) in your VPC. This VM launches the agent and bootstraps the necessary infrastructure. Redpanda then assigns fine-grained IAM policies following least privilege, creating dedicated IAM roles per workload with only the permissions each requires. 2. The agent communicates with the control plane to pull the cluster specifications. After the agent is up and running, it connects to the control plane and starts dequeuing and applying cluster specifications that provision, configure, and maintain clusters. The agent is in constant communication with the control plane, receiving and applying cluster specifications and exchanging cluster metadata. Agents are authenticated and authorized through opaque and ephemeral tokens, and they have dedicated job queues in the control plane. Agents also manage VPC peering networks. ![cloud\_byoc\_apply](https://docs.redpanda.com/cloud-data-platform/shared/_images/byoc_apply.png) > 📝 **NOTE** > > To create a Redpanda cluster in your virtual private cloud (VPC), follow the instructions in the Redpanda Cloud UI. The UI contains the parameters necessary to successfully run `rpk cloud byoc apply` with your cloud provider. > 📝 **NOTE** > > Redpanda Cloud does not support customer access or modifications to any of the internal data plane resources. This restriction allows Redpanda Data to manage all configuration changes internally to ensure a 99.99% service level agreement (SLA) for BYOC clusters. --- # Page 357: Redpanda Cloud Overview **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cloud-overview.md --- # Redpanda Cloud Overview > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Cloud Overview latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cloud-overview page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cloud-overview.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cloud-overview.adoc description: Learn about Redpanda Cloud deployment options including BYOC, Dedicated, and Serverless clusters. page-git-created-date: "2024-06-06" page-git-modified-date: "2026-08-24" --- Redpanda Cloud is a complete data streaming platform delivered as a fully-managed service. It provides automated upgrades and patching, data balancing, and support while continuously monitoring your data to meet strict performance, availability, reliability, and security requirements. All Redpanda Cloud clusters are deployed with an integrated [Redpanda Console](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#redpanda-console), and all clusters have access to unlimited retention and 300+ data connectors with [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#redpanda-connect). ## [](#redpanda-cloud-deployment-options)Redpanda Cloud deployment options Redpanda Cloud applications are supported by three fully-managed deployment options: - **[Serverless](#serverless)**: Fastest way to get started with automatic scaling - **[Dedicated](#dedicated)**: Production clusters in Redpanda’s cloud with enhanced isolation - **[Bring Your Own Cloud (BYOC)](#bring-your-own-cloud-byoc)**: Maximum control and security by deploying in your own cloud environment ### [](#quick-comparison)Quick comparison | | Serverless | Dedicated | BYOC | | --- | --- | --- | --- | | Best for | Starter projects and applications with low or variable traffic | Production clusters requiring cloud hosting, higher throughput, and extra isolation | Production clusters requiring data sovereignty, the highest throughput, and added security | | Deployment | Redpanda’s cloud (AWS/GCP) | Redpanda’s cloud (AWS/Azure/GCP) | Your cloud account (AWS/Azure/GCP) | | Tenancy | Multi-tenant | Single-tenant | Single-tenant | | Cloud SLA | 99.9% | 99.99%, multi-AZ | 99.99%, multi-AZ | | Max throughput (write, read) | Up to 100 MB/s, 300 MB/s | Up to 400 MB/s, 800 MB/s | Up to 2 GB/s, 4 GB/s | | Partitions, pre-replication | Up to 5,000 | Up to 45,600 | Up to 112,500 | | Max message size (MiB) | 8 (default), 20 (max) | 20 (default), 32 (max) | 20 (default), 32 (max) | | Private networking | ✓ | ✓ | ✓ | | SSO authentication | ✓ (GitHub, Google, OIDC) | ✓ (GitHub, Google, OIDC) | ✓ (GitHub, Google, OIDC) | | Redpanda Connect | ✓ | ✓ | ✓ | | Role-based access control (RBAC) & audit logs | ✗ | ✓ | ✓ | | Group-based access control (GBAC) | ✗ | ✓ | ✓ | | Prometheus/OpenMetrics endpoint for cluster metrics | ✓ | ✓ | ✓ | | Multiple availability zones (AZs) | ✗ | ✓ | ✓ | | Cluster properties editing | ✗ | ✓ (AWS/GCP) | ✓ (AWS/GCP) | | Kafka Connect | ✗ | ✓ (disabled by default) | ✓ (disabled by default) | | Redpanda Support | Enterprise support with annual contracts | Enterprise support | Enterprise support for BYOC; Premium support required for BYOVPC/BYOVNet | > 📝 **NOTE** > > - The partition limit is the number of logical partitions before replication occurs. Redpanda Cloud uses a replication factor of three. > > - SSO is configured for your organization and applies to every cluster in it, whatever the cluster type. See [Single sign-on](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#single-sign-on). > > - Enterprise support provides access to streaming experts 24/5, with 24/7 priority escalation for production outages. Premium support provides an enhanced Support SLA. > > - See also: [Serverless vs BYOC/Dedicated](#serverless-vs-byocdedicated) ### [](#serverless)Serverless Serverless is the fastest and easiest way to start data streaming. With Serverless clusters, you host your data in Redpanda’s VPC, and Redpanda handles automatic scaling, provisioning, operations, and maintenance. This is a production-ready deployment option with a cluster available instantly, and you only pay for what you consume. > 📝 **NOTE** > > - Serverless on GCP is currently in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release. #### [](#sign-up-for-serverless)Sign up for Serverless ##### Free trial A [free trial on AWS](https://www.redpanda.com/try-redpanda) is the fastest way to get started with Serverless. Each free-trial customer qualifies for $100 (USD) in credits to spend in the first 30 days. This should be enough to run Redpanda with reasonable throughput. No credit card is required. To continue using Serverless after your trial expires, you can enter a credit card and pay as you go. Any remaining credit balance is used before you are charged. When either the credits expire or the days in the trial expire, the clusters move into a suspended state, and you won’t be able to access your data in either the Redpanda Cloud Console or with the Kafka API. There is a seven-day grace period following the end of the trial when you can add your credit card and restore service. After that, the data is permanently deleted. For questions about the trial, use the **#serverless** [Community Slack](https://redpandacommunity.slack.com/) channel. After you start a trial, Redpanda instantly prepares an account for you. The first time you sign in, you can answer a few quick questions about your project so Redpanda can tailor your experience. Your account includes a `welcome` cluster with a `hello-world` demo topic you can explore. It includes sample data so you can see how real-time messaging works before sending your own data. On that first visit, the **Overview** page shows a **Get started** button with guided ways to [interact with your cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#interact-with-your-cluster): create a Redpanda Connect [pipeline](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#pipeline), use `rpk` from the command line, or connect with your own Kafka client. To get started with `rpk`: 1. Log in with `rpk cloud login`. 2. Consume from the `hello-world` topic with `rpk topic consume hello-world`. 3. In the [Redpanda Cloud Console](https://cloud.redpanda.com), navigate to the **Topics** page and open the `hello-world` topic to see the included messages. ##### Redpanda Sales To request a private offer with possible discounts for annual committed use, contact [Redpanda Sales](https://www.redpanda.com/price-estimator). When you subscribe to Serverless through Redpanda Sales, you gain immediate access to Enterprise support. Redpanda creates a cloud organization for you and sends you a welcome email. ##### AWS Marketplace New subscriptions to Redpanda Cloud through [AWS Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/aws-pay-as-you-go/) receive $300 (USD) in free credits to spend in the first 30 days. AWS Marketplace charges for anything beyond $300, unless you cancel the subscription. After your free credits have been used, you can continue using your cluster without any commitment, only paying for what you consume and canceling anytime. > 📝 **NOTE** > > When you subscribe to Redpanda through AWS Marketplace, you do not have immediate access to Enterprise support, only the [Community Slack](https://redpandacommunity.slack.com/) channel. For Enterprise support, contact [Redpanda Sales](https://www.redpanda.com/price-estimator). Redpanda creates a cloud organization for you and sends you a welcome email. ##### Google Cloud Marketplace New subscriptions to Redpanda Cloud through [Google Cloud Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/gcp-pay-as-you-go/) receive $300 (USD) in free credits to spend in the first 30 days. Google Cloud Marketplace charges for anything beyond $300, unless you cancel the subscription. After your free credits have been used, you can continue using your cluster without any commitment, only paying for what you consume and canceling anytime. > 📝 **NOTE** > > When you subscribe to Redpanda through Google Cloud Marketplace, you do not have immediate access to Enterprise support, only the [Community Slack](https://redpandacommunity.slack.com/) channel. For Enterprise support, contact [Redpanda Sales](https://www.redpanda.com/price-estimator). Redpanda creates a cloud organization for you and sends you a welcome email. ### [](#dedicated)Dedicated With Dedicated clusters, you host your data on Redpanda Cloud resources (AWS, GCP, or Azure), and Redpanda handles provisioning, operations, and maintenance. When you create a Dedicated cluster, you select the supported [tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/dedicated-tiers/) that meets your compute and storage needs. #### [](#sign-up-for-dedicated)Sign up for Dedicated ##### Redpanda Sales To request a private offer with possible discounts for monthly or annual committed use, contact [Redpanda Sales](https://www.redpanda.com/price-estimator). With a usage-based billing commitment, you sign up for a minimum spend amount through [AWS Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/aws-commit/), [Azure Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/azure-commit/), or [Google Cloud Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/gcp-commit/). Redpanda creates a cloud organization for you and sends you a welcome email. You can then provision Dedicated clusters in Redpanda Cloud, and you can view invoices and manage your subscription in the marketplace. ##### AWS Marketplace New subscriptions to Redpanda Cloud through [AWS Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/aws-pay-as-you-go/) receive $300 (USD) in free credits to spend in the first 30 days. AWS Marketplace charges for anything beyond $300, unless you cancel the subscription. After your free credits have been used, you can continue using your cluster without any commitment, only paying for what you consume and canceling anytime. Redpanda creates a cloud organization for you and sends you a welcome email. ### [](#bring-your-own-cloud-byoc)Bring Your Own Cloud (BYOC) With BYOC clusters, you deploy the Redpanda [data plane](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#data-plane) into your existing VPC (for AWS and GCP) or VNet (for Azure), and all data is contained in your own environment. This provides an additional layer of security and isolation. (See [BYOC Architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/).) Redpanda manages provisioning, monitoring, upgrades, and security policies, including the underlying infrastructure and Kubernetes used to run the cluster. Redpanda also manages required resources in your VPC or VNet, including subnets (subnetworks in GCP), IAM roles, and object storage resources (for example, S3 buckets or Azure Storage accounts). For full details, see [Upgrades and Maintenance](https://docs.redpanda.com/cloud-data-platform/manage/maintenance/). #### [](#bring-your-own-vpcvnet-byovpcbyovnet)Bring Your Own VPC/VNet (BYOVPC/BYOVNet) With BYOVPC/BYOVNet clusters, you take full control of the networking lifecycle. Compared to standard BYOC, BYOVPC/BYOVNet provides more security, but the configuration is more complex. See the [shared responsibility model](#shared-responsibility-model) to understand what you manage versus what Redpanda manages. The BYOC infrastructure that Redpanda manages should not be used to deploy any other workloads. For details about the control plane - data plane framework in BYOC, see [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/). #### [](#sign-up-for-byoc)Sign up for BYOC To start using BYOC, contact [Redpanda sales](https://redpanda.com/try-redpanda?section=enterprise-trial) to request a private offer with possible discounts. You are billed directly or through Google Cloud Marketplace or AWS Marketplace. ### [](#serverless-vs-byocdedicated)Serverless vs BYOC/Dedicated Serverless clusters are a good fit for the following use cases: - Quick setup for development or testing - Variable or unpredictable traffic patterns - No upfront cost commitment - Isolated environments for different applications Consider BYOC or Dedicated if you need more control over the deployment or if you have workloads with consistently-high throughput. BYOC and Dedicated clusters offer the following features: - Multiple availability zones (AZs). A multi-AZ cluster provides higher resiliency in the event of a failure in one of the zones. - Role-based access control (RBAC) in the data plane - Group-based access control (GBAC) - Kafka Connect - Higher limits and quotas. See [BYOC usage tiers](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/) and [Dedicated usage tiers](https://docs.redpanda.com/cloud-data-platform/reference/tiers/dedicated-tiers/) compared to [Serverless limits](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#serverless-usage-limits). ## [](#redpanda-cloud-architecture)Redpanda Cloud architecture When you sign up for a Redpanda account, Redpanda creates an organization for you. Your organization contains all your Redpanda resources, including your clusters and networks. Within your organization, Redpanda creates a default resource group to contain your resources. You can rename this resource group, and you can create more resource groups. For example, you may want different resource groups for production and testing. > 💡 **TIP** > > For more detailed information about the Redpanda platform, see [Introduction to Redpanda](https://docs.redpanda.com/cloud-data-platform/get-started/intro-to-events/) and [How Redpanda Works](https://docs.redpanda.com/cloud-data-platform/get-started/architecture/). ## [](#shared-responsibility-model)Shared responsibility model The Redpanda Cloud shared responsibility model lists the security areas owned by Redpanda and the security areas owned by customers. Responsibilities depend on the type of deployment. For a summary of this split alongside the rest of the Redpanda Cloud security model, see [Redpanda Cloud Security Overview](https://docs.redpanda.com/cloud-data-platform/security/cloud-security-overview/). ### BYOC | Resource | Redpanda responsibility | Customer responsibility | | --- | --- | --- | | Redpanda upgrades and hotfixes | ✓ | | | Cost management and attribution | ✓ | ✓ | | Software vulnerability remediation | ✓ | | | Infrastructure vulnerability remediation | ✓ | | | IAM (roles, service accounts, access segmentation) | ✓ | ✓ | | Compute | ✓ | | | Redpanda agent VM maintenance | ✓ | | | VPC (subnets, routing, firewall) | ✓ | ✓ | | VPC peering | | ✓ | | VPC private links (service endpoint) | ✓ | | | VPC private links (consumer endpoint) | | ✓ | | Local storage | ✓ | | | Tiered Storage | ✓ | | | Control plane | ✓ | | | Access controls and audit | ✓ | ✓ | | Managed disaster recovery | | ✓ | | Observability and monitoring (SLOs, SLIs, tracing, alerting, runbooks) | ✓ | | | Availability service-level agreement (SLA) | ✓ (subject to required access to customer resources) | | | Proactive threat detection | ✓ | ✓ | | Static secret rotation | ✓ | | | Incident response | ✓ | | | Resilience verification | ✓ | | | Kafka Connect infrastructure | ✓ | ✓ | | Kafka Connect tasks state | | ✓ | ### BYOVPC/BYOVNet | Resource | Redpanda responsibility | Customer responsibility | | --- | --- | --- | | Redpanda upgrades and hotfixes | ✓ | | | Cost management and attribution | ✓ | ✓ | | Software vulnerability remediation | ✓ | | | Infrastructure vulnerability remediation | ✓ | ✓ | | IAM (roles, service accounts, access segmentation) | | ✓ | | Compute | ✓ | | | Redpanda agent VM maintenance | ✓ | | | VPC (subnets, routing, firewall) | | ✓ | | VPC peering | | ✓ | | VPC private links (service endpoint) | ✓ | | | VPC private links (consumer endpoint) | | ✓ | | Local storage | ✓ | | | Tiered Storage | | ✓ | | Control plane | ✓ | | | Access controls and audit | ✓ | ✓ | | Managed disaster recovery | | ✓ | | Observability and monitoring (SLOs, SLIs, tracing, alerting, runbooks) | ✓ | ✓ (for VPC components and cloud storage buckets/containers managed by customer) | | Availability SLA | ✓ (subject to required access to customer resources) | ✓ | | Proactive threat detection | ✓ | ✓ | | Static secret rotation | ✓ | ✓ | | Incident response | ✓ | | | Resilience verification | ✓ | | | Kafka Connect infrastructure | ✓ | ✓ | | Kafka Connect tasks state | | ✓ | ### Dedicated | Resource | Redpanda responsibility | Customer responsibility | | --- | --- | --- | | Redpanda upgrades and hotfixes | ✓ | | | Cost management and attribution | ✓ | | | Software vulnerability remediation | ✓ | | | Infrastructure vulnerability remediation | ✓ | | | IAM (roles, service accounts, access segmentation) | ✓ | | | Compute | ✓ | | | Redpanda agent VM maintenance | ✓ | | | VPC (subnets, routing, firewall) | ✓ | | | VPC peering | ✓ | | | VPC private links (service endpoint) | ✓ | | | VPC private links (consumer endpoint) | | ✓ | | Local storage | ✓ | | | Tiered Storage | ✓ | | | Control plane | ✓ | | | Access controls and audit | ✓ | | | Managed disaster recovery | | ✓ | | Observability and monitoring (SLOs, SLIs, tracing, alerting, runbooks) | ✓ | | | Availability SLA | ✓ | | | Proactive threat detection | ✓ | | | Static secret rotation | ✓ | | | Incident response | ✓ | | | Resilience verification | ✓ | | | Kafka Connect infrastructure | ✓ | | | Kafka Connect tasks state | | ✓ | ## [](#redpanda-connect-and-kafka-connect)Redpanda Connect and Kafka Connect [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/) lets you compose pipelines from a rich library of inputs, processors, and outputs with strong metrics, logging, and per-pipeline scaling. To try it, see the [quickstart](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/). [Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/) is disabled by default on all new clusters. To unlock this feature for your BYOC or Dedicated cluster, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). When enabled, a Kafka Connect node runs even if no connectors are configured. | | Data transforms | Redpanda Connect | | --- | --- | --- | | Best for | Simple, stateless, per-record normalization inside Redpanda | Enrichment/lookup with external services; multi-stage flows | | External I/O | Not permitted (sandboxed) | Native (HTTP/database/object storage) | | Topology | 1:1 or 1:N (no cross-topic fan-in) | Fan-in and fan-out; multi-step pipelines | | Ordering | Preserves per-partition order | Per-partition order can be preserved; configure parallelism and batching accordingly | | Scale & isolation | Shares broker CPU/memory; best for lightweight operations | Scales independently; isolates heavy work from brokers | | Failure handling | You code routing/error behavior | Built-in retries/backoff and DLQ patterns | > 💡 **TIP** > > - Use data transforms for simple, in-broker, per-record changes with minimal latency. > > - Use Redpanda Connect if your pipeline must talk to external systems (HTTP services, databases, cloud storage), or when you need advanced flow control, such as batching and windowed processing. ### [](#redpanda-connect-vs-data-transforms)Redpanda Connect vs data transforms [Data transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/how-transforms-work/) (Wasm) provide lightweight, per-record changes between Redpanda topics with minimal latency. Transforms run inside the broker, map one input topic to one or more output topics, and are intentionally sandboxed (no external network or disk access). They’re ideal for validation, redaction, format/schema conversion, and simple routing. ## [](#redpanda-cloud-vs-self-managed-feature-compatibility)Redpanda Cloud vs Self-Managed feature compatibility Because Redpanda Cloud is a fully-managed service that provides maintenance, data and partition balancing, upgrades, and recovery, much of the cluster maintenance required for Self-Managed users is not necessary for Redpanda Cloud users. Also, Redpanda Cloud is opinionated about Kafka configurations. For example, automatic topic creation is disabled. Some systems expect the Kafka service to automatically create topics when a message is produced to a topic that doesn’t exist. (You can enable this for BYOC and Dedicated clusters with the `auto_create_topics_enabled` cluster property.) New clusters in Redpanda Cloud generally include functionality added in Self-Managed versions immediately. Existing clusters include new functionality when they get upgraded to the latest version. Redpanda Cloud deployments do not support the following functionality available in Redpanda Self-Managed deployments: - Kafka API OIDC authentication. However, Redpanda Cloud does support [SSO to the Redpanda Cloud UI](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#single-sign-on). - Admin API. - FIPS-compliance mode. - Kerberos authentication. - Redpanda debug bundles. - Redpanda Console topic documentation. - Manual deserialization of Schema Registry - Configuring access to object storage with customer-managed encryption key. - Kubernetes Helm chart and Redpanda Operator functionality. - The following `rpk` commands: - `rpk cluster health` - `rpk cluster license` - `rpk cluster maintenance` - `rpk cluster partitions` - `rpk cluster self-test` - `rpk cluster storage restore` (But `rpk cluster storage` and subcommands for mountable topics are supported in BYOC and Dedicated clusters) - `rpk connect` - `rpk container` - `rpk debug` - `rpk generate app` (This is supported in Serverless clusters only.) - `rpk iotune` - `rpk redpanda` - `rpk topic describe-storage` (All other `rpk topic` commands are supported on both Redpanda Cloud and Self Managed.) > 📝 **NOTE** > > The `rpk cloud` commands are not supported in Self-Managed deployments. ## [](#suggested-videos)Suggested videos - [YouTube - What is Redpanda BYOC? (3 mins)](https://www.youtube.com/watch?v=gVlzsJAYT64&ab_channel=RedpandaData) ## [](#next-steps)Next steps - [Learn about Redpanda Cloud security](https://docs.redpanda.com/cloud-data-platform/security/cloud-security-overview/) - [Learn about upgrades and maintenance](https://docs.redpanda.com/cloud-data-platform/manage/maintenance/) - [Create a Serverless cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) - [Create a BYOC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/) --- # Page 358: Redpanda Cloud Deployment **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types.md --- # Redpanda Cloud Deployment > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Cloud Deployment latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/index.adoc description: Learn about Redpanda Cloud deployments. page-git-created-date: "2024-06-06" page-git-modified-date: "2024-08-01" --- - [Serverless](serverless/) Learn how to create a Serverless cluster and start streaming. - [BYOC](byoc/) Learn how to create a Bring Your Own Cloud (BYOC), Bring Your Own Virtual Private Cloud (BYOVPC), or Bring Your Own Virtual Network (BYOVNet) cluster. - [Dedicated](create-dedicated-cloud-cluster/) Learn how to create a Dedicated cluster and start streaming. --- # Page 359: BYOC **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc.md --- # BYOC > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: BYOC latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/index.adoc description: Learn how to create a Bring Your Own Cloud (BYOC), Bring Your Own Virtual Private Cloud (BYOVPC), or Bring Your Own Virtual Network (BYOVNet) cluster. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-16" --- Bring Your Own Cloud (BYOC) lets you run Redpanda in your own cloud environment while using managed services provided by Redpanda. With BYOC clusters, Redpanda deploys into your existing cloud network: - AWS and GCP: Virtual Private Cloud (VPC) - Azure: Virtual Network (VNet) Your data never leaves your environment, giving you extra security and control. See [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/) for details. Redpanda manages provisioning, monitoring, upgrades, and security policies, and it manages required resources in your VPC or VNet, including subnets (subnetworks in GCP), IAM roles, and object storage resources (for example, S3 buckets or Azure Storage accounts). You get hands-off operations with a 99.99% uptime guarantee while keeping full control of your data. If you want to manage the networking infrastructure yourself, create a Bring Your Own Virtual Private Cloud (BYOVPC) or Bring Your Own Virtual Network (BYOVNet) cluster. With BYOVPC/BYOVNet, the Redpanda agent does not create or change resources in your account. This is ideal for organizations with stringent compliance requirements or existing network configurations, when you need full control over the network lifecycle. Compared to standard BYOC, BYOVPC/BYOVNet provides more security, but the configuration is more complex. See the [shared responsibility model](https://docs.redpanda.com/cloud-data-platform/get-started/cloud-overview/#shared-responsibility-model) to understand what you manage versus what Redpanda manages. > ❗ **IMPORTANT** > > Don’t deploy other workloads on the BYOC infrastructure that Redpanda manages. - [BYOC: AWS](aws/) Learn how to create a BYOC or BYOVPC cluster on AWS. - [BYOC: Azure](azure/) Learn how to create a BYOC or BYOVNet cluster on Azure. - [BYOC: GCP](gcp/) Learn how to create a BYOC or BYOVPC cluster on GCP. - [Create Remote Read Replicas](remote-read-replicas/) Learn how to create a remote read replica topic with BYOC, which is a read-only topic that mirrors a topic on a different cluster. --- # Page 360: BYOC: AWS **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws.md --- # BYOC: AWS > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: "BYOC: AWS" latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/aws/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/aws/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/aws/index.adoc description: Learn how to create a BYOC or BYOVPC cluster on AWS. page-git-created-date: "2024-10-24" page-git-modified-date: "2025-05-07" --- - [Create a BYOC Cluster on AWS](create-byoc-cluster-aws/) Use the Redpanda Cloud UI to create a BYOC cluster on AWS. - [Create a BYOVPC Cluster on AWS](vpc-byo-aws/) Use the Redpanda BYOVPC Terraform module to deploy a BYOVPC cluster on AWS. --- # Page 361: Create a BYOC Cluster on AWS **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/create-byoc-cluster-aws.md --- # Create a BYOC Cluster on AWS > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a BYOC Cluster on AWS latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/aws/create-byoc-cluster-aws page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/aws/create-byoc-cluster-aws.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/aws/create-byoc-cluster-aws.adoc description: Use the Redpanda Cloud UI to create a BYOC cluster on AWS. page-git-created-date: "2024-10-24" page-git-modified-date: "2026-05-21" --- To create a Redpanda cluster in your virtual private cloud (VPC), follow the instructions in the Redpanda Cloud UI. The UI contains the parameters necessary to successfully run `rpk cloud byoc apply`. See also: [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/). > 📝 **NOTE** > > With standard BYOC clusters, Redpanda manages security policies and resources for your VPC, including subnetworks, service accounts, IAM roles, firewall rules, and storage buckets. For the highest level of security, you can manage these resources yourself with a [BYOVPC cluster on AWS](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/vpc-byo-aws/). ## [](#prerequisites)Prerequisites Before you deploy a BYOC cluster on AWS, check that the user creating the cluster has the following prerequisites: - A minimum version of Redpanda `rpk` v24.1. See [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). - The user authenticating to AWS has `AWSAdministratorAccess` access to create the IAM policies specified in [AWS IAM policies](https://docs.redpanda.com/cloud-data-platform/security/authorization/cloud-iam-policies/). - The user has the AWS variables necessary to authenticate. Use either: - `AWS_PROFILE` or - `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` To verify access, you should be able to successfully run `aws sts get-caller-identity` for your region. For more information, see the [AWS CLI reference](https://awscli.amazonaws.com/v2/documentation/api/latest/reference/sts/get-caller-identity.html). ## [](#create-a-byoc-cluster)Create a BYOC cluster 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. On the Clusters page, click **Create cluster**, then click **Create** for BYOC. 3. Enter a cluster name, then select the resource group, provider (AWS), [region, tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/), availability, and Redpanda version. > 📝 **NOTE** > > - If you plan to create a private network in your own VPC, select the region where your VPC is located. > > - Three availability zones provide two backups in case one availability zone goes down. Optionally, click **Advanced settings** to specify up to five key-value custom tags. After the cluster is created, the tags are applied to all AWS resources associated with this cluster. For more information, see the [AWS documentation](https://docs.aws.amazon.com/mediaconnect/latest/ug/tagging-restrictions.html). After the cluster is created, you can [specify more tags with the Cloud API](#manage-custom-tags). 4. Click **Next**. 5. On the Network page, select the connection type: either public or private. For BYOC clusters, private is best-practice. - Your network name is used to identify this network. - For a [CIDR range](https://docs.redpanda.com/cloud-data-platform/networking/cidr-ranges/), choose one that does not overlap with your existing VPCs or your Redpanda network. - Clusters with private networking include a setting for API Gateway network access. Public access exposes endpoints for Redpanda Console, the Data Plane API, and the MCP Server API, but they remain protected by your authentication and authorization controls. Private access restricts endpoint access to your VPC only. > 📝 **NOTE** > > After the cluster is created, you can change the API Gateway access on the Dataplane settings page. If you change from public to private access, users without VPN access to the Redpanda VPC will lose access to these services. > 💡 **TIP** > > To route all cluster egress through your own AWS Transit Gateway and hub VPC instead of a per-VPC NAT Gateway, set the **Transit Gateway ID** field on this page. The field is only available on clusters with a private connection type, and is only visible if centralized egress is enabled for your organization. This option is in beta. See [Configure Centralized Egress with AWS Transit Gateway](https://docs.redpanda.com/cloud-data-platform/networking/byoc/aws/nat-free-egress/). 6. Click **Next**. 7. On the Deploy page, follow the steps to log in to Redpanda Cloud and deploy the agent. As part of agent deployment: - Redpanda assigns the permission required to run the agent. For details about these permissions, see [AWS IAM policies](https://docs.redpanda.com/cloud-data-platform/security/authorization/cloud-iam-policies/). - Redpanda allocates one Elastic IP (EIP) address in AWS for each BYOC cluster. > 📝 **NOTE** > > Redpanda Cloud does not support customer access or modifications to any of the internal data plane resources. This restriction allows Redpanda Data to manage all configuration changes internally to ensure a 99.99% service level agreement (SLA) for BYOC clusters. ## [](#manage-custom-tags)Manage custom tags Your organization might require custom tags for cost allocation, audit compliance, or governance policies. After cluster creation, you can manage tags with the [Cloud Control Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/). The Control Plane API allows up to 16 custom tags in AWS. Make sure you have: - The cluster ID. You can find this in the Redpanda Cloud UI, in the **Details** section of the cluster overview. - A valid bearer token for the Cloud Control Plane API. For details, see [Authenticate to the API](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication). Then complete the following steps: 1. To refresh agent permissions so the Redpanda agent can update tags, run: ```bash export CLUSTER_ID="" rpk cloud byoc aws apply --redpanda-id="$CLUSTER_ID" ``` This step is required because tag management requires additional IAM permissions that may not have been granted during initial cluster creation: - `ec2:DescribeTags` - `ec2:DescribeVolumes` - `ec2:DescribeNetworkInterfaces` - `ec2:CreateTags` - `ec2:DeleteTags` - `iam:TagPolicy` - `iam:UntagPolicy` - `iam:TagInstanceProfile` - `iam:UntagInstanceProfile` 2. To update tags, invoke the Cloud API. First, set your authentication token: ```bash export AUTH_TOKEN="" ``` The `PATCH` call sets the tags specified under `"cloud_provider_tags"`. It replaces the existing tags with the specified tags. Include all desired tags in the request. To remove a single entry, omit it from the map you send. ```bash cluster_patch_body=$(cat <<'JSON' { "cloud_provider_tags": { "Environment": "production", "CostCenter": "engineering" } } JSON ) curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` To remove all tags, send an empty `cloud_provider_tags` object: ```bash cluster_patch_body='{"cloud_provider_tags": {}}' curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` ## [](#next-steps)Next steps [Configure private networking](https://docs.redpanda.com/cloud-data-platform/networking/byoc/aws/) --- # Page 362: Create a BYOVPC Cluster on AWS **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/vpc-byo-aws.md --- # Create a BYOVPC Cluster on AWS > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a BYOVPC Cluster on AWS latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/aws/vpc-byo-aws page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/aws/vpc-byo-aws.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/aws/vpc-byo-aws.adoc description: Use the Redpanda BYOVPC Terraform module to deploy a BYOVPC cluster on AWS. page-topic-type: how-to personas: platform_admin learning-objective-1: Deploy a BYOVPC cluster on AWS using the Redpanda Terraform module learning-objective-2: Configure the Redpanda network and cluster resources using module outputs learning-objective-3: Enable PrivateLink on a BYOVPC cluster page-git-created-date: "2024-12-02" page-git-modified-date: "2026-08-27" --- > ❗ **IMPORTANT** > > BYOVPC/BYOVNet is an add-on feature that requires Premium support. To unlock this feature for your account, contact your Redpanda account team or [Redpanda Sales](https://www.redpanda.com/price-estimator). A Bring Your Own Virtual Private Cloud (BYOVPC) cluster allows you to deploy the Redpanda [data plane](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#data-plane) into your existing VPC and manage the networking lifecycle yourself. Compared to a standard Bring Your Own Cloud (BYOC) setup, where Redpanda manages the networking lifecycle for you, BYOVPC provides more control. For background on the architecture, see [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/). When you create a BYOVPC cluster, you specify your VPC and the IAM role (instance profile) that the Redpanda agent will assume. The Redpanda Cloud agent doesn’t create any new resources or alter any settings in your account. With BYOVPC: - You provide your own VPC in your AWS account. - You maintain more control over your account, because Redpanda requires fewer permissions than standard BYOC clusters. - You control your security resources and policies, including subnets, service accounts, IAM roles, firewall rules, and storage buckets. The [Redpanda BYOVPC Terraform Module](https://registry.terraform.io/modules/redpanda-data/redpanda-byovpc/aws/latest) contains [Terraform](https://developer.hashicorp.com/terraform) code that deploys the resources required for a BYOVPC cluster on AWS. You need to create these resources in advance and provide them to Redpanda during cluster creation. Variables are provided in the code so you can exclude resources that already exist in your environment, such as the VPC. > 📝 **NOTE** > > Secrets management is enabled by default with the Terraform module. It allows you to store and read secrets in your cluster, for example to integrate a REST catalog with Iceberg-enabled topics. > > For existing BYOVPC clusters, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new) to enable secrets management. ## [](#prerequisites)Prerequisites - Access to an AWS account in which you create your cluster. - Minimum permissions in that AWS account. For the actions required by the user who will create the cluster with `terraform apply`, see [`iam_rpk_user.tf`](https://github.com/redpanda-data/terraform-aws-redpanda-byovpc/blob/main/iam_rpk_user.tf). - Each BYOVPC cluster requires one allocated Elastic IP (EIP) address in AWS. - [Terraform](https://developer.hashicorp.com/terraform/tutorials/aws-get-started/install-cli) version 1.8.5 or later. - The [Redpanda Terraform provider](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs) configured with valid credentials. For setup details, see the provider documentation. ## [](#limitations)Limitations - Existing clusters cannot be converted to BYOVPC clusters. - After creating a BYOVPC cluster, you cannot change to a different VPC. - Only primary CIDR ranges are supported for the VPC. > 📝 **NOTE** > > For simplicity, the instructions are based on the assumption that Terraform is configured to use local state. You may want to configure [remote state](https://developer.hashicorp.com/terraform/language/state/remote). ## [](#configure-the-redpanda-byovpc-terraform-module)Configure the Redpanda BYOVPC Terraform module The following example uses the Redpanda BYOVPC Terraform Module to create the resources required to create a BYOVPC cluster. > 📝 **NOTE** > > Redpanda recommends using a VPC in AWS with a CIDR block (10.0.0.0/16) to allow for enough address space. The subnets must be set to /24. ```hcl locals { common_prefix = "abc-stg" region = "us-east-2" zones = ["use2-az1", "use2-az2", "use2-az3"] enable_private_link = false # see example below for enabling private link force_destroy_cloud_storage = false # see example below if using pre-existing VPC and subnets, # otherwise when provided with these cidrs the module will # attempt to create the VPC and subnets vpc_cidr_block = "10.0.0.0/16" public_subnet_cidrs = [ "10.0.1.0/24", "10.0.3.0/24", "10.0.5.0/24", "10.0.7.0/24", "10.0.9.0/24", "10.0.11.0/24" ] private_subnet_cidrs = [ "10.0.0.0/24", "10.0.2.0/24", "10.0.4.0/24", "10.0.6.0/24", "10.0.8.0/24", "10.0.10.0/24" ] # condition_tags restrict the IAM permissions granted by the # module to only those resources with these tags, when using # condition_tags these tags must also be provided to the # redpanda_cluster so that all resources created are given # these tags condition_tags = { "redpanda-managed" : "true" } # default_tags are applied to all resources created by the # module or redpanda_cluster resource default_tags = { "env" : "staging" } # when using a brand new AWS account that has never hosted an # EKS cluster before the EKS node group service linked role # must be created, if it already exists this may be set to false create_eks_nodegroup_service_linked_role = true } module "redpanda_byovpc" { source = "redpanda-data/redpanda-byovpc/aws" common_prefix = local.common_prefix region = local.region zones = local.zones create_rpk_user = false enable_private_link = local.enable_private_link force_destroy_cloud_storage = local.force_destroy_cloud_storage enable_redpanda_connect = true vpc_cidr_block = local.vpc_cidr_block private_subnet_cidrs = local.private_subnet_cidrs public_subnet_cidrs = local.public_subnet_cidrs condition_tags = local.condition_tags default_tags = local.default_tags create_eks_nodegroup_service_linked_role = local.create_eks_nodegroup_service_linked_role } ``` > 📝 **NOTE** > > - To send telemetry back to the Redpanda control plane, the cluster needs outbound internet access. You can provide this through at least one public subnet, or through network peering or a transit gateway to another VPC that routes traffic through a public subnet. The example configuration includes multiple public subnets to allow for future scaling. Standard BYOC clusters can also route egress through a customer-owned hub VPC and Transit Gateway, eliminating the per-VPC NAT Gateway entirely. See [Configure Centralized Egress with AWS Transit Gateway](https://docs.redpanda.com/cloud-data-platform/networking/byoc/aws/nat-free-egress/). > > - The example creates an Internet Gateway and an associated Route Table rule that routes traffic into the VPC, which allows the Redpanda control plane to access the cluster. To disable creation of the Internet Gateway, either remove the configuration and value for `create_internet_gateway` or set `"create_internet_gateway": false`. > > - When using a pre-existing VPC, at least one public subnet must already exist in that VPC. Setting `public_subnet_cidrs = []` only prevents the module from creating new ones. > 💡 **TIP** > > See the full list of zones and tiers available with each provider in the [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-regions-and-usage-tiers). ## [](#configure-the-redpanda-network-and-cluster)Configure the Redpanda network and cluster After provisioning the AWS infrastructure, configure the Redpanda network and cluster resources using the module outputs. ```hcl locals { resource_group_name = "staging" throughput_tier = "tier-1-aws-v3-arm" } data "redpanda_resource_group" "staging" { name = local.resource_group_name } resource "redpanda_network" "network" { name = "${local.common_prefix}-network" resource_group_id = data.redpanda_resource_group.staging.id cloud_provider = "aws" region = local.region cluster_type = "byoc" customer_managed_resources = { aws = { management_bucket = { arn = module.redpanda_byovpc.management_bucket_arn } dynamodb_table = { arn = module.redpanda_byovpc.dynamodb_table_arn } vpc = { arn = module.redpanda_byovpc.vpc_arn } private_subnets = { arns = module.redpanda_byovpc.private_subnet_arns } } } depends_on = [module.redpanda_byovpc] } resource "redpanda_cluster" "cluster" { name = "${local.common_prefix}-cluster" resource_group_id = data.redpanda_resource_group.staging.id cloud_provider = "aws" region = redpanda_network.network.region zones = local.zones network_id = redpanda_network.network.id cluster_type = "byoc" connection_type = "private" throughput_tier = local.throughput_tier allow_deletion = false tags = merge(local.condition_tags, local.default_tags) customer_managed_resources = { aws = { agent_instance_profile = { arn = module.redpanda_byovpc.agent_instance_profile_arn } cloud_storage_bucket = { arn = module.redpanda_byovpc.cloud_storage_bucket_arn } cluster_security_group = { arn = module.redpanda_byovpc.cluster_security_group_arn } connectors_node_group_instance_profile = { arn = module.redpanda_byovpc.connectors_node_group_instance_profile_arn } connectors_security_group = { arn = module.redpanda_byovpc.connectors_security_group_arn } k8s_cluster_role = { arn = module.redpanda_byovpc.k8s_cluster_role_arn } node_security_group = { arn = module.redpanda_byovpc.node_security_group_arn } permissions_boundary_policy = { arn = module.redpanda_byovpc.permissions_boundary_policy_arn } redpanda_agent_security_group = { arn = module.redpanda_byovpc.redpanda_agent_security_group_arn } redpanda_node_group_instance_profile = { arn = module.redpanda_byovpc.redpanda_node_group_instance_profile_arn } redpanda_node_group_security_group = { arn = module.redpanda_byovpc.redpanda_node_group_security_group_arn } utility_node_group_instance_profile = { arn = module.redpanda_byovpc.utility_node_group_instance_profile_arn } utility_security_group = { arn = module.redpanda_byovpc.utility_security_group_arn } redpanda_connect_node_group_instance_profile = { arn = module.redpanda_byovpc.redpanda_connect_node_group_instance_profile_arn } redpanda_connect_security_group = { arn = module.redpanda_byovpc.redpanda_connect_security_group_arn } } } depends_on = [redpanda_network.network] } ``` ## [](#apply-the-terraform-configuration)Apply the Terraform configuration Initialize, plan, and apply Terraform to set up the AWS infrastructure: ```bash terraform init && terraform plan && terraform apply ``` Cluster provisioning can take up to 45 minutes. When provisioning completes, the cluster status updates to `Running`. If the cluster stays in `Creating` status, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). ### [](#validation-checks)Validation checks The `redpanda_cluster` resource performs validation checks before proceeding with provisioning: - RPK user: Checks if the user running the command has sufficient privileges to provision the agent. Any missing permissions are displayed in the output. - IAM instance profile: Checks that `agent_instance_profile`, `connectors_node_group_instance_profile`, `redpanda_node_group_instance_profile`, `redpanda_connect_node_group_instance_profile`, `utility_node_group_instance_profile`, and `k8s_cluster_role` have the minimum required permissions. Any missing permissions are displayed in the output. - Storage: Checks that the `management_bucket` exists and is versioned, checks that the `cloud_storage_bucket` exists and is not versioned, and checks that the `dynamodb_table` exists. - Network: Checks that the VPC exists, checks that the subnets exist and have the expected tags, and checks that the security groups exist and have the desired ingress and egress rules. ## [](#delete-the-cluster)Delete the cluster To delete the cluster and all associated resources, run `terraform destroy`. > ⚠️ **WARNING** > > This also deletes the customer-managed resources created by the module. ```bash terraform destroy ``` ## [](#enable-privatelink)Enable PrivateLink PrivateLink can be enabled during cluster creation or on an already existing cluster. Start by enabling PrivateLink in the Redpanda BYOVPC Terraform module. This adds the permissions required for PrivateLink. ```hcl module "redpanda_byovpc" { # ... enable_private_link = true # ... } ``` Enable PrivateLink on the `redpanda_cluster` resource: ```hcl resource "redpanda_cluster" "cluster" { # ... aws_private_link = { allowed_principals = ["arn:aws:iam::${var.aws_account_id}:root"] enabled = true connect_console = false } # ... } ``` ## [](#deploy-with-pre-existing-vpc-and-subnets)Deploy with pre-existing VPC and subnets If you already have a VPC and subnets in your AWS account, provide their IDs to the module instead of CIDR blocks. ```hcl module "redpanda_byovpc" { # ... # vpc_cidr_block = local.vpc_cidr_block # private_subnet_cidrs = local.private_subnet_cidrs # public_subnet_cidrs = local.public_subnet_cidrs vpc_id = "vpc-0c79b236047faa1ab" private_subnet_ids = [ "subnet-0e58df59b5eb037c3", "subnet-0c74559ab372f5123", "subnet-0525df35c467cad1c", "subnet-09c301e004e96c803", "subnet-0f67e76738572cb8e", "subnet-0cca6892cf789f6ec", ] public_subnet_cidrs = [] # when empty the module will not create any public subnets # ... } ``` ## [](#next-steps)Next steps - [Configure AWS PrivateLink](https://docs.redpanda.com/cloud-data-platform/networking/aws-privatelink/) - [Learn about `rpk` commands](https://docs.redpanda.com/cloud-data-platform/reference/rpk/) - [Enable Redpanda SQL on a BYOVPC Cluster on AWS](https://docs.redpanda.com/cloud-data-platform/sql/get-started/enable-sql-byovpc-aws/) --- # Page 363: BYOC: Azure **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure.md --- # BYOC: Azure > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: "BYOC: Azure" latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/azure/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/azure/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/azure/index.adoc description: Learn how to create a BYOC or BYOVNet cluster on Azure. page-git-created-date: "2024-10-24" page-git-modified-date: "2025-07-30" --- - [Create a BYOC Cluster on Azure](create-byoc-cluster-azure/) Use the Redpanda Cloud UI to create a BYOC cluster on Azure. - [Create a BYOVNet Cluster on Azure](vnet-azure/) Use Terraform to deploy a BYOVNet cluster on Azure. --- # Page 364: Create a BYOC Cluster on Azure **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/create-byoc-cluster-azure.md --- # Create a BYOC Cluster on Azure > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a BYOC Cluster on Azure latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/azure/create-byoc-cluster-azure page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/azure/create-byoc-cluster-azure.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/azure/create-byoc-cluster-azure.adoc description: Use the Redpanda Cloud UI to create a BYOC cluster on Azure. page-git-created-date: "2024-10-24" page-git-modified-date: "2026-07-30" --- To create a Redpanda cluster in your virtual network (VNet), follow the instructions in the Redpanda Cloud UI. The UI contains the parameters necessary to successfully run `rpk cloud byoc apply`. See also: [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/). > 📝 **NOTE** > > With standard BYOC clusters, Redpanda manages security policies and resources for your virtual network (VNet), including subnetworks, managed identities, IAM roles, security groups, and storage accounts. For the most security, you can manage these resources yourself with a [BYOVNet cluster on Azure](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/vnet-azure/). ## [](#prerequisites)Prerequisites Before you deploy a BYOC cluster on Azure, check all prerequisites to ensure that your Azure subscription meets requirements. ### [](#configure-azure-cli)Configure Azure CLI - [Install the Azure CLI](https://learn.microsoft.com/en-us/cli/azure/install-azure-cli). - [Sign in](https://learn.microsoft.com/en-us/cli/azure/authenticate-azure-cli) with the Azure CLI: ```none az login ``` - Set the desired subscription for the Azure CLI: ```none az account set --subscription ``` ### [](#verify-rpk-version)Verify rpk version Confirm you have a minimum version of Redpanda `rpk` v24.1. See [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). ### [](#prepare-your-azure-subscription)Prepare your Azure subscription In the [Azure Portal](https://login.microsoftonline.com/), confirm that the dedicated subscription you intend to use with Redpanda includes the following: - **Role**: The Azure user must have the _Owner_ role in the subscription. - **Resources**: The subscription must be registered for the following resource providers (AKS + common dependencies). See the [Microsoft documentation](https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/resource-providers-and-types). - Microsoft.Compute - Microsoft.ManagedIdentity - Microsoft.Storage - Microsoft.KeyVault - Microsoft.Network - Microsoft.ContainerService To check if a resource provider is registered, run the following command using the Azure CLI or in the Azure Cloud Shell: ```none az provider show -n Microsoft.Compute --query registrationState -o tsv az provider show -n Microsoft.ManagedIdentity --query registrationState -o tsv az provider show -n Microsoft.Storage --query registrationState -o tsv az provider show -n Microsoft.KeyVault --query registrationState -o tsv az provider show -n Microsoft.Network --query registrationState -o tsv az provider show -n Microsoft.ContainerService --query registrationState -o tsv ``` If a resource provider is not registered, run: ```none az provider register --namespace Microsoft.Compute az provider register --namespace Microsoft.ManagedIdentity az provider register --namespace Microsoft.Storage az provider register --namespace Microsoft.KeyVault az provider register --namespace Microsoft.Network az provider register --namespace Microsoft.ContainerService ``` - **Feature**: The subscription must be registered for Microsoft.Compute/EncryptionAtHost. See the [Microsoft documentation](https://learn.microsoft.com/en-us/azure/virtual-machines/linux/disks-enable-host-based-encryption-cli#prerequisites). To register it, run: ```none az feature register --namespace Microsoft.Compute --name EncryptionAtHost # (optional) Wait and verify it shows as Registered az feature show --namespace Microsoft.Compute --name EncryptionAtHost --query properties.state -o tsv # Refresh the provider after enabling a feature az provider register --namespace Microsoft.Compute ``` - **Monitoring**: The subscription must have Azure Network Watcher enabled in the NetworkWatcherRG resource group and the region where you will use Redpanda. Network Watcher lets you monitor and diagnose conditions at a network level. See the [Microsoft documentation](https://learn.microsoft.com/en-us/azure/network-watcher/network-watcher-create?tabs=portaly). To enable it, run: ```none # Create the NetworkWatcherRG resource group az group create --name 'NetworkWatcherRG' --location '' # Enable Network Watcher in az network watcher configure --resource-group 'NetworkWatcherRG' --locations '' --enabled ``` ### [](#check-azure-quota)Check Azure quota Confirm that the Azure subscription has enough virtual CPUs (vCPUs) per instance family and total regional vCPUs in the region where you will use Redpanda: - Standard Ddv5-series vCPUs: 12 (3 Redpanda broker nodes + extra capacity for 3 more nodes that could be utilized temporarily during tier 1 maintenance) - Standard Dadsv5-series vCPUs: 8 (2 Redpanda utility nodes) - Standard Dv3-series vCPUs: 2 (1 Redpanda agent node) See the [Microsoft documentation](https://learn.microsoft.com/en-us/azure/quotas/view-quotas). ### [](#check-azure-sku-restrictions)Check Azure SKU restrictions Ensure your subscription has access to the required VM sizes in the region where you will use Redpanda. For example, using the Azure CLI or in the Azure Cloud Shell, run: ```bash # Replace eastus2 with your target region az vm list-skus -l eastus2 --zone --size Standard_D2d_v5 --output table ``` Example output (no restrictions: good) ```bash ResourceType Locations Name Zones Restrictions --------------- ----------- --------------- ------- ------------ virtualMachines eastus2 Standard_D2d_v5 1,2,3 None ``` Example output (with restrictions: needs attention) ```bash ResourceType Locations Name Zones Restrictions --------------- ----------- --------------- ------- ------------ virtualMachines eastus2 Standard_D2d_v5 1,2,3 NotAvailableForSubscription ``` If you see restrictions, [open a Microsoft support request](https://learn.microsoft.com/en-us/troubleshoot/azure/general/region-access-request-process) to remove them. ### [](#prerequisite-checklist)Prerequisite checklist - Verified `rpk` version - Verified Azure user has Owner role - Registered all required resource providers - Registered EncryptionAtHost feature - Enabled Network Watcher - Verified vCPU quota - Verified no SKU restrictions ## [](#create-a-byoc-cluster)Create a BYOC cluster To create a Redpanda cluster in your Azure VNet, follow the [prerequisites](#prerequisites) then follow the instructions in the Redpanda Cloud UI. The UI contains the parameters necessary to successfully run `rpk cloud byoc apply`. 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. On the Clusters page, click **Create cluster**, then click **Create** for BYOC. 3. Enter a cluster name, then select the resource group, provider (Azure), [region, tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/), availability, and Redpanda version. > 📝 **NOTE** > > - If you plan to create a private network in your own VNet, select the region where your VNet is located. > > - Multi-AZ is the default configuration. Three AZs provide two backups in case one availability zone goes down. Optionally, click **Advanced settings** to specify up to five key-value custom tags. After the cluster is created, the tags are applied to all Azure resources associated with this cluster. For details, see the [Microsoft documentation](https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/tag-resources). After the cluster is created, you can [specify more tags with the Cloud API](#manage-custom-tags). 4. Click **Next**. 5. On the Network page, select the connection type: either public or private. For BYOC clusters, private using Azure Private Link is best-practice. - Your network name is used to identify this network. - For a [CIDR range](https://docs.redpanda.com/cloud-data-platform/networking/cidr-ranges/), choose one that does not overlap with your existing VPCs or your Redpanda network. - Clusters with private networking include a setting for API Gateway network access. Public access exposes endpoints for Redpanda Console, the Data Plane API, and the MCP Server API, but they remain protected by your authentication and authorization controls. Private access restricts endpoint access to your VNet only. Private access incurs an additional cost, since it involves deploying two network load balancers, instead of one. > 📝 **NOTE** > > After the cluster is created, you can change the API Gateway access on the Dataplane settings page. If you change from public to private access, users without VPN access to the Redpanda VPC will lose access to these services. > 💡 **TIP** > > To route all cluster egress through your own Azure Firewall and hub VNet instead of a per-cluster NAT Gateway, enter the **Hub egress VNet ID** and **Firewall private IP** on this page, or set `egress_spec.azure.hub_vnet_id` and `egress_spec.azure.firewall_private_ip` when you create the network with the Cloud API. This option is only available on clusters with a private connection type and a Redpanda-managed VNet, and only if centralized egress is enabled for your organization. This option is in beta. See [Configure Centralized Egress with Azure Firewall](https://docs.redpanda.com/cloud-data-platform/networking/byoc/azure/nat-free-egress/). 6. Click **Next**. 7. On the Deploy page, follow the steps to log in to Redpanda Cloud and deploy the agent. As part of agent deployment, Redpanda assigns the permissions required to run the agent. For details about these permissions, see [Azure IAM policies](https://docs.redpanda.com/cloud-data-platform/security/authorization/cloud-iam-policies-azure/). ## [](#manage-custom-tags)Manage custom tags Your organization might require custom tags for cost allocation, audit compliance, or governance policies. After cluster creation, you can manage tags with the [Cloud Control Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/). The Control Plane API allows up to 16 custom tags in Azure. Make sure you have: - The cluster ID. You can find this in the Redpanda Cloud UI, in the **Details** section of the cluster overview. - A valid bearer token for the Cloud Control Plane API. For details, see [Authenticate to the API](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication). Then complete the following steps: 1. To refresh Redpanda agent permissions in the target subscription, run: ```bash export CLUSTER_ID="" export SUBSCRIPTION_ID="" rpk cloud byoc azure apply --redpanda-id="$CLUSTER_ID" --subscription-id="$SUBSCRIPTION_ID" ``` 2. To update tags, invoke the Cloud API. First, set your authentication token: ```bash export AUTH_TOKEN="" ``` The `PATCH` call sets the tags specified under `"cloud_provider_tags"`. It replaces the existing tags with the specified tags. Include all desired tags in the request. To remove a single entry, omit it from the map you send. ```bash cluster_patch_body=$(cat <<'JSON' { "cloud_provider_tags": { "Environment": "production", "CostCenter": "engineering" } } JSON ) curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` To remove all tags, send an empty `cloud_provider_tags` object: ```bash cluster_patch_body='{"cloud_provider_tags": {}}' curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` ### [](#limitations)Limitations - Nodepool Application Security Groups (ASG): Custom tags are set only when the cluster is created. Tags cannot be updated on these resources after cluster creation. - Private Link network interfaces (Kubernetes API server, Tiered Storage, and Private Link service): Custom tags are set only during cluster creation and cannot be changed later. --- # Page 365: Create a BYOVNet Cluster on Azure **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/vnet-azure.md --- # Create a BYOVNet Cluster on Azure > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a BYOVNet Cluster on Azure latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/azure/vnet-azure page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/azure/vnet-azure.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/azure/vnet-azure.adoc description: Use Terraform to deploy a BYOVNet cluster on Azure. page-topic-type: how-to personas: platform_admin learning-objective-1: Deploy a BYOVNet cluster on Azure using Terraform learning-objective-2: Configure the Redpanda network and cluster resources using the Cloud API learning-objective-3: Manage the lifecycle of a BYOVNet cluster, including creation and deletion page-git-created-date: "2024-11-15" page-git-modified-date: "2026-08-27" --- > ❗ **IMPORTANT** > > BYOVPC/BYOVNet is an add-on feature that requires Premium support. To unlock this feature for your account, contact your Redpanda account team or [Redpanda Sales](https://www.redpanda.com/price-estimator). A Bring Your Own Virtual Network (BYOVNet) cluster allows you to deploy the Redpanda [data plane](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#data-plane) into your existing VNet and manage the networking lifecycle. Compared to a standard Bring Your Own Cloud (BYOC) setup, where Redpanda manages the networking lifecycle for you, BYOVNet provides more control. For background on the architecture, see [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/). When you create a BYOVNet cluster, you specify your VNet and managed identities. The Redpanda Cloud agent doesn’t create any new resources or alter any settings in your account. With a customer-managed VNet: - You provide your own VNet in your Azure account. - You maintain more control over your account, because Redpanda requires fewer permissions than standard BYOC clusters. - You control your security resources and policies, including subnets, user-assigned identities, IAM roles and assignments, security groups, storage accounts, and key vaults. The [Redpanda Cloud Examples repository](https://github.com/redpanda-data/cloud-examples/tree/main/customer-managed/azure/README.md) contains [Terraform](https://developer.hashicorp.com/terraform) code that deploys the resources required for a BYOVNet cluster on Azure. You need to create these resources in advance and provide them to Redpanda during cluster creation. Variables are provided in the code so you can exclude resources that already exist in your environment, such as the VNet. See the code for the complete list of resources required to create and deploy a Redpanda cluster. Customer-managed resources can be broken down into the following groups: - Resource group resources - User-assigned identities - IAM roles and assignments - Network - Storage - Key vaults ## [](#prerequisites)Prerequisites - Access to an Azure subscription where you want to create your cluster - Knowledge of your internal VNet and subnet configuration - Permission to call the [Redpanda Cloud API](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview) - Permission to create, modify, and delete the resources described by Terraform - [Terraform](https://developer.hashicorp.com/terraform/install) version 1.8.5 or later - [jq](https://jqlang.org/download/), which is used to parse JSON values from API responses ## [](#limitations)Limitations - Existing clusters cannot be moved to a BYOVNet cluster. - After creating a BYOVNet cluster, you cannot change to a different VNet. - Only primary CIDR ranges are supported for the VNet. ## [](#set-environment-variables)Set environment variables Set environment variables for the resource group, VNet name, and Azure region. For example: ```bash export AZURE_RESOURCE_GROUP_NAME=sample-redpanda-rg export AZURE_VNET_NAME="sample-vnet" export AZURE_REGION=centralus ``` ## [](#create-azure-resource-group-and-vnet)Create Azure resource group and VNet 1. Create a resource group to contain all resources, and then create a VNet with your address and subnet prefixes. The following example uses the environment variables to create the `sample-redpanda-rg` resource group and the `sample-vnet` virtual network with an address space of `10.0.0.0/16`. ```bash az group create --name ${AZURE_RESOURCE_GROUP_NAME} --location ${AZURE_REGION} az network vnet create \ --name ${AZURE_VNET_NAME} \ --resource-group ${AZURE_RESOURCE_GROUP_NAME} \ --location ${AZURE_REGION} \ --address-prefix 10.0.0.0/16 ``` 2. Set additional environment variables for Azure resources. For example: ```bash export AZURE_SUBSCRIPTION_ID= export AZURE_TENANT_ID= export AZURE_ZONES='["centralus-az1", "centralus-az2", "centralus-az3"]' export AZURE_RESOURCE_PREFIX=sample- export REDPANDA_CLUSTER_NAME= export REDPANDA_RG_ID= export REDPANDA_THROUGHPUT_TIER=tier-1-azure-v3-x86 export REDPANDA_VERSION= export REDPANDA_MANAGEMENT_STORAGE_ACCOUNT_NAME=rpmgmtsa export REDPANDA_MANAGEMENT_STORAGE_CONTAINER_NAME=rpmgmtsc export REDPANDA_0_PODS_SUBNET_NAME=snet-rp-0-pods export REDPANDA_0_VNET_SUBNET_NAME=snet-rp-0-vnet export REDPANDA_1_PODS_SUBNET_NAME=snet-rp-1-pods export REDPANDA_1_VNET_SUBNET_NAME=snet-rp-1-vnet export REDPANDA_2_PODS_SUBNET_NAME=snet-rp-2-pods export REDPANDA_2_VNET_SUBNET_NAME=snet-rp-2-vnet export REDPANDA_CONNECT_PODS_SUBNET_NAME=snet-connect-pods export REDPANDA_CONNECT_VNET_SUBNET_NAME=snet-connect-vnet export KAFKA_CONNECT_PODS_SUBNET_NAME=snet-kafka-connect-pods export KAFKA_CONNECT_VNET_SUBNET_NAME=snet-kafka-connect-vnet export SYSTEM_PODS_SUBNET_NAME=snet-system-pods export SYSTEM_VNET_SUBNET_NAME=snet-system-vnet export REDPANDA_AGENT_SUBNET_NAME=snet-agent-private export REDPANDA_EGRESS_SUBNET_NAME=snet-agent-public export REDPANDA_MANAGEMENT_KEY_VAULT_NAME=redpanda-vault export REDPANDA_CONSOLE_KEY_VAULT_NAME=rp-console-vault export REDPANDA_AKS_SUBNET_CIDR="10.0.15.0/24" export REDPANDA_IAM_RESOURCE_GROUP_NAME=sample-redpanda-rg export REDPANDA_NETWORK_RESOURCE_GROUP_NAME=sample-redpanda-rg export REDPANDA_RESOURCE_GROUP_NAME=sample-redpanda-rg export REDPANDA_STORAGE_RESOURCE_GROUP_NAME=sample-redpanda-rg export REDPANDA_SECURITY_GROUP_NAME=redpanda-nsg export REDPANDA_TIERED_STORAGE_ACCOUNT_NAME=tieredsa export REDPANDA_TIERED_STORAGE_CONTAINER_NAME=tieredsc export REDPANDA_AGENT_USER_ASSIGNED_IDENTITY_NAME=agent-uai export REDPANDA_AKS_USER_ASSIGNED_IDENTITY_NAME=aks-uai export REDPANDA_CERT_MANAGER_USER_ASSIGNED_IDENTITY_NAME=cert-manager-uai export REDPANDA_EXTERNAL_DNS_USER_ASSIGNED_IDENTITY_NAME=external-dns-uai export REDPANDA_CLUSTER_USER_ASSIGNED_IDENTITY_NAME=cluster-uai export REDPANDA_CONSOLE_USER_ASSIGNED_IDENTITY_NAME=console-uai export KAFKA_CONNECT_USER_ASSIGNED_IDENTITY_NAME=kafka-connect-uai export REDPANDA_CONNECT_USER_ASSIGNED_IDENTITY_NAME=redpanda-connect-uai export REDPANDA_CONNECT_API_USER_ASSIGNED_IDENTITY_NAME=redpanda-connect-api-uai export REDPANDA_OPERATOR_USER_ASSIGNED_IDENTITY_NAME=redpanda-operator-uai ``` > 📝 **NOTE** > > Replace `` with a `major.minor` version supported by Redpanda Cloud, which may be earlier than the latest Redpanda release. Supported versions come from Redpanda Cloud’s certified install packs. You can also omit `redpanda_version` from the cluster request body, in which case Redpanda Cloud deploys its default version. ## [](#configure-terraform)Configure Terraform > 📝 **NOTE** > > For simplicity, these instructions assume that Terraform is configured to use local state. You may want to configure [remote state](https://developer.hashicorp.com/terraform/language/state/remote). Create a JSON file called `byovnet.auto.tfvars.json` inside the Terraform directory to configure variables for your specific needs: Show script ```bash cat > byovnet.auto.tfvars.json < 💡 **TIP** > > To get the Redpanda authentication credentials, follow the [authentication guide](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/authentication). ## [](#create-the-network)Create the network To create the Redpanda network: 1. Define a JSON file called `redpanda-network.json` to configure the network for Redpanda with details about VNet, subnets, and storage. Show script ```bash cat > redpanda-network.json < redpanda-cluster.json < 💡 **TIP** > > See the full list of zones and tiers available with each provider in the [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-regions-and-usage-tiers). 2. Make a Cloud API call to create a Redpanda cluster and get the network ID from the response in JSON `.operation.metadata.network_id`. ```bash export REDPANDA_ID=$(curl -X POST "https://api.redpanda.com/v1/clusters" \ -H "accept: application/json"\ -H "content-type: application/json" \ -H "authorization: Bearer ${BEARER_TOKEN}" \ --data-binary @redpanda-cluster.json | jq -r '.operation.resource_id') ``` ## [](#create-the-cluster-resources)Create the cluster resources To create the initial cluster resources, first log in to Redpanda Cloud, then run `rpk cloud byoc azure apply`: ```bash rpk cloud login \ --save \ --client-id=${REDPANDA_CLIENT_ID} \ --client-secret=${REDPANDA_CLIENT_SECRET} \ --no-profile ``` ```bash rpk cloud byoc azure apply --redpanda-id="${REDPANDA_ID}" --subscription-id="${AZURE_SUBSCRIPTION_ID}" ``` The Redpanda Cloud agent now is running and handles the remaining steps. This can take up to 45 minutes. When provisioning completes, the cluster status updates to `Running`. If the cluster remains in `Creating` status after 45 minutes, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). ## [](#check-the-cluster-status)Check the cluster status Cluster creation is an example of an operation that can take a longer period of time to complete. You can check the operation state with the Cloud API, or check the Redpanda Cloud UI for cluster status. Example using the returned `operation_id`: ```bash curl -X GET "https://api.redpanda.com/v1/operations/" \ -H "accept: application/json"\ -H "content-type: application/json" \ -H "authorization: Bearer ${BEARER_TOKEN}" ``` Example retrieving cluster: ```bash curl -X GET "https://api.redpanda.com/v1/clusters/" \ -H "accept: application/json"\ -H "content-type: application/json" \ -H "authorization: Bearer ${BEARER_TOKEN}" ``` ## [](#delete-the-cluster)Delete the cluster To delete the cluster, first send a DELETE request to the Cloud API, and retrieve the `resource_id` of the DELETE operation. Then run the `rpk` command to destroy the cluster identified by the `resource_id`. ```bash export REDPANDA_ID=$(curl -X DELETE "https://api.redpanda.com/v1/clusters/${REDPANDA_ID}" \ -H "accept: application/json"\ -H "content-type: application/json" \ -H "authorization: Bearer ${BEARER_TOKEN}" | jq -r '.operation.resource_id') ``` After that completes, run: ```bash rpk cloud byoc azure destroy --redpanda-id ${REDPANDA_ID} ``` > 📝 **NOTE** > > Redpanda Cloud does not support customer access or modifications to any of the internal data plane resources. This restriction allows Redpanda Data to manage all configuration changes internally to ensure a 99.99% service level agreement (SLA) for BYOC clusters. ## [](#manage-custom-tags)Manage custom tags Your organization might require custom tags for cost allocation, audit compliance, or governance policies. After cluster creation, you can manage tags with the [Cloud Control Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/). The Control Plane API allows up to 16 custom tags in Azure. Make sure you have: - The cluster ID. You can find this in the Redpanda Cloud UI, in the **Details** section of the cluster overview. - A valid bearer token for the Cloud Control Plane API. For details, see [Authenticate to the API](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication). Then complete the following steps: 1. To refresh Redpanda agent permissions in the target subscription, run: ```bash export CLUSTER_ID="" export SUBSCRIPTION_ID="" rpk cloud byoc azure apply --redpanda-id="$CLUSTER_ID" --subscription-id="$SUBSCRIPTION_ID" ``` 2. To update tags, invoke the Cloud API. First, set your authentication token: ```bash export AUTH_TOKEN="" ``` The `PATCH` call sets the tags specified under `"cloud_provider_tags"`. It replaces the existing tags with the specified tags. Include all desired tags in the request. To remove a single entry, omit it from the map you send. ```bash cluster_patch_body=$(cat <<'JSON' { "cloud_provider_tags": { "Environment": "production", "CostCenter": "engineering" } } JSON ) curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` To remove all tags, send an empty `cloud_provider_tags` object: ```bash cluster_patch_body='{"cloud_provider_tags": {}}' curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` ### [](#limitations-2)Limitations - Nodepool Application Security Groups (ASG): Custom tags are set only when the cluster is created. Tags cannot be updated on these resources after cluster creation. - Private Link network interfaces (Kubernetes API server, Tiered Storage, and Private Link service): Custom tags are set only during cluster creation and cannot be changed later. > 📝 **NOTE** > > For BYOVNet clusters, custom tags are not applied to the customer-managed resources that are deployed by the customer. ## [](#next-steps)Next steps - [Configure Azure Private Link](https://docs.redpanda.com/cloud-data-platform/networking/azure-private-link/) - [Learn about `rpk` commands](https://docs.redpanda.com/cloud-data-platform/reference/rpk/) --- # Page 366: BYOC: GCP **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp.md --- # BYOC: GCP > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: "BYOC: GCP" latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/gcp/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/gcp/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/gcp/index.adoc description: Learn how to create a BYOC or BYOVPC cluster on GCP. page-git-created-date: "2024-10-24" page-git-modified-date: "2025-05-07" --- - [Create a BYOC Cluster on GCP](create-byoc-cluster-gcp/) Use the Redpanda Cloud UI to create a BYOC cluster on GCP. - [Create a BYOVPC Cluster on GCP](vpc-byo-gcp/) Connect Redpanda Cloud to your existing VPC for additional security. - [Enable Redpanda Connect on an Existing BYOVPC Cluster on GCP](enable-rpcn-byovpc-gcp/) Add Redpanda Connect to your existing BYOVPC cluster. - [Enable Secrets Management on an Existing BYOVPC Cluster on GCP](enable-secrets-byovpc-gcp/) Store and read secrets in your existing BYOVPC cluster. --- # Page 367: Create a BYOC Cluster on GCP **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/create-byoc-cluster-gcp.md --- # Create a BYOC Cluster on GCP > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a BYOC Cluster on GCP latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/gcp/create-byoc-cluster-gcp page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/gcp/create-byoc-cluster-gcp.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/gcp/create-byoc-cluster-gcp.adoc description: Use the Redpanda Cloud UI to create a BYOC cluster on GCP. page-git-created-date: "2024-10-24" page-git-modified-date: "2026-07-08" --- To create a Redpanda cluster in your virtual private cloud (VPC), follow the instructions in the Redpanda Cloud UI. The UI contains the parameters necessary to successfully run `rpk cloud byoc apply`. See also: [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/). > 📝 **NOTE** > > With standard BYOC clusters, Redpanda manages security policies and resources for your VPC, including subnetworks, service accounts, IAM roles, firewall rules, and storage buckets. For the highest level of security, you can manage these resources yourself with a [BYOVPC cluster on GCP](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/vpc-byo-gcp/). If your clients need to connect from different GCP regions than where your cluster will be deployed, you must enable global access during cluster creation using the Cloud API. To create a BYOC cluster with global access enabled, see [Enable Global Access](https://docs.redpanda.com/cloud-data-platform/networking/byoc/gcp/enable-global-access/). ## [](#prerequisites)Prerequisites Before you deploy a BYOC cluster on GCP, verify the following prerequisites: - A minimum version of Redpanda `rpk` v24.1. See [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). - Assign the `roles/editor` role (or higher, such as `roles/owner`) to the GCP user or service account that runs the bootstrap on the target GCP project. This grants the permissions needed to create VPC networks, GKE clusters, service accounts, and other infrastructure during the initial bootstrap. These bootstrap permissions are separate from the [agent permissions](https://docs.redpanda.com/cloud-data-platform/security/authorization/cloud-iam-policies-gcp/) that Redpanda assigns after bootstrap. - The user has the [Google Cloud CLI](https://cloud.google.com/sdk/docs/install) installed and authenticated, with the target project selected. To verify, run: ```bash gcloud auth list gcloud config get-value project ``` ### [](#gcp-quotas)GCP quotas Ensure at least three nodes of headroom in the relevant GCP quotas in the same region as your cluster. During maintenance, Redpanda may temporarily create extra nodes. Quotas such as vCPUs per VM family (for example, N2D) and Local SSD total per VM family (quota key: `LOCAL_SSD_TOTAL_GB_PER_VM_FAMILY`) are listed for each tier on the **Create BYOC cluster** page in the Redpanda Cloud UI. Headroom formulas: - vCPU spare = `3 x (vCPUs per node)` - Local SSD spare (GB) = `3 x (Storage size per node in GB)` For example, with per-node storage **1500 GB** (4 × 375 GB Local SSD) and machine type **n2d-standard-4** (4 vCPUs), keep **4500 GB** Local SSD and **12 vCPUs** of spare quota. ## [](#create-a-byoc-cluster)Create a BYOC cluster 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. On the Clusters page, click **Create cluster**, then click **Create** for BYOC. Enter a cluster name, then select the resource group, provider (GCP), [region, tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/), availability, and Redpanda version. > 📝 **NOTE** > > - If you plan to create a private network in your own VPC, select the region where your VPC is located. > > - Three availability zones provide two backups in case one availability zone goes down. Optionally, click **Advanced settings** to specify up to five key-value custom GCP labels. If a label key starts with `gcp.network-tag.`, then the agent interprets it as a request to apply the `` [network tag](https://cloud.google.com/vpc/docs/add-remove-network-tags) to GCE instances in the cluster. Use labels for organization/metadata; use network tags to target firewall rules and routes. After the cluster is created, labels are applied to applicable GCP resources (for example, instances and disks), and network tags are applied to instances. For more information, see the [GCP documentation](https://cloud.google.com/compute/docs/labeling-resources). You can also [specify more labels and network tags on the Dataplane settings page or with the Control Plane API](#manage-custom-resource-labels-and-network-tags). 3. Click **Next**. 4. On the Network page, select the connection type: either public or private. For BYOC clusters, private is best-practice. - Your network name is used to identify this network. - For a [CIDR range](https://docs.redpanda.com/cloud-data-platform/networking/cidr-ranges/), choose one that does not overlap with your existing VPCs or your Redpanda network. - Clusters with private networking include a setting for API Gateway network access. Public access exposes endpoints for Redpanda Console, the Data Plane API, but they remain protected by your authentication and authorization controls. Private access restricts endpoint access to your VPC only. > 📝 **NOTE** > > After the cluster is created, you can change the API Gateway access on the Dataplane settings page. If you change from public to private access, users without VPN access to the Redpanda VPC will lose access to these services. > 💡 **TIP** > > To route all cluster egress through your own GCP hub VPC and NAT VM instead of a per-cluster Cloud NAT, enter the **Hub VPC name** and **Hub project ID** on this page. These fields are only available on clusters with a private connection type, and are only visible if centralized egress is enabled for your organization. This option is in beta. See [Configure Centralized Egress with GCP VPC Peering](https://docs.redpanda.com/cloud-data-platform/networking/byoc/gcp/nat-free-egress/). 5. Click **Next**. 6. On the Deploy page, follow the steps to log in to Redpanda Cloud and deploy the agent. As part of agent deployment, Redpanda assigns the permissions required to run the agent. For details about these permissions, see [GCP IAM permissions](https://docs.redpanda.com/cloud-data-platform/security/authorization/cloud-iam-policies-gcp/). > 📝 **NOTE** > > Redpanda Cloud does not support customer access or modifications to any of the internal data plane resources. This restriction allows Redpanda Data to manage all configuration changes internally to ensure a 99.99% service level agreement (SLA) for BYOC clusters. ## [](#manage-custom-resource-labels-and-network-tags)Manage custom resource labels and network tags Your organization might require custom resource labels and network tags for cost allocation, audit compliance, or governance policies. After cluster creation, you can manage labels and network tags on your cluster’s **Dataplane settings** page in the Redpanda Cloud UI, or with the [Control Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/). > ⚠️ **CAUTION** > > Do not add labels or network tags directly to the GCP node pools of a BYOC cluster, either in the GCP console or with commands such as `gcloud container node-pools update`. Google Kubernetes Engine (GKE) treats a node pool label or tag update as a node replacement and aggressively replaces all nodes, rather than performing a controlled rolling upgrade. This bypasses the failover process that Redpanda requires for safe node cycling. It can leave persistent volume claims (PVCs) in a pending state and may require recovering the entire cluster. > > The Dataplane settings page and the Control Plane API are the supported methods. With both, Redpanda applies labels directly to the Google Compute Engine (GCE) instances and disks in your cluster, and network tags directly to the instances. Because nothing is applied to the node pools, GKE does not replace any nodes. ### [](#use-the-dataplane-settings-page)Use the Dataplane settings page To manage labels and network tags in the Redpanda Cloud UI: 1. In the Redpanda Cloud UI, open your [cluster](https://cloud.redpanda.com/clusters), and click **Dataplane settings**. 2. Under **Manage resource labels and network tags for your cluster**, click **Add**, and enter a key and value for each label. Keys and values are case-sensitive. To apply a network tag to the GCE instances in the cluster, use a key that starts with `gcp.network-tag.`. For example, the key `gcp.network-tag.web-servers` applies the `web-servers` [network tag](https://cloud.google.com/vpc/docs/add-remove-network-tags). You can add up to 10 entries on this page. To manage up to 16, use the [Control Plane API](#use-the-control-plane-api). 3. Click **Save**. ### [](#use-the-control-plane-api)Use the Control Plane API The Control Plane API allows up to 16 custom resource labels and network tags in GCP. Make sure you have: - The cluster ID. You can find this in the Redpanda Cloud UI, in the **Details** section of the cluster overview. - A valid bearer token for the Control Plane API. For details, see [Authenticate to the API](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication). Then complete the following steps: 1. To refresh agent permissions so the Redpanda agent can update labels and network tags, run: ```bash export CLUSTER_ID="" export PROJECT_ID="" rpk cloud byoc gcp apply --redpanda-id="$CLUSTER_ID" --project-id="$PROJECT_ID" ``` This step is required because label/tag management requires additional IAM permissions that may not have been granted during initial cluster creation: - `compute.disks.get` - `compute.disks.list` - `compute.disks.setLabels` - `compute.instances.setLabels` 2. To update labels and network tags, invoke the Control Plane API. First, set your authentication token: ```bash export AUTH_TOKEN="" ``` The `PATCH` call sets the labels and network tags specified under `"cloud_provider_tags"`. It replaces the existing labels and tags with the specified labels and tags. Include all desired labels and tags in the request. To remove a single entry, omit it from the map you send. ```bash cluster_patch_body=$(cat <<'JSON' { "cloud_provider_tags": { "environment": "production", "cost-center": "engineering", "gcp.network-tag.web-servers": "true", "gcp.network-tag.database-access": "true" } } JSON ) curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` To remove all labels and network tags, send an empty `cloud_provider_tags` object: ```bash cluster_patch_body='{"cloud_provider_tags": {}}' curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` ## [](#service-account-credential-rotation)Service account credential rotation To rotate service account credentials for your BYOC cluster, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new) with your cluster ID, the service accounts that require rotation, and your target timeline. > ⚠️ **WARNING** > > GCP service account credential rotation for BYOC clusters is not self-service. Rotating these credentials without coordinating with Redpanda can disrupt agent connectivity, monitoring, and Tiered Storage uploads, and can leave the cluster stuck and unable to complete future operations. ## [](#next-steps)Next steps [Configure private networking](https://docs.redpanda.com/cloud-data-platform/networking/byoc/gcp/) --- # Page 368: Enable Redpanda Connect on an Existing BYOVPC Cluster on GCP **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/enable-rpcn-byovpc-gcp.md --- # Enable Redpanda Connect on an Existing BYOVPC Cluster on GCP > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Enable Redpanda Connect on an Existing BYOVPC Cluster on GCP latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/gcp/enable-rpcn-byovpc-gcp page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/gcp/enable-rpcn-byovpc-gcp.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/gcp/enable-rpcn-byovpc-gcp.adoc description: Add Redpanda Connect to your existing BYOVPC cluster. page-git-created-date: "2025-04-04" page-git-modified-date: "2026-05-26" --- > ❗ **IMPORTANT** > > BYOVPC is an add-on feature that may require an additional purchase. To unlock this feature for your account, contact your Redpanda account team or [Redpanda Sales](https://www.redpanda.com/price-estimator). To enable Redpanda Connect on an existing BYOVPC cluster, you must update your configuration. You can also create [a new BYOVPC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/vpc-byo-gcp/) with Redpanda Connect already enabled. Replace all `` with your own values. 1. Create two new service accounts with the necessary permissions and roles. Show commands ```bash # Account used to check for and read secrets, which are required to create Redpanda Connect pipelines. gcloud iam service-accounts create redpanda-connect-api \ --display-name="Redpanda Connect API Service Account" cat << EOT > redpanda-connect-api.role { "name": "redpanda_connect_api_role", "title": "Redpanda Connect API Role", "description": "Redpanda Connect API Role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.secrets.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_connect_api_role --project= --file redpanda-connect-api.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-connect-api@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_connect_api_role" ``` ```bash # Account used to retrieve secrets and create Redpanda Connect pipelines. gcloud iam service-accounts create redpanda-connect \ --display-name="Redpanda Connect Service Account" cat << EOT > redpanda-connect.role { "name": "redpanda_connect_role", "title": "Redpanda Connect Role", "description": "Redpanda Connect Role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_connect_role --project= --file redpanda-connect.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-connect@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_connect_role" ``` 2. Bind the service accounts. The account ID of the GCP service account is used to configure service account bindings. This account ID is the local part of the email address for the GCP service account. For example, if the GCP service account is `my-gcp-sa@my-project.iam.gserviceaccount.com`, then the account ID is `my-gcp-sa`. Show commands ```none gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-connect/]" ``` ```none gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-connect/]" ``` 3. Make a [`PATCH /v1/clusters/{cluster-id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) request to update the cluster configuration. Show request ```bash export CLUSTER_PATCH_BODY=`cat << EOF { "customer_managed_resources": { "gcp": { "redpanda_connect_api_service_account": { "email": "@.iam.gserviceaccount.com" }, "redpanda_connect_service_account": { "email": "@.iam.gserviceaccount.com" } } } } EOF` curl -v -X PATCH \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$CLUSTER_PATCH_BODY" $PUBLIC_API_ENDPOINT/v1/clusters/ ``` 4. Check Redpanda Connect is available in the Cloud UI. 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Go to the **Connect** page and you should see Redpanda Connect. ## [](#next-steps)Next steps - Choose [connectors for your use case](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/about/). - Learn how to [configure, test, and run a data pipeline locally](https://docs.redpanda.com/connect/get-started/quickstarts/rpk/). - Try the [Redpanda Connect quickstart](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/). - Try one of our [Redpanda Connect cookbooks](https://docs.redpanda.com/cloud-data-platform/develop/connect/cookbooks/). - Learn how to [add secrets to your pipeline](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/). --- # Page 369: Enable Secrets Management on an Existing BYOVPC Cluster on GCP **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/enable-secrets-byovpc-gcp.md --- # Enable Secrets Management on an Existing BYOVPC Cluster on GCP > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Enable Secrets Management on an Existing BYOVPC Cluster on GCP page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/gcp/enable-secrets-byovpc-gcp page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/gcp/enable-secrets-byovpc-gcp.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/gcp/enable-secrets-byovpc-gcp.adoc description: Store and read secrets in your existing BYOVPC cluster. # Beta release status page-beta: "true" page-git-created-date: "2025-06-06" page-git-modified-date: "2025-08-20" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- > ❗ **IMPORTANT** > > BYOVPC is an add-on feature that may require an additional purchase. To unlock this feature for your account, contact your Redpanda account team or [Redpanda Sales](https://www.redpanda.com/price-estimator). Storing secrets in your cluster allows you to keep your cloud infrastructure secure as you integrate your data across different systems, for example, REST catalogs with your Iceberg-enabled topics. If you do not have secrets management enabled on an existing BYOVPC cluster, you can do so by following the steps on this page to update your cluster configuration. You can also create [a new BYOVPC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/vpc-byo-gcp/) with secrets management already enabled. Replace all `` with your own values. 1. Create one new service account with the necessary permissions and roles. Show commands ```bash # Account used to check for and read secrets gcloud iam service-accounts create redpanda-operator \ --display-name="Redpanda Operator Service Account" cat << EOT > redpanda-operator.role { "name": "redpanda_operator_role", "title": "Redpanda Operator Role", "description": "Redpanda Operator Role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.secrets.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_operator_role --project= --file redpanda-operator.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-operator@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_operator_role" ``` 2. Update the existing Redpanda cluster service account with the necessary permissions to read secrets. Show commands ```bash cat << EOT > redpanda-cluster.role { "name": "redpanda_cluster_role", "title": "Redpanda Cluster Role", "description": "Redpanda Cluster Role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.secrets.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_cluster_role --project= --file redpanda-cluster.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-cluster@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_cluster_role" ``` 3. Bind the new service account. The account ID of the GCP service account is used to configure service account bindings. This account ID is the local part of the email address for the GCP service account. For example, if the GCP service account is `my-gcp-sa@my-project.iam.gserviceaccount.com`, then the account ID is `my-gcp-sa`. Show commands ```none gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-system/]" ``` 4. Make a [`PATCH /v1/clusters/{cluster-id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) request to update the cluster configuration. Show request ```bash export CLUSTER_PATCH_BODY=`cat << EOF { "customer_managed_resources": { "gcp": { "redpanda_operator_service_account": { "email": "@.iam.gserviceaccount.com" } } } } EOF` curl -v -X PATCH \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$CLUSTER_PATCH_BODY" $PUBLIC_API_ENDPOINT/v1/clusters/ ``` 5. Check secrets management is available in the Cloud UI. 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Go to the **Secrets Store** page of your cluster. You should be able to create a new secret. ## [](#next-steps)Next steps - [Reference a secret in a cluster property](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/#set-cluster-configuration-properties). - [Integrate a catalog](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/) for querying Iceberg topics in your cluster. --- # Page 370: Create a BYOVPC Cluster on GCP **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/vpc-byo-gcp.md --- # Create a BYOVPC Cluster on GCP > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create a BYOVPC Cluster on GCP latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/gcp/vpc-byo-gcp page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/gcp/vpc-byo-gcp.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/gcp/vpc-byo-gcp.adoc description: Connect Redpanda Cloud to your existing VPC for additional security. page-git-created-date: "2024-10-24" page-git-modified-date: "2026-08-03" --- > ❗ **IMPORTANT** > > BYOVPC/BYOVNet is an add-on feature that requires Premium support. To unlock this feature for your account, contact your Redpanda account team or [Redpanda Sales](https://www.redpanda.com/price-estimator). A Bring Your Own Virtual Private Cloud (BYOVPC) cluster allows you to deploy the Redpanda [data plane](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#data-plane) into your existing VPC and manage the networking lifecycle. Compared to a standard Bring Your Own Cloud (BYOC) setup, where Redpanda manages the networking lifecycle for you, BYOVPC provides more control. See also: [BYOC architecture](https://docs.redpanda.com/cloud-data-platform/get-started/byoc-arch/). When you create a BYOVPC cluster, you specify your VPC and service account. The Redpanda Cloud agent doesn’t create any new resources or alter any settings in your account. With BYOVPC: - You provide your own VPC in your Google Cloud account. - You maintain more control of your Google Cloud account, because Redpanda requires fewer permissions than standard BYOC clusters. - You control your security resources and policies, including subnets, service accounts, IAM roles, firewall rules, and storage buckets. If your clients need to connect from different GCP regions than where your cluster will be deployed, you must enable global access during cluster creation. To create a BYOVPC cluster with global access enabled, see [Enable Global Access](https://docs.redpanda.com/cloud-data-platform/networking/byoc/gcp/enable-global-access/). ## [](#prerequisites)Prerequisites - A standalone GCP project is recommended. If your host project (where your VPC project is created) and your service project (where your Redpanda cluster is created) are in different projects, you must first provision a shared VPC in Google Cloud. For more information, see the [Google shared VPC documentation](https://cloud.google.com/vpc/docs/provisioning-shared-vpc). - Redpanda creates a private Google Kubernetes Engine (GKE) cluster in your VPC. The subnet and secondary IP ranges you provide must allow public internet access. The configuration requires you to provide reserved CIDR ranges for the subnet and GKE Pods, Services, and master IP addresses. See the [GKE service account documentation](https://cloud.google.com/kubernetes-engine/docs/how-to/service-accounts) and [Configure your VPC](#configure-your-vpc). - Only primary CIDR ranges are supported for the VPC. - Redpanda requires access to certain Google APIs, storage buckets, and service accounts. See [Configure the service project](#configure-the-service-project). ### [](#gcp-quotas)GCP quotas Ensure at least three nodes of headroom in the relevant GCP quotas in the same region as your cluster. During maintenance, Redpanda may temporarily create extra nodes. Quotas such as vCPUs per VM family (for example, N2D) and Local SSD total per VM family (quota key: `LOCAL_SSD_TOTAL_GB_PER_VM_FAMILY`) are listed for each tier on the **Create BYOC cluster** page in the Redpanda Cloud UI. Headroom formulas: - vCPU spare = `3 x (vCPUs per node)` - Local SSD spare (GB) = `3 x (Storage size per node in GB)` For example, with per-node storage **1500 GB** (4 × 375 GB Local SSD) and machine type **n2d-standard-4** (4 vCPUs), keep **4500 GB** Local SSD and **12 vCPUs** of spare quota. ## [](#limitations)Limitations - Existing clusters cannot be moved to a BYOVPC cluster. - After creating a BYOVPC cluster, you cannot change to a different VPC. ## [](#configure-your-vpc)Configure your VPC 1. Create the primary and secondary subnets in your VPC using CIDR notation. Redpanda clusters require one subnet, and that subnet should have two secondary IP ranges: - Subnet IP range should be at least /24 CIDR, such as 10.0.0.0/24. - Secondary IP for GKE Pods is a /21 CIDR, such as 10.0.8.0/21. - Secondary IP for GKE Services is a /24 CIDR, such as 10.0.1.0/24. Replace all `` with your own values. ```bash gcloud compute networks subnets create \ --project \ --network \ --range 10.0.0.0/24 \ --region \ --secondary-range =10.0.8.0/21,=10.0.1.0/24 ``` Additionally, a /28 CIDR is required for the GKE master IP addresses. This CIDR is not used in the GCP networking configuration, but is input into the Redpanda UI; for example, 10.0.7.240/28. 2. To enable egress, create a cloud router and NAT at the host project: ```bash gcloud compute routers create \ --project \ --region \ --network gcloud compute addresses create --region gcloud compute routers nats create \ --project \ --router \ --region \ --nat-all-subnet-ip-ranges \ --nat-external-ip-pool \ --enable-endpoint-independent-mapping ``` 3. Create VPC firewall rules. - Redpanda ingress: ```bash gcloud compute firewall-rules create redpanda-ingress \ --description="Allow access to Redpanda cluster" \ --network="" \ --project="" \ --direction="INGRESS" \ --target-tags="redpanda-node" \ --source-ranges="10.0.0.0/8,172.16.0.0/12,192.168.0.0/16,100.64.0.0/10" \ --allow="tcp:9092-9094,tcp:30081,tcp:30082,tcp:30092" ``` - Master webhooks: ```bash gcloud compute firewall-rules create gke-redpanda-cluster-webhooks \ --description="Allow master to hit pods for admission controllers/webhooks" \ --network="" \ --project="" \ --direction="INGRESS" \ --source-ranges="" \ --allow="tcp:9443,tcp:8443,tcp:6443" ``` Replace `` with a /28 CIDR. For example: 172.16.0.32/28. For information about the master CIDR, and how to set it using `--master-ipv4-cidr`, see the **gcloud** tab in [Creating a private cluster with no client access to the public endpoint](https://cloud.google.com/kubernetes-engine/docs/how-to/legacy/network-isolation#private_cp) 4. Grant permission to read the VPC and related resources. If the host project and service project are in different projects, it’s helpful for the Redpanda team to have read access to the VPC and related resources in the host project. If your host project and service project are the same, you can skip this step. - Redpanda Agent custom role: ```bash cat << EOT > redpanda-agent.role { "name": "redpanda_agent_role", "title": "Redpanda Agent Role", "description": "A role granting the redpanda agent permissions to view network resources in the project of the vpc.", "includedPermissions": [ "compute.firewalls.get", "compute.subnetworks.get", "resourcemanager.projects.get", "compute.networks.getRegionEffectiveFirewalls", "compute.networks.getEffectiveFirewalls" ] } EOT gcloud iam roles create redpanda_agent_role --project= --file redpanda-agent.role ``` ## [](#configure-the-service-project)Configure the service project 1. Enable Google APIs in the service project: ```bash gcloud services enable cloudresourcemanager.googleapis.com --project gcloud services enable dns.googleapis.com --project gcloud services enable secretmanager.googleapis.com --project gcloud services enable compute.googleapis.com --project gcloud services enable iam.googleapis.com --project gcloud services enable storage-api.googleapis.com --project gcloud services enable container.googleapis.com --project gcloud services enable serviceusage.googleapis.com --project ``` 2. Create storage buckets at the service project in the same region as the cluster: ```bash gcloud storage buckets create gs:// \ --location="" \ --uniform-bucket-level-access gcloud storage buckets create gs:// \ --location="" \ --uniform-bucket-level-access gcloud storage buckets update gs:// --versioning ``` - Redpanda uses the tiered storage bucket for writing log segments. This should not be versioned. - Redpanda uses the management storage bucket to store cluster metadata. This can have versioning enabled. 3. Create service accounts with necessary permissions and roles. - Redpanda Cloud agent service account Show commands ```bash gcloud iam service-accounts create redpanda-agent \ --display-name="Redpanda Agent Service Account" cat << EOT > redpanda-agent.role { "name": "redpanda_agent_role", "title": "Redpanda Agent Role", "description": "A role comprising general permissions allowing the agent to manage Redpanda cluster resources.", "includedPermissions": [ "compute.firewalls.get", "compute.disks.get", "compute.globalOperations.get", "compute.instanceGroupManagers.get", "compute.instanceGroupManagers.delete", "compute.instanceGroups.delete", "compute.instances.list", "compute.instanceTemplates.delete", "compute.networks.getRegionEffectiveFirewalls", "compute.networks.getEffectiveFirewalls", "compute.projects.get", "compute.subnetworks.get", "compute.zoneOperations.get", "compute.zoneOperations.list", "compute.zones.get", "compute.zones.list", "dns.changes.create", "dns.changes.get", "dns.changes.list", "dns.managedZones.create", "dns.managedZones.delete", "dns.managedZones.get", "dns.managedZones.list", "dns.managedZones.update", "dns.projects.get", "dns.resourceRecordSets.create", "dns.resourceRecordSets.delete", "dns.resourceRecordSets.get", "dns.resourceRecordSets.list", "dns.resourceRecordSets.update", "iam.roles.get", "iam.roles.list", "iam.serviceAccounts.actAs", "iam.serviceAccounts.get", "iam.serviceAccounts.getIamPolicy", "resourcemanager.projects.get", "resourcemanager.projects.getIamPolicy", "serviceusage.services.list", "storage.buckets.get", "storage.buckets.getIamPolicy", "compute.subnetworks.use", "compute.instances.use", "compute.networks.use", "compute.regionOperations.get", "compute.serviceAttachments.create", "compute.serviceAttachments.delete", "compute.serviceAttachments.get", "compute.serviceAttachments.list", "compute.serviceAttachments.update", "compute.forwardingRules.use", "compute.forwardingRules.create", "compute.forwardingRules.delete", "compute.forwardingRules.get", "compute.forwardingRules.setLabels", "compute.forwardingRules.setTarget", "compute.forwardingRules.pscCreate", "compute.forwardingRules.pscDelete", "compute.forwardingRules.pscSetLabels", "compute.forwardingRules.pscSetTarget", "compute.forwardingRules.pscUpdate", "compute.regionBackendServices.create", "compute.regionBackendServices.delete", "compute.regionBackendServices.get", "compute.regionBackendServices.use", "compute.regionNetworkEndpointGroups.create", "compute.regionNetworkEndpointGroups.delete", "compute.regionNetworkEndpointGroups.get", "compute.regionNetworkEndpointGroups.use", "compute.regionNetworkEndpointGroups.attachNetworkEndpoints", "compute.regionNetworkEndpointGroups.detachNetworkEndpoints", "compute.disks.list", "compute.disks.setLabels", "compute.instanceGroupManagers.update", "compute.instances.delete", "compute.instances.get", "compute.instances.setLabels" ] } EOT gcloud iam roles create redpanda_agent_role --project= --file redpanda-agent.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-agent@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_agent_role" gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-agent@.iam.gserviceaccount.com" \ --role="roles/container.admin" gcloud storage buckets add-iam-policy-binding gs:// \ --member="serviceAccount:redpanda-agent@.iam.gserviceaccount.com" \ --role="roles/storage.objectAdmin" # skip this step if host project and service project are the same gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-agent@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_agent_role" ``` - Redpanda cluster service account Show commands ```bash cat << EOT > redpanda-cluster.role { "name": "redpanda_cluster_role", "title": "Redpanda Cluster Role", "description": "Redpanda Cluster role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.secrets.get", "secretmanager.versions.access" ] } EOT gcloud iam service-accounts create redpanda-cluster \ --display-name="Redpanda Cluster Service Account" gcloud storage buckets add-iam-policy-binding gs:// \ --member="serviceAccount:redpanda-cluster@.iam.gserviceaccount.com" \ --role="roles/storage.objectAdmin" gcloud iam roles create redpanda_cluster_role --project= --file redpanda-cluster.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-cluster@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_cluster_role" ``` - Redpanda operator service account Show commands ```bash gcloud iam service-accounts create redpanda-operator \ --display-name="Redpanda Operator Service Account" cat << EOT > redpanda-operator.role { "name": "redpanda_operator_role", "title": "Redpanda Operator Role", "description": "Redpanda Operator role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.secrets.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_operator_role --project= --file redpanda-operator.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-operator@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_operator_role" ``` - Redpanda Connect service accounts Show commands ```bash # Account used to check for and read secrets, which are required to create Redpanda Connect pipelines. gcloud iam service-accounts create redpanda-connect-api \ --display-name="Redpanda Connect API Service Account" cat << EOT > redpanda-connect-api.role { "name": "redpanda_connect_api_role", "title": "Redpanda Connect API Role", "description": "Redpanda Connect API role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.secrets.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_connect_api_role --project= --file redpanda-connect-api.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-connect-api@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_connect_api_role" ``` ```bash # Account used to retrieve secrets and create Redpanda Connect pipelines. gcloud iam service-accounts create redpanda-connect \ --display-name="Redpanda Connect Service Account" cat << EOT > redpanda-connect.role { "name": "redpanda_connect_role", "title": "Redpanda Connect Role", "description": "Redpanda Connect role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_connect_role --project= --file redpanda-connect.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-connect@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_connect_role" ``` - Redpanda Cloud secret manager Show commands ```bash gcloud iam service-accounts create redpanda-console \ --display-name="Redpanda Cloud Secret Manager" cat << EOT > redpanda-console.role { "name": "redpanda_console_secret_manager_role", "title": "Redpanda Cloud Secret Manager Writer", "description": "Redpanda Cloud Secret Manager Writer", "includedPermissions": [ "secretmanager.secrets.get", "secretmanager.secrets.create", "secretmanager.secrets.delete", "secretmanager.secrets.list", "secretmanager.secrets.update", "secretmanager.versions.add", "secretmanager.versions.destroy", "secretmanager.versions.disable", "secretmanager.versions.enable", "secretmanager.versions.list", "iam.serviceAccounts.getAccessToken" ] } EOT gcloud iam roles create redpanda_console_secret_manager_role --project= --file redpanda-console.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-console@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_console_secret_manager_role" ``` - Kafka Connect service account Show commands ```bash gcloud iam service-accounts create redpanda-connectors \ --display-name="Kafka Connect Service Account" cat << EOT > redpanda-connectors.role { "name": "redpanda_connectors_role", "title": "Kafka Connect Custom Role", "description": "Kafka Connect custom role", "includedPermissions": [ "resourcemanager.projects.get", "secretmanager.versions.access" ] } EOT gcloud iam roles create redpanda_connectors_role --project= --file redpanda-connectors.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-connectors@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_connectors_role" ``` - Redpanda GKE cluster service account Show commands ```bash gcloud iam service-accounts create redpanda-gke \ --display-name="Redpanda GKE cluster default node service account" cat << EOT > redpanda-gke.role { "name": "redpanda_gke_utility_role", "title": "Redpanda cluster utility node role", "description": "Redpanda cluster utility node role", "includedPermissions": [ "artifactregistry.dockerimages.get", "artifactregistry.dockerimages.list", "artifactregistry.files.get", "artifactregistry.files.list", "artifactregistry.locations.get", "artifactregistry.locations.list", "artifactregistry.mavenartifacts.get", "artifactregistry.mavenartifacts.list", "artifactregistry.npmpackages.get", "artifactregistry.npmpackages.list", "artifactregistry.packages.get", "artifactregistry.packages.list", "artifactregistry.projectsettings.get", "artifactregistry.pythonpackages.get", "artifactregistry.pythonpackages.list", "artifactregistry.repositories.downloadArtifacts", "artifactregistry.repositories.get", "artifactregistry.repositories.list", "artifactregistry.repositories.listEffectiveTags", "artifactregistry.repositories.listTagBindings", "artifactregistry.repositories.readViaVirtualRepository", "artifactregistry.tags.get", "artifactregistry.tags.list", "artifactregistry.versions.get", "artifactregistry.versions.list", "logging.logEntries.create", "logging.logEntries.route", "monitoring.metricDescriptors.create", "monitoring.metricDescriptors.get", "monitoring.metricDescriptors.list", "monitoring.monitoredResourceDescriptors.get", "monitoring.monitoredResourceDescriptors.list", "monitoring.timeSeries.create", "cloudnotifications.activities.list", "monitoring.alertPolicies.get", "monitoring.alertPolicies.list", "monitoring.dashboards.get", "monitoring.dashboards.list", "monitoring.groups.get", "monitoring.groups.list", "monitoring.notificationChannelDescriptors.get", "monitoring.notificationChannelDescriptors.list", "monitoring.notificationChannels.get", "monitoring.notificationChannels.list", "monitoring.publicWidgets.get", "monitoring.publicWidgets.list", "monitoring.services.get", "monitoring.services.list", "monitoring.slos.get", "monitoring.slos.list", "monitoring.snoozes.get", "monitoring.snoozes.list", "monitoring.timeSeries.list", "monitoring.uptimeCheckConfigs.get", "monitoring.uptimeCheckConfigs.list", "opsconfigmonitoring.resourceMetadata.list", "resourcemanager.projects.get", "stackdriver.projects.get", "stackdriver.resourceMetadata.list", "dns.changes.create", "dns.changes.get", "dns.changes.list", "dns.managedZones.list", "dns.resourceRecordSets.create", "dns.resourceRecordSets.delete", "dns.resourceRecordSets.get", "dns.resourceRecordSets.list", "dns.resourceRecordSets.update", "secretmanager.versions.access", "stackdriver.resourceMetadata.write", "storage.objects.get", "storage.objects.list", "compute.instances.use", "iam.serviceAccounts.getAccessToken", "compute.regionNetworkEndpointGroups.create", "compute.regionNetworkEndpointGroups.delete", "compute.regionNetworkEndpointGroups.get", "compute.regionNetworkEndpointGroups.use", "compute.regionNetworkEndpointGroups.attachNetworkEndpoints", "compute.regionNetworkEndpointGroups.detachNetworkEndpoints" ] } EOT gcloud iam roles create redpanda_gke_utility_role --project= --file redpanda-gke.role gcloud projects add-iam-policy-binding \ --member="serviceAccount:redpanda-gke@.iam.gserviceaccount.com" \ --role="projects//roles/redpanda_gke_utility_role" ``` 4. Bind the service accounts. The account ID of the GCP service account is used to configure service account bindings. This account ID is the local part of the email address for the GCP service account. For example, if the GCP service account is `my-gcp-sa@my-project.iam.gserviceaccount.com`, then the account ID is `my-gcp-sa`. - Redpanda cluster service account Show command ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda/rp-]" ``` - Redpanda operator service account Show command ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-system/]" ``` - Redpanda Console service account Show command ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda/console-]" ``` - Redpanda Connect service accounts Show command ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-connect/]" ``` ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-connect/]" ``` - Kafka Connect service account Show command ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-connectors/connectors-]" ``` - Cert-manager and external-DNS service accounts Show commands ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[cert-manager/cert-manager]" gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[external-dns/external-dns]" ``` - Private Service Connect Controller service account Show commands ```bash gcloud iam service-accounts add-iam-policy-binding @.iam.gserviceaccount.com \ --role roles/iam.workloadIdentityUser \ --member "serviceAccount:.svc.id.goog[redpanda-psc/psc-controller]" ``` ## [](#create-cluster)Create cluster Log in to the [Redpanda Cloud UI](https://cloud.redpanda.com), and follow the steps to [create a BYOC cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/create-byoc-cluster-gcp/), with the following exceptions: 1. On the **Network** page, select the **BYOVPC** connection type, and enter the network, service account, storage bucket information, and GKE master CIDR range you created. 2. With customer-managed networks, you must grant yourself (the user deploying the cluster with `rpk`) the following permissions: Expand permissions - `compute.disks.create` - `compute.disks.setLabels` - `compute.instanceGroupManagers.create` - `compute.instanceGroupManagers.delete` - `compute.instanceGroupManagers.get` - `compute.instanceGroups.create` - `compute.instanceGroups.delete` - `compute.instanceTemplates.create` - `compute.instanceTemplates.delete` - `compute.instanceTemplates.get` - `compute.instanceTemplates.useReadOnly` - `compute.instances.create` - `compute.instances.setLabels` - `compute.instances.setMetadata` - `compute.instances.setTags` - `compute.subnetworks.get` - `compute.subnetworks.use` - `compute.zones.list` - `iam.roles.get` - `iam.serviceAccounts.actAs` - `iam.serviceAccounts.get` - `resourcemanager.projects.get` - `resourcemanager.projects.getIamPolicy` - `serviceusage.services.list` - `storage.buckets.get` - `storage.buckets.getIamPolicy` - `storage.objects.create` - `storage.objects.delete` - `storage.objects.get` - `storage.objects.list` This can be done through a Google account, a service account, or any principal identity supported by GCP. - If running `rpk` from a Google account, the user must acquire new user credentials to use for [Application Default Credentials](https://cloud.google.com/sdk/gcloud/reference/auth/application-default/login). - If running `rpk` from a service account, the user must create a [service account key](https://cloud.google.com/iam/docs/keys-create-delete#creating), then [export GOOGLE\_APPLICATION\_CREDENTIALS](https://cloud.google.com/docs/authentication/application-default-credentials#GAC) and [set the account as the default in gcloud](https://cloud.google.com/sdk/gcloud/reference/config/set): ```bash export GOOGLE_APPLICATION_CREDENTIALS= gcloud config set account $SERVICE_ACCOUNT@$PROJECT_ID.iam.gserviceaccount.com ``` 3. To validate your configuration, run: ```bash rpk cloud byoc gcp apply --redpanda-id='' --project-id='' --validate-only ``` 4. Click **Next**. 5. On the **Deploy** page, similar to standard BYOC clusters, log in to Redpanda Cloud and deploy the agent. > 📝 **NOTE** > > Redpanda Cloud does not support customer access or modifications to any of the internal data plane resources. This restriction allows Redpanda Data to manage all configuration changes internally to ensure a 99.99% service level agreement (SLA) for BYOC clusters. ## [](#delete-cluster)Delete cluster You can delete the cluster in the Cloud UI. 1. Log in to [Redpanda Cloud](https://cloud.redpanda.com). 2. Select your cluster. 3. Go to the **Dataplane settings** page and click **Delete**, then confirm your deletion. ## [](#manage-custom-resource-labels-and-network-tags)Manage custom resource labels and network tags Your organization might require custom resource labels and network tags for cost allocation, audit compliance, or governance policies. After cluster creation, you can manage labels and network tags on your cluster’s **Dataplane settings** page in the Redpanda Cloud UI, or with the [Control Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/). > ⚠️ **CAUTION** > > Do not add labels or network tags directly to the GCP node pools of a BYOC cluster, either in the GCP console or with commands such as `gcloud container node-pools update`. Google Kubernetes Engine (GKE) treats a node pool label or tag update as a node replacement and aggressively replaces all nodes, rather than performing a controlled rolling upgrade. This bypasses the failover process that Redpanda requires for safe node cycling. It can leave persistent volume claims (PVCs) in a pending state and may require recovering the entire cluster. > > The Dataplane settings page and the Control Plane API are the supported methods. With both, Redpanda applies labels directly to the Google Compute Engine (GCE) instances and disks in your cluster, and network tags directly to the instances. Because nothing is applied to the node pools, GKE does not replace any nodes. ### [](#use-the-dataplane-settings-page)Use the Dataplane settings page To manage labels and network tags in the Redpanda Cloud UI: 1. In the Redpanda Cloud UI, open your [cluster](https://cloud.redpanda.com/clusters), and click **Dataplane settings**. 2. Under **Manage resource labels and network tags for your cluster**, click **Add**, and enter a key and value for each label. Keys and values are case-sensitive. To apply a network tag to the GCE instances in the cluster, use a key that starts with `gcp.network-tag.`. For example, the key `gcp.network-tag.web-servers` applies the `web-servers` [network tag](https://cloud.google.com/vpc/docs/add-remove-network-tags). You can add up to 10 entries on this page. To manage up to 16, use the [Control Plane API](#use-the-control-plane-api). 3. Click **Save**. ### [](#use-the-control-plane-api)Use the Control Plane API The Control Plane API allows up to 16 custom resource labels and network tags in GCP. Make sure you have: - The cluster ID. You can find this in the Redpanda Cloud UI, in the **Details** section of the cluster overview. - A valid bearer token for the Control Plane API. For details, see [Authenticate to the API](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication). Then complete the following steps: 1. To refresh agent permissions so the Redpanda agent can update labels and network tags, run: ```bash export CLUSTER_ID="" export PROJECT_ID="" rpk cloud byoc gcp apply --redpanda-id="$CLUSTER_ID" --project-id="$PROJECT_ID" ``` This step is required because label/tag management requires additional IAM permissions that may not have been granted during initial cluster creation: - `compute.disks.get` - `compute.disks.list` - `compute.disks.setLabels` - `compute.instances.setLabels` 2. To update labels and network tags, invoke the Control Plane API. First, set your authentication token: ```bash export AUTH_TOKEN="" ``` The `PATCH` call sets the labels and network tags specified under `"cloud_provider_tags"`. It replaces the existing labels and tags with the specified labels and tags. Include all desired labels and tags in the request. To remove a single entry, omit it from the map you send. ```bash cluster_patch_body=$(cat <<'JSON' { "cloud_provider_tags": { "environment": "production", "cost-center": "engineering", "gcp.network-tag.web-servers": "true", "gcp.network-tag.database-access": "true" } } JSON ) curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` To remove all labels and network tags, send an empty `cloud_provider_tags` object: ```bash cluster_patch_body='{"cloud_provider_tags": {}}' curl -X PATCH "https://api.redpanda.com/v1/clusters/$CLUSTER_ID" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $AUTH_TOKEN" \ -d "$cluster_patch_body" ``` > 📝 **NOTE** > > For BYOVPC clusters, custom labels are not applied to the customer-managed resources that are deployed by the customer. ## [](#next-steps)Next steps - [Configure private networking](https://docs.redpanda.com/cloud-data-platform/networking/byoc/gcp/) - [Enable Redpanda SQL on a BYOVPC Cluster on GCP](https://docs.redpanda.com/cloud-data-platform/sql/get-started/enable-sql-byovpc-gcp/) --- # Page 371: Create Remote Read Replicas **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/remote-read-replicas.md --- # Create Remote Read Replicas > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Create Remote Read Replicas latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/byoc/remote-read-replicas page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/byoc/remote-read-replicas.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/byoc/remote-read-replicas.adoc description: Learn how to create a remote read replica topic with BYOC, which is a read-only topic that mirrors a topic on a different cluster. page-git-created-date: "2024-08-01" page-git-modified-date: "2026-07-06" --- A remote read replica topic is a read-only topic that mirrors a topic on a different cluster. You can create a separate remote cluster just for consumers of this topic and populate its topics from object storage. A read-only topic on a remote cluster can serve any consumer, without increasing the load on the source cluster. Because these read-only topics access data directly from object storage, there’s no impact to the performance of the cluster. Remote read replica topics do not store any data. When a cluster running a remote read replica is terminated, the topic data only exists on the origin cluster. Redpanda Cloud supports remote read replica topics in BYOC clusters on AWS or GCP. These clusters can be ephemeral; that is, created temporarily to handle specific or transient workloads, but they don’t have to be. The ability to make them ephemeral provides flexibility and cost efficiency: you can scale resources up or down as needed and pay only for what you use. > ❗ **IMPORTANT** > > Creating the remote read replica topic is only one of several required steps. To set up a remote read replica, you must complete all of the following: > > 1. **BYOVPC only**: Grant the reader cluster’s service account read access to the source cluster’s storage bucket. See [BYOVPC: Grant storage permissions](#byovpc-storage-permissions). > > 2. Link the reader cluster to the source cluster by setting `read_replica_cluster_ids`. See [Link the reader cluster to the source cluster](#link-clusters). Creating the topic alone does _not_ link the clusters. > > 3. Create the remote read replica topic. See [Create remote read replica topic](#create-remote-read-replica-topic). ## [](#prerequisites)Prerequisites To use remote read replicas, you need: - A BYOC reader cluster in Ready state. This separate reader cluster must exist in the same Redpanda organization as the source cluster. - AWS: The reader cluster must be in the same region and the same account as the source cluster. - GCP: The reader cluster can be in the same or a different region as the source cluster. The reader cluster must be in the same project as the source cluster. - Azure: Remote read replicas are not supported. ### [](#byovpc-storage-permissions)BYOVPC: Grant storage permissions > 📝 **NOTE** > > This prerequisite only applies to BYOVPC deployments. Skip this step if you’re enabling remote read replicas on standard BYOC clusters. #### GCP To grant additional permissions to the cloud storage manager of the reader cluster, run: ```bash gcloud storage buckets add-iam-policy-binding \ gs:// \ --member="serviceAccount:" \ --role="roles/storage.objectViewer" ``` #### AWS To grant additional permissions to the cloud storage manager of the reader cluster, set the `source_cluster_bucket_names` and `reader_cluster_id` variables in [cloud-examples](https://github.com/redpanda-data/cloud-examples/blob/main/customer-managed/aws/terraform/variables.tf). This should be done in the Terraform of the reader cluster. If the reader cluster’s service account does not have read access to the source cluster’s storage bucket, creating the remote read replica topic fails with an error like `UNKNOWN_SERVER_ERROR: Unable to perform requested topic operation`. ## [](#link-clusters)Link the reader cluster to the source cluster Linking is a required step. It does more than record the relationship between the clusters: - It orders upgrades safely across the linked clusters. - It prevents the source cluster from being deleted while it is linked to a reader cluster. The way linking interacts with storage permissions depends on your deployment type: - Standard BYOC: Linking the clusters grants the reader cluster read access to the source cluster’s storage bucket, so you cannot create a remote read replica topic until the clusters are linked. - BYOVPC: You grant the reader cluster read access to the bucket manually (see [BYOVPC: Grant storage permissions](#byovpc-storage-permissions)), so creating the topic succeeds even when the clusters are not linked. This makes the linking step easy to miss. > ⚠️ **WARNING** > > Creating a remote read replica topic does _not_ link the reader cluster to the source cluster. You must also set `read_replica_cluster_ids` as described in this section. Without this link, upgrades are not ordered safely across the clusters, and the source cluster is not protected from deletion, even if remote read replica topics exist. Add or remove reader clusters to a source cluster in Redpanda Cloud with the [Cloud Control Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/controlplane/). For information on accessing the Cloud API, see the [authentication guide](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication). 1. To update your source cluster to add one or more reader cluster IDs, make a [`PATCH /v1/clusters/{cluster.id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) request. The full list of clusters is expected on every call. If an ID is removed from the list, it is removed as a reader cluster. ```bash export SOURCE_CLUSTER_ID=....... export READER_CLUSTER_ID=....... curl -X PATCH $API_HOST/v1/clusters/$SOURCE_CLUSTER_ID \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $API_TOKEN" \ -d @- << EOF { "read_replica_cluster_ids": ["$READER_CLUSTER_ID"] } EOF ``` 2. Optional: To see the list of reader clusters on a given source cluster, make a [`GET /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_getcluster) request: ```bash export SOURCE_CLUSTER_ID=....... curl -X GET $API_HOST/v1/clusters/$SOURCE_CLUSTER_ID \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $API_TOKEN" ``` > 📝 **NOTE** > > A source cluster cannot be deleted while any reader cluster IDs are set in its `read_replica_cluster_ids` list. This deletion protection is based on the link, not on the existence of remote read replica topics. When you delete a reader cluster, that cluster’s ID is automatically removed from any existing source cluster `read_replica_cluster_ids` lists. ## [](#create-remote-read-replica-topic)Create remote read replica topic To create a remote read replica topic, run: ```bash rpk topic create -c redpanda.remote.readreplica= --tls-enabled ``` - For ``, use the same name as the original topic. - For ``, use the bucket specified in the `cloud_storage_bucket` properties for the origin cluster. For standard BYOC clusters, the source cluster bucket name follows the pattern: `redpanda-cloud-storage-${SOURCE_CLUSTER_ID}` ## [](#optional-tune-for-live-topics)Optional: Tune for live topics For remote read replicas reading from a live topic (that is, a topic that’s being actively written to by a source cluster), it may be advantageous to control how often segments are flushed to object storage. By default, this is set to 60 minutes. To tune `cloud_storage_segment_max_upload_interval_sec` on the source cluster, contact [Redpanda support](https://support.redpanda.com/hc/en-us/requests/new). (For cold topics, where segments are closed and older than 60 minutes, this configuration is unnecessary: the data is already uploaded to object storage.) --- # Page 372: Dedicated **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster.md --- # Dedicated > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Dedicated latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/create-dedicated-cloud-cluster page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/create-dedicated-cloud-cluster.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/create-dedicated-cloud-cluster.adoc description: Learn how to create a Dedicated cluster and start streaming. page-git-created-date: "2025-04-01" page-git-modified-date: "2026-06-04" --- After you log in to [Redpanda Cloud](https://cloud.redpanda.com), you land on the **Clusters** page. This page lists all the clusters in your organization. ## [](#create-a-dedicated-cluster)Create a Dedicated cluster 1. On the Clusters page, click **Create cluster**, then click **Create** for Dedicated. Enter a cluster name, then select the resource group, cloud provider (AWS, GCP, or Azure), [region, tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/dedicated-tiers/), availability, and Redpanda version. > ❗ **IMPORTANT** > > Dedicated on Azure is in [limited availability](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#limited-availability). It is production-ready and covered by Redpanda Support for early adopters. > 📝 **NOTE** > > - If you plan to create a private network in your own VPC, select the region where your VPC is located. > > - Three availability zones provide two backups in case one availability zone goes down. 2. Click **Next**. 3. On the Network page, enter the connection type: public or private. For private networks: - Your network name is used to identify this network. - For a [CIDR range](https://docs.redpanda.com/cloud-data-platform/networking/cidr-ranges/), choose one that does not overlap with your existing VPCs or your Redpanda network. Private networks require either a VPC peering connection or a private connectivity service, such as [AWS PrivateLink](https://docs.redpanda.com/cloud-data-platform/networking/configure-privatelink-in-cloud-ui/), [GCP Private Service Connect](https://docs.redpanda.com/cloud-data-platform/networking/configure-private-service-connect-in-cloud-ui/), or [Azure Private Link](https://docs.redpanda.com/cloud-data-platform/networking/azure-private-link/). - Clusters with private networking include a setting for API Gateway network access. Public access exposes endpoints for Redpanda Console, the Data Plane API, and the MCP Server API, but they remain protected by your authentication and authorization controls. Private access restricts endpoint access to your VPC/VNet only. On Azure, private access incurs an additional cost, since it involves deploying two network load balancers, instead of one. > 📝 **NOTE** > > After the cluster is created, you can change the API Gateway access on the Dataplane settings page. If you change from public to private access, users without VPN access to the Redpanda VPC will lose access to these services. 4. Click **Create**. After the cluster is created, you can select the cluster on the **Clusters** page to see the overview for it. ## [](#start-streaming-example)Start streaming: example Use `rpk`, Redpanda’s CLI, to build a basic streaming application that creates a topic, produces messages to it, and consumes messages from it. To learn about `rpk`, see the [Introduction to rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/intro-to-rpk/). 1. Login to Redpanda Cloud, and select your resource group using the interactive prompt. ```bash rpk cloud login ``` 2. On the **Overview** page, copy your bootstrap server address and set it as an environment variable on your local machine: ```bash export REDPANDA_BROKERS="" ``` 3. Go to **Security** > **Users**, click **Create user**, and create a user called **redpanda-chat-account** that uses the SCRAM-SHA-256 mechanism. 4. In the **User created successfully** dialog, copy the password and set the following environment variables on your local machine: ```bash export REDPANDA_SASL_USERNAME="redpanda-chat-account" export REDPANDA_SASL_PASSWORD="" export REDPANDA_SASL_MECHANISM="SCRAM-SHA-256" ``` 5. Click **Go to user details**. 6. Under **ACLs**, click **\+ Add ACL**, and define the following rule to grant the user full access to the `chat-room` topic: - **Resource Type**: Topic - **Pattern Type**: Literal - **Resource Name**: `chat-room` - **Operation**: All - **Permission**: Allow - **Host**: `*` 7. Click **Add ACL**. 8. Use `rpk` on your local machine to authenticate to Redpanda as the **redpanda-chat-account** user and get information about the cluster: ```bash rpk cluster info -X tls.enabled=true ``` 9. Create a topic called `chat-room`. You granted permissions to the **redpanda-chat-account** user to access only this topic. ```bash rpk topic create chat-room -X tls.enabled=true ``` Output: TOPIC STATUS chat-room OK 10. Produce a message to the topic: ```bash rpk topic produce chat-room -X tls.enabled=true ``` 11. Enter a message, then press Enter: ```text Pandas are fabulous! ``` Example output: Produced to partition 0 at offset 0 with timestamp 1663282629789. 12. Press Ctrl+C to finish producing messages to the topic. 13. Consume one message from the topic: ```bash rpk topic consume chat-room --num 1 -X tls.enabled=true ``` Your message is displayed along with its metadata: ```json { "topic": "chat-room", "value": "Pandas are fabulous!", "timestamp": 1663282629789, "partition": 0, "offset": 0 } ``` ### [](#explore-your-topic)Explore your topic In Redpanda Cloud, go to **Topics** > **chat-room**. The message that you produced to the topic is displayed along with some other details about the topic. ### [](#clean-up)Clean up If you don’t want to continue experimenting with your cluster, you can delete it. Go to **Dataplane settings** and click **Delete cluster**. ## [](#next-steps)Next steps - [Learn more about Redpanda Cloud](https://docs.redpanda.com/cloud-data-platform/get-started/cloud-overview/) - [Learn about private networking](https://docs.redpanda.com/cloud-data-platform/networking/dedicated/) --- # Page 373: Serverless **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless.md --- # Serverless > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Serverless latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-types/serverless page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-types/serverless.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/cluster-types/serverless.adoc description: Learn how to create a Serverless cluster and start streaming. page-topic-type: overview personas: evaluator, app_developer, platform_admin learning-objective-1: Identify the use cases and usage limits for Serverless clusters learning-objective-2: Describe how to create a Serverless cluster and connect a client learning-objective-3: Recognize which features are supported and unsupported on Serverless page-git-created-date: "2024-06-06" page-git-modified-date: "2026-08-27" --- Serverless is the fastest and easiest way to start data streaming. With Serverless clusters, you host your data in Redpanda’s VPC, and Redpanda handles automatic scaling, provisioning, operations, and maintenance. This is a production-ready deployment option with a cluster available instantly, and you only pay for what you consume. You can view detailed billing activity for each cluster and edit payment methods on the **Billing** page. After reading this page, you will be able to: - Identify the use cases and usage limits for Serverless clusters - Describe how to create a Serverless cluster and connect a client - Recognize which features are supported and unsupported on Serverless ## [](#serverless-usage-limits)Serverless usage limits Each Serverless cluster has the following maximum usage limits: - **Ingress**: 100 MB/s - **Egress**: 300 MB/s - **Partitions**: 5,000 - **Message size**: 20 MiB - **Retention**: unlimited - **Storage**: unlimited - **Users**: 30 - **ACLs**: 120 - **Consumer groups**: 200 - **Connections**: 10,000 - **Producer IDs**: 250 - **Schema Registry**: - **Max schemas**: 500 - **Max subjects**: 500 - **Rate limit**: 100 requests/s - **Redpanda Connect pipelines**: 100 > 📝 **NOTE** > > The partition limit is the number of logical partitions before replication occurs. Redpanda Cloud uses a replication factor of 3. ## [](#prerequisites)Prerequisites Make sure you have the latest version of `rpk`, the Redpanda CLI. See [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). ## [](#get-started-with-serverless)Get started with Serverless > 📝 **NOTE** > > Serverless on GCP is currently in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release. Choose the option that fits how you want to subscribe: ### Free trial A [free trial on AWS](https://www.redpanda.com/try-redpanda) is the fastest way to get started with Serverless. Each free-trial customer qualifies for $100 (USD) in credits to spend in the first 30 days. This should be enough to run Redpanda with reasonable throughput. No credit card is required. To continue using Serverless after your trial expires, you can enter a credit card and pay as you go. Any remaining credit balance is used before you are charged. When either the credits expire or the days in the trial expire, the clusters move into a suspended state, and you won’t be able to access your data in either the Redpanda Cloud Console or with the Kafka API. There is a seven-day grace period following the end of the trial when you can add your credit card and restore service. After that, the data is permanently deleted. For questions about the trial, use the **#serverless** [Community Slack](https://redpandacommunity.slack.com/) channel. After you start a trial, Redpanda instantly prepares an account for you. The first time you sign in, you can answer a few quick questions about your project so Redpanda can tailor your experience. Your account includes a `welcome` cluster with a `hello-world` demo topic you can explore. It includes sample data so you can see how real-time messaging works before sending your own data. On that first visit, the **Overview** page shows a **Get started** button with guided ways to [interact with your cluster](#interact-with-your-cluster): create a Redpanda Connect [pipeline](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#pipeline), use `rpk` from the command line, or connect with your own Kafka client. To get started with `rpk`: 1. Log in with `rpk cloud login`. 2. Consume from the `hello-world` topic with `rpk topic consume hello-world`. 3. In the [Redpanda Cloud Console](https://cloud.redpanda.com), navigate to the **Topics** page and open the `hello-world` topic to see the included messages. ### Redpanda Sales To request a private offer with possible discounts for annual committed use, contact [Redpanda Sales](https://www.redpanda.com/price-estimator). When you subscribe to Serverless through Redpanda Sales, you gain immediate access to Enterprise support. Redpanda creates a cloud organization for you and sends you a welcome email. ### AWS Marketplace New subscriptions to Redpanda Cloud through [AWS Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/aws-pay-as-you-go/) receive $300 (USD) in free credits to spend in the first 30 days. AWS Marketplace charges for anything beyond $300, unless you cancel the subscription. After your free credits have been used, you can continue using your cluster without any commitment, only paying for what you consume and canceling anytime. > 📝 **NOTE** > > When you subscribe to Redpanda through AWS Marketplace, you do not have immediate access to Enterprise support, only the [Community Slack](https://redpandacommunity.slack.com/) channel. For Enterprise support, contact [Redpanda Sales](https://www.redpanda.com/price-estimator). Redpanda creates a cloud organization for you and sends you a welcome email. ### Google Cloud Marketplace New subscriptions to Redpanda Cloud through [Google Cloud Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/gcp-pay-as-you-go/) receive $300 (USD) in free credits to spend in the first 30 days. Google Cloud Marketplace charges for anything beyond $300, unless you cancel the subscription. After your free credits have been used, you can continue using your cluster without any commitment, only paying for what you consume and canceling anytime. > 📝 **NOTE** > > When you subscribe to Redpanda through Google Cloud Marketplace, you do not have immediate access to Enterprise support, only the [Community Slack](https://redpandacommunity.slack.com/) channel. For Enterprise support, contact [Redpanda Sales](https://www.redpanda.com/price-estimator). Redpanda creates a cloud organization for you and sends you a welcome email. ## [](#create-a-serverless-cluster)Create a Serverless cluster To create a Serverless cluster: 1. In the [Redpanda Cloud Console](https://cloud.redpanda.com), on the **Clusters** page, click **Create cluster**, then click **Create** for Serverless. 2. Enter a cluster name, then select the resource group. If you don’t have an existing resource group, you can create one. Refresh the page to see newly-created resource groups. 3. Select a cloud provider: AWS or GCP. (GCP is currently in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release.) 4. Select a [region](https://docs.redpanda.com/cloud-data-platform/reference/tiers/serverless-regions/). For best performance, select the region closest to your applications. Redpanda expects your applications to be deployed in the same cloud provider and region as your Serverless cluster. 5. **AWS only**: Clusters on AWS can enable private access between their VPC and Redpanda, so data does not traverse the public internet. Private connectivity is implemented using AWS PrivateLink for secure traffic. - When you enable both public access and private access on the cluster, you can choose between the public address or the private address. When the public address is used the data flows over the public internet. - You can either create a new PrivateLink or use an existing one from the same resource group. - You can enable or disable private access at any time on the cluster’s **Dataplane settings** page. - Enabling private access incurs additional charges. > 📝 **NOTE** > > After private access is disabled, attempts to reach the private endpoints will fail. However, the PrivateLink endpoint in your AWS account and the PrivateLink resource in Redpanda Cloud both remain provisioned and continue to incur charges until you explicitly delete them. 6. Click **Create cluster**. 7. To start working with your cluster, go to the **Topics** page to create a topic and produce messages to it. Add team members on the **Security** > **Users** page, then click into a user to add ACLs from their detail page. ## [](#interact-with-your-cluster)Interact with your cluster > 💡 **TIP** > > The cluster’s **Overview** page includes a **Get started** guide to help you start streaming data into and out of Redpanda. See also: [Redpanda Connect Quickstart](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/) The **Overview** page lists your bootstrap server URL and security settings in the **How to connect - Kafka API** tab. Here you can add a Kafka client to interact with your cluster. Or, Redpanda can generate a sample application to interact with your cluster. Run [`rpk generate app`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-generate/rpk-generate-app/), and select Go as the language. Follow the commands in the terminal to run the application, create a demo topic, produce to the topic, and consume the data back. The first time you sign in, the **Overview** page shows a **Get started** button whose **Connect from your terminal** option walks you through using `rpk`. You can also use these `rpk` commands at any time: - [`rpk cloud login`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-login/): Use this to log in to Redpanda Cloud or to refresh the session. - [`rpk topic`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-topic/rpk-topic/): Use this to manage topics, produce data, and consume data. - [`rpk profile print`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-profile/rpk-profile-print/): Use this to view your `rpk` configuration and see the URL for your Serverless cluster. - [`rpk security user`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-security/rpk-security-user/): Use this to manage users and permissions. > 📝 **NOTE** > > Redpanda Serverless is opinionated about Kafka configurations. For example, automatic topic creation is disabled. Some systems expect the Kafka service to automatically create topics when a message is produced to a topic that doesn’t exist. Create topics on the **Topics** page or with `rpk topic create`. ## [](#supported-features)Supported features - Redpanda Serverless supports the Kafka API. Serverless clusters work with all Kafka clients. See [Kafka Compatibility](https://docs.redpanda.com/cloud-data-platform/develop/kafka-clients/). - Serverless clusters support all major Apache Kafka messages for managing topics, producing/consuming data (including transactions), managing groups, managing offsets, and managing ACLs. (User management is available in the [Redpanda Cloud Console](https://cloud.redpanda.com) or with `rpk security acl`.) ### [](#unsupported-features)Unsupported features Not all features included in BYOC clusters are available in Serverless. For example, the following features are not supported: - HTTP Proxy API - Multiple availability zones (AZs) - Role-based access control (RBAC) in the data plane and mTLS authentication for Kafka API clients - Group-based access control (GBAC) - Kafka Connect - Configurable maintenance windows ## [](#maintenance-and-upgrades)Maintenance and upgrades Redpanda manages all maintenance for Serverless clusters. Because Serverless runs on shared, multi-tenant infrastructure, you cannot configure a maintenance window or schedule upgrades for an individual cluster. Redpanda may run maintenance operations on Serverless clusters at any time. Continuous operations are integral to keeping Serverless clusters up to date and secure. Operations run in a rolling fashion and are designed to be non-disruptive. Mainstream Kafka client libraries reconnect automatically when broker connections restart. If you need control over when maintenance runs on your cluster, use a Dedicated or BYOC cluster, both of which support configurable maintenance windows. For more information, see [Upgrades and Maintenance](https://docs.redpanda.com/cloud-data-platform/manage/maintenance/). ## [](#next-steps)Next steps - [Set up private access for Serverless clusters](https://docs.redpanda.com/cloud-data-platform/networking/serverless/aws/) - [Manage Redpanda Cloud with Terraform](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) - [Learn more about Redpanda Cloud](https://docs.redpanda.com/cloud-data-platform/get-started/cloud-overview/) - [Manage topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/config-topics/) - [Learn about billing](https://docs.redpanda.com/cloud-data-platform/billing/billing/) --- # Page 374: Introduction to Redpanda **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/intro-to-events.md --- # Introduction to Redpanda > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Introduction to Redpanda latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: intro-to-events page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: intro-to-events.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/intro-to-events.adoc description: Learn about Redpanda event streaming. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- Distributed systems often require data and system updates to happen as quickly as possible. In software architecture, these updates can be handled with either messages or events. - With messages, updates are sent directly from one component to another to trigger an action. - With events, updates indicate that an action occurred at a specific time, and are not directed to a specific recipient. An event is simply a record of something changing state. For example, the event of a credit card transaction includes the product purchased, the payment, the delivery, and the time of the purchase. The event occurred in the purchasing component, but it also impacted the inventory, the payment processing, and the shipping components. In an event-driven architecture, all actions are defined and packaged as events to precisely identify individual actions and how they’re processed throughout the system. Instead of processing updates in consecutive order, event-driven architecture lets components process events at their own pace. This helps developers build fast and scalable systems. ## [](#what-is-redpanda)What is Redpanda? Redpanda is an event streaming platform: it provides the infrastructure for streaming real-time data. Producers are client applications that send data to Redpanda in the form of events. Redpanda safely stores these events in sequence and organizes them into topics, which represent a replayable log of changes in the system. Consumers are client applications that subscribe to Redpanda topics to asynchronously read events. Consumers can store, process, or react to the events. Redpanda decouples producers from consumers to allow for asynchronous event processing, event tracking, event manipulation, and event archiving. Producers and consumers interact with Redpanda using the Apache Kafka® API. ![Producers and consumers in a cluster](https://docs.redpanda.com/cloud-data-platform/shared/_images/cluster.png) | Event-driven architecture (Redpanda) | Message-driven architecture | | --- | --- | | Producers send events to an event processing system (Redpanda) that acknowledges receipt of the write. This guarantees that the write is durable within the system and can be read by multiple consumers. | Producers send messages directly to each consumer. The producer must wait for acknowledgement that the consumer received the message before it can continue with its processes. | Event streaming lets you extract value out of each event by analyzing, mining, or transforming it for insights. You can: - Take one event and consume it in multiple ways. - Replay events from the past and route them to new processes in your application. - Run transformations on the data in real-time or historically. - Integrate with other event processing systems that use the Kafka API. ## [](#redpanda-differentiators)Redpanda differentiators Redpanda is less complex and less costly than any other commercial mission-critical event streaming platform. It’s fast, it’s easy, and it keeps your data safe. - Redpanda is designed for maximum performance on any data streaming workload. It can scale up to use all available resources on a single machine and scale out to distribute performance across multiple nodes. Built on C++, Redpanda delivers greater throughput and up to 10x lower p99 latencies than other platforms. This enables previously unimaginable use cases that require high throughput, low latency, and a minimal hardware footprint. - Redpanda is packaged as a single binary: it doesn’t rely on any external systems. It’s compatible with the Kafka API, so it works with the full ecosystem of tools and integrations built on Kafka. Redpanda can be deployed on bare metal, containers, or virtual machines in a data center or in the cloud. And Redpanda Console makes it easy to set up, manage, and monitor your clusters. Additionally, Tiered Storage lets you offload log segments to object storage in near real-time, providing long-term data retention and topic recovery. - Redpanda uses the [Raft consensus algorithm](https://raft.github.io/) throughout the platform to coordinate writing data to log files and replicating that data across multiple servers. Raft facilitates communication between the nodes in a Redpanda cluster to make sure that they agree on changes and remain in sync, even if a minority of them are in a failure state. This allows Redpanda to tolerate partial environmental failures and deliver predictable performance, even at high loads. - Redpanda provides data sovereignty. With the Bring Your Own Cloud (BYOC) offering, you deploy Redpanda in your own virtual private cloud, and all data is contained in your environment. Redpanda handles provisioning, monitoring, and upgrades, but you manage your streaming data without Redpanda’s control plane ever seeing it. ## [](#redpanda-streaming-versions)Redpanda Streaming versions You can deploy Redpanda in a self-hosted environment (Redpanda Streaming) or as a fully managed cloud service (Redpanda Cloud). Redpanda Streaming version numbers follow the convention AB.C.D, where AB is the two-digit year, C is the feature release, and D is the patch release. For example, version 22.3.1 indicates the first patch release on the third feature release of the year 2022. Patch releases include bug fixes and minor improvements, with no change to user-facing behavior. New and enhanced features are documented with each feature release. Redpanda Cloud releases on a continuous basis and uptakes Redpanda Streaming versions. --- # Page 375: Partner Integrations **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/partner-integration.md --- # Partner Integrations > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Partner Integrations latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: partner-integration page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: partner-integration.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/partner-integration.adoc description: Learn about Redpanda integrations built and supported by our partners. page-git-created-date: "2024-07-25" page-git-modified-date: "2026-05-26" --- Learn about Redpanda integrations built and supported by our partners. | Partner | Description | More information | | --- | --- | --- | | Superstream | Superstream optimizes and improves Redpanda (and other Kafka platforms) for cost reduction, increased reliability, and improved visibility. | Superstream for Redpanda | | Aklivity Zilla | Zilla is a multi-protocol proxy that abstracts Redpanda for non-native clients, such as browsers and IoT devices, by exposing Redpanda topics using user-defined REST, Server-Sent Events (SSE), MQTT, or gRPC API entry points. | Modern Eventing with CQRS, Redpanda and Zilla | | Bytewax | Bytewax is an open source framework and distributed stream processing engine in Python. | Enriching streaming data with Bytewax and Redpanda | | ClickHouse | ClickHouse is a high-performance, column-oriented SQL database management system (DBMS) for online analytical processing (OLAP). | Building an OLAP database with ClickHouse and Redpanda | | Conduktor | Conduktor provides simple, flexible, and powerful tooling for Kafka developers and infrastructure. | Conduktor & Redpanda: Best of breed Kafka experience | | Decodable | Decodable is a real-time data processing platform powered by Apache Flink and Debezium. | Decodable + Redpanda | | ElastiFlow | ElastiFlow captures and analyzes flow and SNMP data to provide detailed insights into network performance and security. | Leveraging Redpanda for Enhanced Network Observability: ElastiFlow Integration | | Materialize | Materialize is a data warehouse purpose-built for operational workloads where an analytical data warehouse would be too slow, and a stream processor would be too complicated. | Ingesting data from Redpanda with Materialize | | PeerDB | PeerDB provides a fast, simple, and cost-effective way to replicate data from Postgres to data warehouses, queues and storage. | Quickstart guide | | Pinecone | Pinecone is a vector database for building accurate and performant AI applications at scale. The Pinecone connector for Redpanda Connect provides a production-ready integration from many existing data sources through simple YAML configuration. | Redpanda Connect integration | | RisingWave | RisingWave is a distributed SQL streaming database that enables simple, efficient, and reliable processing of streaming data. | Ingesting data from Redpanda with Risingwave | | Timeplus | Timeplus is a stream processor that provides powerful end-to-end capabilities, leveraging the open source streaming engine Proton. | Realizing low latency streaming analytics with Timeplus and Redpanda | | Tinybird | Tinybird is a data platform for data and engineering teams to solve complex real-time, operational, and user-facing analytics use cases at any scale. | Building a complete IoT backend with Redpanda and Tinybird | | Quix | Quix is a complete platform for building, deploying, and monitoring stream processing pipelines in Python. | Integrating Redpanda with Quix | | Yugabyte | YugabyteDB is an open-source, distributed SQL database that combines the capabilities of relational databases with the scalability of NoSQL systems. | How to Integrate Yugabyte CDC Connector with Redpanda | ## [](#how-to-contribute-to-this-page)How to contribute to this page To request a partner integration with Redpanda Data, reach out to ([partners@redpanda.com](mailto:partners@redpanda.com\)). Provide a link to your product documentation or a blogpost explaining how your product integrates with Redpanda. After meeting these requirements, you can [contribute to this page](https://github.com/redpanda-data/docs/edit/main/modules/get-started/pages/partner-integration.adoc). --- # Page 376: What’s New in Redpanda Cloud **URL**: https://docs.redpanda.com/cloud-data-platform/get-started/whats-new-cloud.md --- # What’s New in Redpanda Cloud > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: What’s New in Redpanda Cloud latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: whats-new-cloud page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: whats-new-cloud.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/get-started/pages/whats-new-cloud.adoc description: Summary of new features in Redpanda Cloud. page-git-created-date: "2024-06-06" page-git-modified-date: "2026-08-25" --- This page lists new features added to Redpanda Cloud. ## [](#august-2026)August 2026 ### [](#migrate-schemas-from-confluent-schema-registry)Migrate schemas from Confluent Schema Registry BYOC and Dedicated clusters can now use Shadowing’s API-mode Schema Registry replication to migrate schemas from a Confluent Schema Registry. A shadow link polls the source registry over HTTP and continuously imports subjects, versions, and compatibility settings, with filtering by context or subject and source-to-destination context mapping. While the link is active, the destination contexts it owns are read-only; pause replication to cut applications over to Redpanda. Configure the link in the Cloud UI, or with `rpk`, the Control Plane API, or the Terraform provider. See [Migrate Schemas from Confluent Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). ### [](#fetch-read-coalescing)Fetch read coalescing BYOC and Dedicated clusters now support [fetch read coalescing](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/fetch-read-coalescing/), which reduces read CPU and fetch-response memory for high consumer fan-out workloads: the broker reads and serializes each unique read once and shares the result with every concurrent consumer of the same data. Fetch read coalescing is disabled by default, and the cluster property that controls it is managed by Redpanda. Request it only for high fan-out workloads. ### [](#redpanda-sql-on-gcp)Redpanda SQL on GCP Redpanda SQL is now available on GCP, for both BYOC and BYOVPC clusters. Run real-time SQL queries on your Redpanda topics, including the Iceberg history of Iceberg-enabled topics through a GCP Lakehouse (formerly BigLake) catalog, using standard PostgreSQL syntax. - On a BYOC cluster, enable the SQL engine from the Cloud Console, Cloud API, or Terraform provider. See [Enable Redpanda SQL](https://docs.redpanda.com/cloud-data-platform/sql/get-started/deploy-sql-cluster/). - On a BYOVPC cluster, provision the SQL-specific GCP resources with the Redpanda BYOVPC Terraform module, then supply them as customer-managed resources when enabling the engine. See [Enable Redpanda SQL on a BYOVPC Cluster on GCP](https://docs.redpanda.com/cloud-data-platform/sql/get-started/enable-sql-byovpc-gcp/). ### [](#redpanda-sql-byovpc-support-on-aws)Redpanda SQL: BYOVPC support on AWS Redpanda SQL is now available on BYOVPC clusters on AWS. Provision the required SQL-specific AWS resources using the Redpanda BYOVPC Terraform module, then supply them as customer-managed resources when enabling the SQL engine via the Cloud Console, Cloud API, or Terraform provider. See [Enable Redpanda SQL on a BYOVPC Cluster on AWS](https://docs.redpanda.com/cloud-data-platform/sql/get-started/enable-sql-byovpc-aws/). ### [](#translate-iceberg-keys-values-and-headers-and-query-them-with-sql)Translate Iceberg keys, values, and headers, and query them with SQL On BYOC and BYOVPC clusters with Iceberg-enabled topics, you can now control how Redpanda translates record keys, values, and headers into the Iceberg table — including decoding keys and header values using a schema or as UTF-8 strings instead of the default raw bytes — and read the decoded keys and headers with Redpanda SQL. - Use the section-based syntax of the `redpanda.iceberg.mode` topic property to independently control how Redpanda translates the record key, value, and headers into the Iceberg table. See [Configure key, value, and header translation](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/#configure-key-value-and-header-translation). - Set `key_decode_mode` and `header_value_type` on `CREATE TABLE`, matching the topic’s translation, to expose the decoded key and header values when you query the topic with Redpanda SQL. See [Query decoded keys and headers](https://docs.redpanda.com/cloud-data-platform/sql/query-data/query-iceberg-topics/#query-decoded-keys-and-headers). ### [](#redpanda-sql-read-uuid-columns-as-the-uuid-type)Redpanda SQL: read UUID columns as the `uuid` type Redpanda SQL now reads UUID columns as the `uuid` type, exposed to PostgreSQL clients as `uuid`. This applies to UUID columns in a Redpanda topic, whether in the topic’s Avro-encoded live records or in its history committed to an Iceberg table. Because text and string functions don’t operate on `uuid` values directly, cast a UUID to text explicitly (`id::text`) to get the 36-character canonical form. Previously, these columns were read as `text` (and returned garbled values when read from Iceberg), so the switch to `uuid` can change how queries that relied on the old `text` columns behave. If a query breaks, cast the column to text (`id::text`) to restore the previous behavior. See [UUID](https://docs.redpanda.com/cloud-data-platform/reference/sql/sql-data-types/uuid/) and [Query Iceberg-enabled topics](https://docs.redpanda.com/cloud-data-platform/sql/query-data/query-iceberg-topics/). ### [](#resolve-iceberg-topic-schemas-within-a-schema-registry-context)Resolve Iceberg topic schemas within a Schema Registry context On BYOC and BYOVPC clusters, you can now set the `redpanda.schema.registry.context` topic property on an Iceberg-enabled topic to resolve its schemas within a specific [Schema Registry context](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/) instead of the default context. See [Resolve schemas within a Schema Registry context](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/#resolve-schemas-within-a-context). ### [](#expanded-kafka-client-validation)Expanded Kafka client validation Redpanda now validates additional non-Java Kafka clients at their current versions, aligned with Kafka 4.x: confluent-kafka-go, Sarama, and confluent-kafka-python. See [Kafka Compatibility](https://docs.redpanda.com/cloud-data-platform/develop/kafka-clients/). ### [](#visual-tab-in-the-redpanda-connect-pipeline-editor)Visual tab in the Redpanda Connect pipeline editor The pipeline editor has a new **Visual** tab alongside the existing **YAML** editor. The **Visual** tab renders a pipeline as an interactive node diagram, so you can inspect and edit inputs, processors, outputs, and control-flow constructs (`switch`, `branch`, `try`/`catch`, and others) without hand-writing YAML. The two tabs stay in sync. See [Explore the Visual editor](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/#visual-editor). ## [](#july-2026)July 2026 ### [](#schema-registry-contexts-enabled-by-default)Schema Registry contexts enabled by default On BYOC and Dedicated clusters, [Schema Registry contexts](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/) are now enabled by default, so you can register and manage schemas in isolated namespaces without an enablement step. See [Upgrade considerations](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/#upgrade-considerations) if any of your existing subject names start with `:.`. ### [](#serdes-client-support-for-schema-registry-contexts)SerDes client support for Schema Registry contexts Any SerDes client, in any language, can now target a Schema Registry context through the base URL alone: append `/contexts/{context}` to your cluster’s Schema Registry URL in the client’s `schema.registry.url` setting. Schema ID lookups sent through the prefix are scoped to that context automatically. Previously, only the Java Confluent SerDes could target non-default contexts. See [Use contexts with SerDes clients](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/#serdes-clients). ### [](#egress-allowlist-for-redpanda-connect-pipelines)Egress allowlist for Redpanda Connect pipelines You can now allow Redpanda Connect pipelines on BYOC and Dedicated clusters to open outbound connections to destinations that the data plane firewall blocks by default. Add up to 16 CIDR and port range combinations to a cluster with the `redpanda_connect.allowed_destination_cidr_ports` field in the Cloud API, or with the same attribute on a `redpanda_cluster` resource in the Redpanda Terraform provider (v2.1.1+). See [Configure Egress for Redpanda Connect Pipelines](https://docs.redpanda.com/cloud-data-platform/networking/connect-egress-allowlist/). ### [](#cloud-topics-enabled-by-default)Cloud Topics enabled by default [Cloud Topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/cloud-topics/) are now enabled by default: create a Cloud Topic by setting its storage mode during topic creation, with no cluster-level enablement step. You can use Cloud Topics exclusively or together with standard topics on the same cluster. ### [](#new-aws-regions-for-byoc)New AWS regions for BYOC For [BYOC clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/#byoc-supported-regions), Redpanda added support for the following AWS regions: - ap-northeast-2 (Seoul) - ap-south-2 (Hyderabad) ### [](#centralized-egress-for-byoc-on-azure-beta)Centralized egress for BYOC on Azure: beta You can route all Azure BYOC cluster egress through your own Azure Firewall and hub VNet instead of a per-cluster NAT Gateway, so outbound traffic exits through your centralized inspection point. This is useful for regulated environments that require a single, predictable public IP for outbound allowlisting or that prohibit per-cluster NAT Gateways. You can set this up in the Cloud UI or Cloud API when you create the network, or add or change it on an existing network through the Cloud API. Centralized egress is in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release and is enabled per organization. Contact your account team for access. See [Configure Centralized Egress with Azure Firewall](https://docs.redpanda.com/cloud-data-platform/networking/byoc/azure/nat-free-egress/). ### [](#improved-serverless-trial-onboarding)Improved Serverless trial onboarding New Serverless free-trial users now get a guided start in the Redpanda Cloud UI. After you sign up, a few quick questions help Redpanda tailor your experience. The welcome cluster’s **Overview** page summarizes your trial credits, and the **Get started** button offers options to start streaming: create a Redpanda Connect pipeline, use `rpk` from the command line, or connect with your own Kafka client. See [Get started with Serverless](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#get-started-with-serverless). ### [](#sql-editor-in-redpanda-cloud-console)SQL editor in Redpanda Cloud Console Redpanda Cloud Console now includes a built-in SQL editor. For SQL-enabled environments, you can write and run `SELECT` queries against your Redpanda topics directly in the Console without installing a separate PostgreSQL client. The editor supports syntax highlighting, autocomplete, query history, and CSV and JSON export. For Iceberg-enabled topics, an indicator shows when results span both live topic data and Iceberg-committed history. To enable Redpanda SQL, see [Enable Redpanda SQL](https://docs.redpanda.com/cloud-data-platform/sql/get-started/deploy-sql-cluster/). To use the editor, see [Use the SQL Editor](https://docs.redpanda.com/cloud-data-platform/sql/query-data/sql-editor/). ### [](#self-service-organization-deletion-for-serverless)Self-service organization deletion for Serverless You can now permanently delete a Serverless organization (free trial and pay-as-you-go plans) directly from the Cloud UI, without contacting Redpanda Support. The new **Manage organization** page, available from your profile icon, shows your plan and organization-wide resource counts, and walks you through the deletion prerequisites: delete all clusters, delete all private links, and settle any pending or outstanding invoices. See [Manage Your Organization](https://docs.redpanda.com/cloud-data-platform/manage/manage-organization/). ### [](#redpanda-connect-updates)Redpanda Connect updates - Inputs: - [jira](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/jira/): Streams Jira issues, comments, or changelog entries via JQL with incremental polling. ## [](#june-2026)June 2026 ### [](#terraform-provider-redpanda-sql-support)Terraform provider: Redpanda SQL support The Redpanda Terraform provider (v2.1.0+) now supports enabling and managing Redpanda SQL on BYOC clusters on AWS. Use the `rpsql` block on a `redpanda_cluster` resource to enable SQL, configure compute replicas, and retrieve the SQL endpoint URL. See [Enable Redpanda SQL on a BYOC cluster](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/#enable-redpanda-sql-on-a-byoc-cluster). ### [](#terraform-provider-secrets-management)Terraform provider: Secrets management The Redpanda Terraform provider (v2.0.0+) now supports managing cluster secrets with the new `redpanda_secret` resource. The secret value is a write-only attribute, so Terraform never stores it in your state file, and scopes control which features can use the secret, such as Redpanda Connect pipelines. See [Manage cluster secrets](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/#manage-cluster-secrets). ### [](#terraform-provider-shadow-link-management)Terraform provider: Shadow link management The Redpanda Terraform provider (v2.0.0+) now supports managing shadow links for disaster recovery on BYOC and Dedicated clusters with the new `redpanda_shadow_link` resource. Define the link between source and shadow clusters, client connection settings, and topic, consumer group, ACL, and Schema Registry synchronization options in code, with the source cluster password stored as a cluster secret rather than in your state file. See the [`redpanda_shadow_link` reference](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/shadow_link) and [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/). ### [](#centralized-egress-for-byoc-on-gcp-beta)Centralized egress for BYOC on GCP: beta You can route all GCP BYOC cluster egress through your own GCP hub VPC and NAT VM instead of a per-cluster Cloud NAT, so outbound traffic exits through your centralized inspection point. This is useful for regulated environments that require a single, predictable public IP for outbound allowlisting or that prohibit per-cluster Cloud NAT. Centralized egress is in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release and is enabled per organization. Contact your account team for access. See [Configure Centralized Egress with GCP VPC Peering](https://docs.redpanda.com/cloud-data-platform/networking/byoc/gcp/nat-free-egress/). ### [](#gcp-lakehouse-catalog-for-iceberg-topics)GCP Lakehouse catalog for Iceberg topics BYOC clusters on GCP can now use GCP Lakehouse as an Iceberg REST catalog. See [Use Iceberg Topics with GCP Lakehouse](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-gcp-biglake/). ### [](#customer-managed-default-topic-settings)Customer-managed default topic settings You can now set cluster-wide defaults for new topics on BYOC and Dedicated clusters. The [`default_topic_partitions`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#default_topic_partitions), [`log_retention_ms`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#log_retention_ms), and [`retention_bytes`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#retention_bytes) cluster properties are now customer-managed. The default topic retention period (`log_retention_ms`) previously could only be changed by Redpanda support. See [Configure Cluster Properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/). ### [](#redpanda-connect-updates-2)Redpanda Connect updates - Processors: - [try\_catch](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/try_catch/): Combines the behavior of the `try` and `catch` processors into a single block with an explicit recovery path. ## [](#may-2026)May 2026 ### [](#redpanda-console-redesigned-security-page)Redpanda Console: redesigned Security page Redpanda Console has a redesigned Security page with three tabs (**Users**, **Roles**, and **Permissions**). Each user and role has a detail page for managing its permissions. - The **Users** tab lists each user with their assigned roles and a count of their ACLs. Filter the list by name using regular expressions; for example, `^prod-` matches every user starting with `prod-`. - Open a user or role to manage permissions on its detail page. The **ACLs** section shows one row per rule, with columns for type, resource, operation, permission, and host, and supports three actions: - **\+ Add ACL** opens a focused form where you specify the resource type, pattern type, resource name, operation, permission, and host. - **Allow all operations** grants full wildcard access across all resource types in a single step. Use this for testing only; it is too broad for production. - Select rows with the checkboxes and click **Delete selected** to remove ACLs in bulk. - The **Permissions** tab is a unified, cluster-wide view of every principal with ACLs. Each row shows direct ACL counts and ACLs inherited from roles, with a red badge highlighting any principal that has Deny rules. Expand a row to see all of that principal’s ACLs in one table: direct rules first, then sections labeled **VIA ROLE: ** for each role they inherit from. Search across principals, resources, and roles, or click **Create ACL** to add a rule from scratch. See [Configure ACLs](https://docs.redpanda.com/cloud-data-platform/security/authorization/acl/) for the full ACL reference and [Configure RBAC in the Data Plane](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac_dp/) for role management. ### [](#redpanda-sql)Redpanda SQL Redpanda SQL is available on BYOC clusters running on AWS. Run real-time SQL queries on Redpanda topic data, including the Iceberg history of Iceberg-enabled topics, using standard PostgreSQL syntax. Connect with `psql` or any PostgreSQL driver. See the [Quickstart](https://docs.redpanda.com/cloud-data-platform/sql/get-started/sql-quickstart/) and [Overview](https://docs.redpanda.com/cloud-data-platform/sql/get-started/overview/). ### [](#centralized-egress-for-byoc-on-aws-beta)Centralized egress for BYOC on AWS: beta You can route all BYOC cluster egress through your own AWS Transit Gateway and hub VPC instead of a per-VPC NAT Gateway, so outbound traffic exits through your centralized inspection point. This is useful for regulated environments that prohibit per-VPC NAT Gateways and for consolidating egress behind a single, predictable public IP for outbound allowlisting. Centralized egress is in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release and is enabled per organization. Contact your account team for access. See [Configure Centralized Egress with AWS Transit Gateway](https://docs.redpanda.com/cloud-data-platform/networking/byoc/aws/nat-free-egress/). ### [](#schema-registry-authorization-enabled-by-default)Schema Registry Authorization enabled by default Schema Registry Authorization is now enabled automatically on all new BYOC and Dedicated clusters. The [`schema_registry_enable_authorization`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#schema_registry_enable_authorization) cluster property is set to `true` at provisioning, and the predefined Admin, Writer, and Reader roles include Schema Registry permissions for the `subject` and `registry` ACL resource types. You can use ACLs and RBAC roles to grant fine-grained access to schemas and subjects without any additional setup. See [Schema Registry Authorization](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-authorization/) and [Predefined roles](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/#predefined-roles). ### [](#account-impersonation-schema-registry-support)Account impersonation: Schema Registry support [Account impersonation](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#account-impersonation) now supports Schema Registry in addition to the Kafka API. With Schema Registry impersonation enabled, the schemas and subjects users see in the Redpanda Cloud UI match exactly what they can access with the Cloud API or `rpk`. You can enable impersonation independently for each subsystem from the **Dataplane settings** page. ### [](#redpanda-connect-updates-3)Redpanda Connect updates - Inputs: - [salesforce](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce/): Executes a SOQL query against Salesforce and emits one message per record. - [salesforce\_cdc](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_cdc/): Captures change data from Salesforce objects using the Pub/Sub API. - [salesforce\_graphql](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_graphql/): Executes GraphQL queries against Salesforce. - Outputs: - [gcp\_bigquery\_write\_api](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/gcp_bigquery_write_api/): Writes records to BigQuery using the Storage Write API for higher throughput and lower latency than the streaming insert API. - Metrics: - [open\_telemetry\_collector](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/metrics/open_telemetry_collector/): Pushes metrics using the OpenTelemetry Protocol (OTLP) over HTTP or gRPC. - Tracers: - [open\_telemetry\_collector](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/open_telemetry_collector/): Pushes tracing data using the OpenTelemetry Protocol (OTLP) over HTTP or gRPC. - Behavior changes: - [`gcp_bigquery_write_api`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/gcp_bigquery_write_api/) output: The default value of `max_in_flight` has been reduced from `64` to `4`. If you rely on the previous behavior for throughput, set `max_in_flight` explicitly in your pipeline configuration. The output also gained new fields for tuning Storage Write API behavior: `max_cached_streams`, `schema_resolve_timeout`, and `schema_evolution_timeout`. - New field support: - Kafka producer tuning fields on the [`kafka_franz`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/kafka_franz/) and [`redpanda`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda/) outputs: `acks`, `max_in_flight_requests`, `max_buffered_records`, `max_buffered_bytes`, `record_retries`, and `record_delivery_timeout`. Use these to align pipeline producers with broker durability and back-pressure requirements. - Removed components: - `salesforce` processor (introduced in April 2026): Replaced with the dedicated [`salesforce`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce/), [`salesforce_cdc`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_cdc/), and [`salesforce_graphql`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_graphql/) inputs. The [`salesforce_sink`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/salesforce_sink/) output remains available. ### [](#extended-serverless-free-trial)Extended Serverless free trial The free trial for Redpanda Serverless now lasts 30 days, up from 14 days. The $100 (USD) credit allowance and 7-day grace period are unchanged. Sign up at [redpanda.com](https://www.redpanda.com/try-data-streaming). See [Serverless clusters](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/). ### [](#service-account-token-rate-limits)Service account token rate limits A daily limit now applies to service account access token requests for each organization. Clients that exceed the limit receive `HTTP 429` responses. Cache tokens until close to expiry to stay within the limit, and contact Redpanda Support if your workload requires a higher daily limit. See [Service account token rate limits](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#service-account-token-rate-limits). ## [](#april-2026)April 2026 ### [](#self-service-sign-up-through-google-cloud-marketplace)Self-service sign-up through Google Cloud Marketplace You can now subscribe to Redpanda Cloud directly through Google Cloud Marketplace with pay-as-you-go billing, with no sales contact required. Self-service sign-up provisions Serverless and Dedicated clusters only. New subscriptions receive $300 (USD) in free credits to spend in the first 30 days. See [Use GCP Pay As You Go](https://docs.redpanda.com/cloud-data-platform/billing/gcp-pay-as-you-go/). ### [](#iceberg-configurable-table-namespace)Iceberg: Configurable table namespace You can now set a custom namespace for Iceberg tables instead of the default `redpanda` namespace, using the `iceberg_default_catalog_namespace` cluster property. A custom namespace is useful when multiple clusters write to the same catalog provider (such as AWS Glue), because each cluster must use a distinct namespace to avoid table name collisions. This property must be set when you first enable Iceberg and cannot be changed afterward. See [Enable Iceberg integration](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/#enable-iceberg-integration). ### [](#group-based-access-control-gbac)Group-based access control (GBAC) - With [GBAC in the control plane](https://docs.redpanda.com/cloud-data-platform/security/authorization/gbac/gbac/), you can manage access to organization-level resources using OIDC groups from your identity provider. Assign OIDC groups to roles so that users inherit access based on their group membership. - With [GBAC in the data plane](https://docs.redpanda.com/cloud-data-platform/security/authorization/gbac/gbac_dp/), you can configure cluster-level permissions for provisioned users at scale using OIDC groups. Because group membership is managed by your identity provider, onboarding and offboarding require no changes in Redpanda. GBAC is available for BYOC and Dedicated clusters. In addition to the predefined roles (including Reader, Writer, and Admin) that you cannot modify or delete, you can now create custom roles. ### [](#remote-mcp-deprecated)Remote MCP: Deprecated Remote MCP has been deprecated and removed from Redpanda Cloud. ### [](#increased-serverless-limits-for-redpanda-connect-pipelines)Increased Serverless limits for Redpanda Connect pipelines Serverless clusters now support up to 100 Redpanda Connect pipelines. See [Serverless usage limits](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/#_serverless_usage_limits). ### [](#terraform-provider-write-only-attributes-for-sensitive-fields)Terraform provider: Write-only attributes for sensitive fields The Redpanda Terraform provider (v1.6.0+) now supports [Terraform 1.11+ write-only attributes](https://developer.hashicorp.com/terraform/plugin/framework/resources/write-only-arguments) for sensitive fields such as user passwords and pipeline client secrets. Use the new `password_wo` and `password_wo_version` attributes (and equivalents for other sensitive fields) to keep credentials out of your `.tfstate` file. See [Manage sensitive attributes with write-only fields](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/#manage-sensitive-attributes-with-write-only-fields). ### [](#redpanda-connect-updates-4)Redpanda Connect updates - The Redpanda Connect pipeline creation and editing workflow has been simplified. The new UI replaces the previous multi-page wizard with a visual pipeline diagram, an IDE-like configuration editor, slash commands for inserting variables, and inline links to component documentation. See the [Redpanda Connect quickstart](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/) to try it out. - Outputs: - [arc](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/arc/): Send data to an Arc columnar analytical database using its high-performance MessagePack ingestion endpoint. - [salesforce\_sink](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/salesforce_sink/): Write messages to Salesforce, routing each Kafka topic to its own sObject configuration. Supports both realtime (sObject Collections REST API) and bulk modes (Bulk API 2.0). - Processors: - [string\_split](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/string_split/): Splits strings into multiple parts using a delimiter, creating new messages or fields for each part. - `salesforce` (deprecated May 2026): Replaced with dedicated [salesforce](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce/), [salesforce\_cdc](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_cdc/), and [salesforce\_graphql](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/salesforce_graphql/) inputs. ## [](#march-2026)March 2026 Redpanda Console now supports paginating past the previous 500-record cap when you browse topic messages, so you can inspect large topics without being limited to the initial result set. See [Paginate Messages in Redpanda Console](https://docs.redpanda.com/cloud-data-platform/develop/consume-data/paginate-messages-events/). ### [](#redpanda-connect-updates-5)Redpanda Connect updates - Inputs: - [oracledb\_cdc](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/oracledb_cdc/): Stream changes from an Oracle database for Change Data Capture (CDC). - [aws\_cloudwatch\_logs](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_cloudwatch_logs/): Consume log events from AWS CloudWatch Logs. Supports filtering by log streams, CloudWatch filter patterns, and configurable start times. - [aws\_dynamodb\_cdc](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/aws_dynamodb_cdc/): Consume item-level changes from DynamoDB Streams with automatic checkpointing and shard management. - Outputs: - [iceberg](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/iceberg/): Write data to Apache Iceberg tables using the REST catalog. - Bloblang methods: - [`escape_url_path`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#escape_url_path): Escapes a string for safe use in URL path segments using percent-encoding. - [`parse_logfmt`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#parse_logfmt): Parses a logfmt-encoded string into an object of key-value pairs. - [`unescape_url_path`](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/bloblang/methods/#unescape_url_path): Unescapes a URL path segment, converting percent-encoded sequences back to their original characters. - Removed components: - `legacy_redpanda_migrator` input and output - `legacy_redpanda_migrator_offsets` input and output - `redpanda_migrator_bundle` input and output Use the unified [`redpanda_migrator`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/redpanda_migrator/) input and [`redpanda_migrator`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/redpanda_migrator/) output instead. ### [](#cloud-topics)Cloud Topics [Cloud Topics](https://docs.redpanda.com/cloud-data-platform/develop/topics/cloud-topics/) are now available, making it possible to use durable cloud storage (S3, ADLS, GCS) as the primary backing store instead of local disk, eliminating over 90% of cross-AZ replication costs. This makes them ideal for latency-tolerant, high-throughput workloads such as observability streams, analytics pipelines, and AI/ML training data feeds, where cross-AZ networking charges are the dominant cost driver. You can use Cloud Topics exclusively or in combination with standard topics on a cluster supporting low-latency workloads. ### [](#schema-registry-metadata-properties)Schema Registry metadata properties [Schema Registry metadata properties](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/#metadata-properties) let you store and retrieve arbitrary key-value pairs alongside schemas. Properties such as `owner`, `team`, or `application.version` travel with the schema through its lifecycle, making it easier to track ownership and lineage without modifying the schema itself. You can set metadata when registering a schema using the `POST /subjects/{subject}/versions` endpoint or with the `--metadata-properties` flag in `rpk registry schema create`. Metadata is returned in API responses and viewable with `rpk registry schema get --print-metadata` or in Redpanda Cloud Console. ### [](#schema-registry-contexts)Schema Registry Contexts [Schema Registry contexts](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/) provide isolated namespaces that separate schemas, subjects, and configuration within a single Schema Registry instance. Each context maintains its own schema ID counter, mode settings, and compatibility settings. On Serverless clusters, Redpanda uses contexts internally for per-tenant isolation. Contexts are not exposed to end users on Serverless. On BYOC and Dedicated clusters, contexts are available and user-configurable. ### [](#user-based-throughput-quotas)User-based throughput quotas Redpanda now supports throughput quotas based on authenticated user principals. Unlike client-based quotas (which rely on self-declared `client-id` values), [user-based quotas](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/#set-user-based-quotas) enforce limits using verified identities from SASL, mTLS, or OIDC authentication. You can set quotas for individual users, default users, or fine-grained user/client combinations. ### [](#iceberg-expanded-json-schema-support)Iceberg: Expanded JSON Schema support Redpanda now supports additional JSON Schema patterns when translating to Iceberg tables: - `$ref` support: Internal references using `$ref` (for example, `"$ref": "#/definitions/myType"`) are resolved from schema resources declared in the same document. External references are not yet supported. - Map type from `additionalProperties`: `additionalProperties` objects that contain subschemas now translate to Iceberg `map`. - `oneOf` nullable pattern: The `oneOf` keyword is now supported for the standard nullable pattern if exactly one branch is `{"type":"null"}` and the other is a non-null schema. See [Specify Iceberg Schema](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/#how-iceberg-modes-translate-to-table-format) for JSON types mapping and updated requirements. ### [](#ordered-rack-preference-for-leader-pinning)Ordered rack preference for leader pinning [Leader pinning](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/leader-pinning/) now supports the `ordered_racks` configuration value, which lets you specify preferred racks in priority order. Unlike `racks`, which distributes leaders uniformly across all listed racks, `ordered_racks` places leaders in the highest-priority available rack and fails over to subsequent racks only when higher-priority racks become unavailable. ### [](#byovpc-on-aws-ga)BYOVPC on AWS: GA [BYOVPC on AWS](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/vpc-byo-aws/) is now generally available (GA). With Bring Your Own VPC (BYOVPC), you deploy the Redpanda data plane into your own VPC and manage security policies and resources yourself, including subnets, IAM roles, firewall rules, and storage buckets. The Redpanda BYOVPC Terraform Module contains Terraform code that deploys the resources required for a BYOVPC cluster on AWS. Secrets management is enabled by default with the Terraform module. ### [](#iceberg-topics-with-snowflake-open-catalog-ga)Iceberg topics with Snowflake Open Catalog: GA The [Snowflake and Open Catalog integration](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/redpanda-topics-iceberg-snowflake-catalog/) for Iceberg topics is now generally available (GA). ### [](#billing-notifications)Billing notifications Redpanda Cloud now sends email notifications to organization admins when credit or commit balances reach spending thresholds (50%, 30%, 10%, and 0% remaining). You can manage your notification preferences or opt out at any time. See [Manage Billing Notifications](https://docs.redpanda.com/cloud-data-platform/billing/billing-notifications/). ## [](#february-2026)February 2026 ### [](#serverless-on-aws-ga)Serverless on AWS: GA [Serverless](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) on AWS is now generally available (GA). This release includes private networking with AWS PrivateLink. You can use the Cloud Console, the Cloud API, or the Redpanda Terraform provider to create and manage Serverless private links. Serverless is the easiest and fastest way to begin streaming data with Redpanda. ### [](#enable-schema-id-validation)Enable schema ID validation You can now enable [schema ID validation](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-id-validation/) by [configuring the `enable_schema_id_validation` cluster property](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/). This controls whether or not Redpanda validates schema IDs in records and which topic properties are enforced. Use caution when enabling this property, because it could cause decompression across topics and increase CPU load. ### [](#cross-region-aws-privatelink)Cross-region AWS PrivateLink AWS PrivateLink now supports cross-region connectivity, allowing clients in different AWS regions to connect to your Redpanda cluster through PrivateLink. Configure supported regions in the [Cloud Console](https://docs.redpanda.com/cloud-data-platform/networking/configure-privatelink-in-cloud-ui/#cross-region-privatelink), using the [Cloud API](https://docs.redpanda.com/cloud-data-platform/networking/aws-privatelink/#cross-region-privatelink), or with the Redpanda Terraform provider (v1.7.0+) using the `supported_regions` attribute of the `aws_private_link` block. See [Configure cross-region AWS PrivateLink](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/#configure-cross-region-aws-privatelink). This feature requires multi-AZ cluster deployments. ## [](#january-2026)January 2026 ### [](#redpanda-connect-updates-6)Redpanda Connect updates - Inputs: - [otlp\_grpc](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/otlp_grpc/): Receive OpenTelemetry traces, logs, and metrics via OTLP/gRPC protocol. Exposes an OpenTelemetry Collector gRPC receiver that accepts traces, logs, and metrics, converting them to individual Redpanda OTEL v1 protobuf messages optimized for Kafka partitioning. - [otlp\_http](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/otlp_http/): Receive OpenTelemetry traces, logs, and metrics via OTLP/HTTP protocol. Supports both protobuf and JSON formats at standard OTLP endpoints, converting telemetry data to individual messages with embedded Resource and Scope metadata. - Outputs: - [otlp\_grpc](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/otlp_grpc/): Send OpenTelemetry traces, logs, and metrics via OTLP/gRPC protocol. Accepts batches of Redpanda OTEL v1 protobuf messages and converts them to OTLP format for transmission to OpenTelemetry collectors. - [otlp\_http](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/otlp_http/): Send OpenTelemetry traces, logs, and metrics via OTLP/HTTP protocol. Supports both protobuf and JSON content types for flexible integration with OpenTelemetry backends. ### [](#redpanda-connect-and-roles-in-terraform-provider)Redpanda Connect and Roles in Terraform provider The [Redpanda Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) now supports managing roles and Redpanda Connect pipelines. Use the provider to create and manage role-based access control and data pipelines in Redpanda Cloud. ## [](#december-2025)December 2025 ### [](#shadowing)Shadowing Redpanda Cloud now supports [Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/), a disaster recovery solution that provides asynchronous, offset-preserving replication between distinct Redpanda clusters. Shadowing enables cross-region data protection by replicating topic data, configurations, consumer group offsets, ACLs, and Schema Registry data with byte-level fidelity. The shadow cluster operates in read-only mode while continuously receiving updates from the source cluster. During a disaster, you can failover individual topics or an entire shadow link to make resources fully writable for production traffic. Shadowing is supported on BYOC and Dedicated clusters running Redpanda version 25.3 and later. ### [](#metrics-for-serverless)Metrics for Serverless You can now view and export metrics from Serverless clusters to third-party monitoring systems like Prometheus and Grafana. See [Monitor Redpanda Cloud](https://docs.redpanda.com/cloud-data-platform/manage/monitor-cloud/) for details on configuring monitoring for your Serverless cluster and [Metrics Reference](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/) for a list of metrics available in Serverless. ### [](#account-impersonation)Account impersonation BYOC and Dedicated clusters now support unified authentication and authorization between the Redpanda Cloud UI and Redpanda with [account impersonation](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#account-impersonation). This means you can authenticate to fine-grained access within Redpanda using the same credentials you use to authenticate to Redpanda Cloud. With account impersonation (originally called user impersonation), the topics users see in the UI are identical to what they can access with the Cloud API or `rpk`, ensuring consistent permissions across all interfaces and clear auditing of data plane user actions. ### [](#redpanda-connect-updates-7)Redpanda Connect updates - Tracers: - [Redpanda](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/tracers/redpanda/): The Redpanda tracer exports distributed tracing data to a Redpanda topic, enabling you to monitor and debug your Redpanda Connect pipelines. Traces are exported in OpenTelemetry format as JSON, allowing integration with observability platforms like Jaeger, Grafana Tempo, or custom trace consumers. ## [](#november-2025)November 2025 ### [](#serverless-on-gcp-beta)Serverless on GCP: beta You can now create [Serverless clusters](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) on Google Cloud Platform (GCP). Serverless on GCP is in a [beta](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#beta) release. ### [](#support-for-additional-regions)Support for additional regions [BYOC clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/#byoc-supported-regions) on Azure now support the Sweden Central and Germany West Central regions. ### [](#connected-client-monitoring)Connected client monitoring You can view details about Kafka client connections using `rpk` or the Data Plane API. This allows you to view detailed information about active client connections on a cluster, and identify and troubleshoot problematic clients. For more information, see the [connected client details](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/#view-connected-client-details) example in the Manage Throughput guide. ### [](#increased-message-size-limit)Increased message size limit Redpanda Cloud increased the [message size limit](https://docs.redpanda.com/cloud-data-platform/develop/topics/create-topic/) on newly-created topics. BYOC and Dedicated clusters have a default message size limit of 20 MiB with a maximum of 32 MiB. Serverless clusters have a default message size limit of 8 MiB with a maximum of 20 MiB. Configure the message size limit with the `max.message.bytes` topic property. The message size setting on existing topics is not changed, but the message size limit on existing topics can only be updated to the new maximum. ### [](#redpanda-connect-updates-8)Redpanda Connect updates Redpanda Connect provides a simplified [quickstart](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/) experience in the UI that helps you to start building data pipelines. The quickstart creates pipelines to stream data into and out of Redpanda using the pipeline editor. ### [](#get-started-with-serverless)Get Started with Serverless A Serverless cluster’s **Overview** page now provides a **Get Started** guide to help you start streaming your own data with a [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/) pipeline. It lets you stream data into and out of Redpanda without writing producer/consumer code. ### [](#remote-read-replicas-ga)Remote read replicas: GA [Remote read replicas](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/remote-read-replicas/) are now generally available (GA) for BYOC clusters on AWS and GCP. This feature allows you to create read-only topics that mirror a topic on a different cluster, providing greater flexibility and scalability for your data streaming needs. ### [](#schema-registry-and-acls-in-terraform-provider)Schema Registry and ACLs in Terraform provider The [Redpanda Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) now supports managing schemas and Schema Registry ACLs. You can use the provider to register schemas in formats such as Avro, Protobuf, or JSON Schema, and control access to Schema Registry subjects and operations through ACLs. ## [](#october-2025)October 2025 ### [](#api-gateway-access)API Gateway access BYOC and Dedicated clusters with private networking now allow control of API Gateway network access, independent of the Redpanda cluster. When you create a cluster, you can choose either public or private access for the API Gateway: - Public access exposes Redpanda Console, Data Plane API, and MCP Server API endpoints over the internet, although they remain protected by your authentication and authorization controls. - Private access restricts endpoint access to your private network (VPC or VNet) only. After the cluster is created, you can change the API Gateway access on the Dataplane settings page. If you change from public to private access, users without VPN access to the Redpanda VPC will lose access to these services. ### [](#redpanda-connect-updates-9)Redpanda Connect updates - Inputs: - [Microsoft SQL Server CDC](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/microsoft_sql_server_cdc/): Streams change data from a Microsoft SQL Server database into Redpanda Connect using Change Data Capture (CDC). - Outputs: - [CyborgDB](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/cyborgdb/): Write vectors to a CyborgDB encrypted index. CyborgDB provides end-to-end encrypted vector storage with automatic dimension detection and index optimization. - Processors: - [`jira`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/processors/jira/): Executes Jira API queries based on input messages and returns structured results. The processor handles pagination, retries, and field expansion automatically. - Deprecated components: - `redpanda_migrator` input and output (renamed to `legacy_redpanda_migrator`) - `redpanda_migrator_offsets` input and output (renamed to `legacy_redpanda_migrator_offsets`) Migrate from these deprecated components to the new unified `redpanda_migrator` input/output pair. For detailed migration instructions, see [Migrate to the Unified Redpanda Migrator](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/migrate-unified-redpanda-migrator/). - `redpanda_migrator_bundle` input and output (these are part of the legacy migration architecture and internally depend on the deprecated `legacy_redpanda_migrator` and `legacy_redpanda_migrator_offsets` components) - `kafka`, `kafka_franz`, and `redpanda_common` inputs and outputs. These components have been consolidated into the unified `redpanda` input and output components. Migrate existing configurations to use the new `redpanda` components for continued support and access to the latest features. For detailed information about recent component updates, see [What’s New in Redpanda Connect](https://docs.redpanda.com/connect/get-started/whats-new/). ## [](#september-2025)September 2025 ### [](#multi-factor-authentication)Multi-factor authentication Enable multi-factor authentication (MFA) to add an extra layer of security to your Redpanda Cloud account. After you enable MFA, you’ll enter your credentials, then be prompted for a one-time code from your authenticator app when you log in. Administrators can also [enforce MFA](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#multi-factor-authentication-mfa) for all members of an organization. ### [](#redpanda-cloud-management-mcp-server-beta)Redpanda Cloud Management MCP Server: beta Connect AI assistants like Claude directly to your Redpanda Cloud account with the new Redpanda Cloud Management MCP Server. This server runs on your computer and provides AI tools for managing clusters, topics, and other cloud resources through natural language commands. Ask your AI assistant to "Create a new topic called user-events" or "List all clusters in my account" and it will handle the technical details automatically. Get started with the [rpk cloud mcp install](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cloud/rpk-cloud-mcp-install/) command. The Redpanda Cloud Management MCP Server uses the Model Context Protocol (MCP) to extend AI assistants with Redpanda-specific capabilities, making cloud operations more accessible through conversational interfaces. ### [](#automatic-topic-creation-and-topic-limit)Automatic topic creation and topic limit For BYOC and Dedicated clusters, you can now configure the `auto_create_topics_enabled` cluster property to automatically create a topic if a client produces to a non-existent topic. For all clusters: each cluster now has a limit of 40,000 topics. ## [](#august-2025)August 2025 ### [](#manage-custom-resource-tags-in-byoc)Manage custom resource tags in BYOC After cluster creation, you can manage custom cloud provider tags and labels on BYOC and BYOVPC/BYOVNet clusters for [AWS](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/create-byoc-cluster-aws/#manage-custom-tags), [Azure](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/create-byoc-cluster-azure/#manage-custom-tags), and [GCP](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/create-byoc-cluster-gcp/#manage-custom-resource-labels-and-network-tags) using the Cloud Control Plane API. This involves refreshing Redpanda agent permissions with `rpk cloud byoc` due to new IAM permissions. ### [](#iceberg-topics-with-aws-glue)Iceberg topics with AWS Glue A new [integration with AWS Glue Data Catalog](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-aws-glue/) allows you to add Redpanda topics as Iceberg tables in your data lakehouse. The AWS Glue catalog integration is available in BYOC clusters with Redpanda version 25.2 and later. See [Integrate with REST Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/rest-catalog/) for supported Iceberg REST catalog integrations. ### [](#manage-throughput)Manage throughput Redpanda Cloud now lets you [manage throughput](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/) configuration at the broker and client levels. You can manage client quotas with [`rpk cluster quotas`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-quotas/) or with the Kafka API. When no quotas apply, the client has unlimited throughput. ## [](#july-2025)July 2025 ### [](#iceberg-topics-in-redpanda-cloud-ga)Iceberg topics in Redpanda Cloud: GA [Iceberg topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/) are now generally available (GA) in Redpanda Cloud. ### [](#byoc-on-azure-ga)BYOC on Azure: GA [BYOC for Azure](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/create-byoc-cluster-azure/) is now generally available (GA). ### [](#schema-registry-authorization)Schema Registry Authorization You can now use [Schema Registry Authorization](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-authorization/) to control access to Schema Registry subjects and operations. Schema Registry Authorization offers more granular control over who can do what with your Redpanda Schema Registry resources. ACLs used for Schema Registry access also support RBAC roles. ### [](#kafka-connect-disabled-on-new-clusters)Kafka Connect disabled on new clusters [Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/) is now disabled by default on all new clusters. To unlock this feature for your account, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). If you previously enabled Kafka Connect on a cluster and want to [disable it](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/), you can use the Cloud API. ### [](#allowlist-nat-gateway-ip)Allowlist NAT gateway IP The [Redpanda NAT gateway IP address](https://docs.redpanda.com/cloud-data-platform/networking/cloud-security-network/#nat-gateways) is now provided in the Cloud UI and the Cloud API for BYOC and Dedicated clusters. If necessary, you can use this IP address to allowlist egress traffic from your Redpanda Connect data sources. ### [](#mtls-and-sasl-authentication-for-kafka-api-on-aws)mTLS and SASL authentication for Kafka API on AWS You can now enable mTLS and SASL authentication simultaneously for the Kafka API on AWS clusters. If you enable both mTLS and SASL on AWS clusters, Redpanda creates two distinct listeners: an mTLS listener operating on one port and a SASL listener operating on a different port. See [Authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#service-authentication) for details on available authentication methods in Redpanda Cloud. ### [](#azure-private-link-in-the-ui-ga)Azure Private Link in the UI: GA You can now [configure Azure Private Link](https://docs.redpanda.com/cloud-data-platform/networking/azure-private-link-in-ui/) for a new BYOC or Dedicated cluster using the Cloud UI. The Azure Private Link service is generally available (GA) in both the Cloud UI and the Cloud API. ### [](#redpanda-connect-in-redpanda-cloud-ga)Redpanda Connect in Redpanda Cloud: GA [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/) is now generally available (GA) in all Redpanda Cloud clusters: BYOC (including BYOVPC/BYOVNet), Dedicated, and Serverless. ### [](#redpanda-connect-updates-10)Redpanda Connect updates Redpanda Connect includes the following updates: - The [GCP Spanner CDC](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gcp_spanner_cdc/) component lets you capture changes from Google Cloud Spanner and stream them into Redpanda. You can use it to ingest data from GCP Spanner databases, enabling real-time data processing and analytics. - The [Slack Reaction](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/outputs/slack_reaction/) component lets you send messages to a Slack channel in response to events in Redpanda. You can use it to create alerts, notifications, or other automated responses based on data changes in Redpanda. - The [Redpanda Cache](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/caches/redpanda/) component lets you cache data in Redpanda, improving performance and reducing latency for data access. You can use it to store frequently accessed data, such as configuration settings or user profiles, in Redpanda. For more detailed information about recent component updates, see [What’s New in Redpanda Connect](https://docs.redpanda.com/connect/get-started/whats-new/). ### [](#serverless-client-connections)Serverless client connections [Serverless](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) clusters have a new usage limit of 10,000 connections. ## [](#june-2025)June 2025 ### [](#schema-registry-ui-for-serverless)Schema Registry UI for Serverless The [Schema Registry UI](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-ui/) is now available for Serverless clusters. ### [](#amazon-vpc-transit-gateway)Amazon VPC Transit Gateway For BYOC and BYOVPC clusters on AWS, you can set up an [Amazon VPC Transit Gateway](https://docs.redpanda.com/cloud-data-platform/networking/byoc/aws/transit-gateway/) to connect VPCs to Redpanda services while maintaining control over network traffic. ### [](#support-for-additional-regions-2)Support for additional regions Serverless clusters now support the following new [regions on AWS](https://docs.redpanda.com/cloud-data-platform/reference/tiers/serverless-regions/): ap-northeast-1 (Tokyo), ap-southeast-1 (Singapore), and eu-west-2 (London). ### [](#http-gateway)HTTP gateway The [`gateway`](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/inputs/gateway/) component is now available in Redpanda Connect for Redpanda Cloud. This component allows you to create an HTTP endpoint that can receive data from any HTTP client and stream it into Redpanda. You can use the gateway to ingest data from IoT devices, web applications, or any other HTTP-based source. See the [Ingest Real-Time Sensor Telemetry with the HTTP Gateway](https://docs.redpanda.com/cloud-data-platform/develop/connect/guides/cloud/gateway/) guide for more information. ## [](#may-2025)May 2025 ### [](#redpanda-connect-for-byovnet-on-azure-beta)Redpanda Connect for BYOVNet on Azure: beta [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/) is now enabled when you create a BYOVNet cluster on [Azure](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/vnet-azure/). ### [](#secrets-management-for-byovpc-clusters-on-aws-and-gcp)Secrets management for BYOVPC clusters on AWS and GCP You can now create new BYOVPC clusters with secrets management enabled by default on [AWS](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/vpc-byo-aws/) and [GCP](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/vpc-byo-gcp/). You can also enable secrets management for existing BYOVPC clusters on AWS and GCP. For GCP, see [Enable Secrets Management for BYOVPC Clusters on GCP](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/enable-secrets-byovpc-gcp/). For AWS, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). ### [](#serverless-standard-deprecated)Serverless Standard: deprecated Serverless Standard is deprecated. All existing clusters will be migrated to the new [Serverless](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) platform (with higher usage limits, 99.9% SLA, and additional regions) on August 31, 2025. - Retirement date: August 30, 2025 ### [](#cloud-api-beta-versions-deprecated)Cloud API beta versions: deprecated The Cloud Control Plane API versions v1beta1 and v1beta2, and Data Plane API versions v1alpha1 and v1alpha2 are deprecated. These Cloud API versions will be removed in a future release and are not recommended for use. The deprecation timeline is: - Announcement date: May 27, 2025 - End-of-support date: November 28, 2025 - Retirement date: May 28, 2026 See the [Cloud API Deprecation Policy](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-deprecation-policy) for more information. ### [](#read-only-cluster-configuration-properties)Read-only cluster configuration properties You can now [view the value of read-only cluster configuration properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/#view-cluster-property-values) with `rpk cluster config` or with the Cloud API. Available properties are listed in [Cluster Properties](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/) and [Object Storage Properties](https://docs.redpanda.com/cloud-data-platform/reference/properties/object-storage-properties/). ### [](#iceberg-topics-in-azure-beta)Iceberg topics in Azure: beta [Iceberg topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/) are now supported for BYOC clusters in Azure. ### [](#support-for-additional-region)Support for additional region [BYOC clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/#byoc-supported-regions) on GCP now support the us-west2 (Los Angeles) region. ### [](#redpanda-terraform-provider-ga)Redpanda Terraform provider: GA The [Redpanda Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) is now generally available (GA). The provider lets you create and manage resources in Redpanda Cloud, such as clusters, topics, users, ACLs, networks, and resource groups. ## [](#april-2025)April 2025 ### [](#mtls-and-sasl-authentication-for-kafka-api-on-gcp)mTLS and SASL authentication for Kafka API on GCP You can now enable mTLS and SASL authentication simultaneously for the Kafka API on GCP clusters. If you enable both mTLS and SASL on GCP clusters, Redpanda creates two distinct listeners: an mTLS listener operating on one port and a SASL listener operating on a different port. See [Authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#service-authentication) for details on available authentication methods in Redpanda Cloud. ### [](#increased-number-of-supported-partitions)Increased number of supported partitions The number of partitions (pre-replication) Redpanda Cloud supports for each [usage tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/) has been doubled. For example, the number of supported partitions in tier 1 went from 1,000 to 2,000, and tier 5 went from 22,800 to 45,600. ### [](#iceberg-topics-beta)Iceberg topics: beta The [Iceberg integration for Redpanda](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/) allows you to store topic data in the cloud in the Iceberg open table format. This makes your streaming data immediately available in downstream analytical systems without setting up and maintaining additional ETL pipelines. You can also integrate your data directly into commonly-used big data processing frameworks, standardizing and simplifying the consumption of streams as tables in a wide variety of data analytics pipelines. Iceberg topics are supported for BYOC clusters in AWS and GCP. ### [](#cluster-configuration)Cluster configuration You can now [configure certain cluster properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/) with `rpk cluster config` or with the Cloud API. For example, you can enable and manage [Iceberg topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/), [data transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/), and [audit logging](https://docs.redpanda.com/cloud-data-platform/manage/audit-logging/). Available properties are listed in [Cluster Configuration Properties](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/). Iceberg topics properties are available for clusters running Redpanda version 25.1 or later. ### [](#manage-secrets-for-cluster-configuration)Manage secrets for cluster configuration Redpanda Cloud now supports managing secrets that you can reference in cluster properties, for example, to configure Iceberg topics. You can create, update, and delete secrets and reference a secret in cluster properties using `rpk` or the Cloud API. See also: - Manage secrets using [`rpk security secret`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-security/rpk-security-secret/) - Manage secrets using the [Data Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dataplane-api/#manage-secrets) - Reference a secret in a cluster property using [`rpk cluster config set`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-config-set/) - Reference a secret in a cluster property using the [Control Plane API](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/) ### [](#data-transforms-ga)Data transforms: GA WebAssembly [data transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/) are now generally available in Redpanda Cloud. Data transforms let you run common data streaming tasks within Redpanda, like filtering, scrubbing, and transcoding. Data transforms are supported for BYOC and Dedicated clusters running Redpanda version 24.3 and later. ### [](#redpanda-connect-for-byovpc-on-aws-and-gcp-beta)Redpanda Connect for BYOVPC on AWS and GCP: beta Redpanda Connect is now enabled when you create a BYOVPC cluster on [AWS](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/vpc-byo-aws/) or [GCP](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/vpc-byo-gcp/). You can also add Redpanda Connect to an [existing BYOVPC GCP cluster](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/gcp/enable-rpcn-byovpc-gcp/). ## [](#march-2025)March 2025 ### [](#serverless)Serverless For a better customer experience, the Serverless Standard and Serverless Pro products have merged into a single offering. [Serverless clusters](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) now include the higher usage limits, 99.9% SLA, additional AWS regions, and the free trial. ### [](#cloud-api-ga)Cloud API: GA The Cloud API is now generally available. It includes endpoints for [managing Serverless clusters](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-serverless-controlplane-api/), configuring RBAC in [BYOC](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/#manage-rbac), [Serverless](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-serverless-controlplane-api/#manage-rbac), and [Dedicated](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dedicated-controlplane-api/#manage-rbac) clusters, and [using Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dataplane-api/#use-redpanda-connect). To get started, see the [Redpanda Cloud API overview](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview) or try the [Cloud API Quickstart](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-quickstart). For full reference documentation, see [Control Plane API](https://docs.redpanda.com/api/doc/cloud-controlplane/) and [Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/). ### [](#support-for-additional-regions-3)Support for additional regions [BYOC clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/#byoc-supported-regions) on GCP now support the europe-southwest1 (Madrid) region. ### [](#byovpc-support-in-the-redpanda-terraform-provider-0-14-0-beta)BYOVPC support in the Redpanda Terraform provider 0.14.0: Beta The [Redpanda Terraform provider](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/cluster#byovpc) now supports BYOVPC clusters on AWS and GCP. You can use the provider to create and manage BYOVPC clusters in Redpanda Cloud. ## [](#february-2025)February 2025 ### [](#role-based-access-control-rbac)Role-based access control (RBAC) With [RBAC in the control plane](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/), you can manage access to organization-level resources like clusters, resource groups, and networks. For example, you could grant everyone access to clusters in a development resource group while limiting access to clusters in a production resource group. Or, you could limit access to geographically-dispersed clusters in accordance with data residency laws. With [RBAC in the data plane](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac_dp/), you can configure cluster-level permissions for provisioned users at scale. ### [](#improved-private-service-connect-support-with-az-affinity)Improved Private Service Connect support with AZ affinity The latest version of the Redpanda [GCP Private Service Connect](https://docs.redpanda.com/cloud-data-platform/networking/gcp-private-service-connect/) service provides the ability to allow requests from Private Service Connect endpoints to stay within the same availability zone, avoiding additional networking costs. The service is now fully supported (GA). To upgrade, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). > ❗ **IMPORTANT** > > Deprecated: The original GCP Private Service Connect service is deprecated and will be removed in a future release. ### [](#serverless-pro-usage-limits-increased)Serverless Pro usage limits increased Usage limits for Serverless Pro clusters increased to: ingress = 100 MBps, egress = 300 MBps, partitions = 5000. ### [](#cloud-api-reference)Cloud API reference The Cloud API reference is now provided as separate references for the [Control Plane API](https://docs.redpanda.com/api/doc/cloud-controlplane/) and [Data Plane APIs](https://docs.redpanda.com/api/doc/cloud-dataplane/). The Control Plane API and Data Plane APIs follow separate OpenAPI specifications, so the reference is updated to better reflect the structure of the Cloud APIs and to improve usability of the documentation. See also: [Cloud API Overview](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview). ## [](#january-2025)January 2025 ### [](#new-tiers-and-regions-on-azure)New tiers and regions on Azure [Tiers 1-5](https://docs.redpanda.com/cloud-data-platform/reference/tiers/) are now supported for BYOC and Dedicated clusters running on Azure. Also, the following [regions](https://docs.redpanda.com/cloud-data-platform/reference/tiers/dedicated-tiers/#dedicated-supported-regions) were added for Dedicated clusters: Central US, East US 2, Norway East. ### [](#serverless-pro-la)Serverless Pro: LA Serverless Pro is a new enterprise-level cluster option. It is similar to Serverless Standard, but with higher usage limits and Enterprise support. This is a limited availability (LA) release. To start using Serverless Pro, contact [Redpanda Sales](https://redpanda.com/try-redpanda?section=enterprise-trial). ### [](#aws-privatelink-ga)AWS PrivateLink: GA AWS PrivateLink is now generally available for private networking in the [Cloud UI](https://docs.redpanda.com/cloud-data-platform/networking/configure-privatelink-in-cloud-ui/) and the [Cloud API](https://docs.redpanda.com/cloud-data-platform/networking/aws-privatelink/). ## [](#december-2024)December 2024 ### [](#support-for-additional-regions-4)Support for additional regions For [BYOC clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/#byoc-supported-regions), Redpanda added support for the following regions: - GCP: europe-west9 (Paris), southamerica-west1 (Santiago) - AWS: ap-southeast-3 (Jakarta), eu-north-1 (Stockholm), eu-south-1 (Milan), eu-west-3 (Paris) ### [](#redpanda-connect-updates-11)Redpanda Connect updates Redpanda Connect is now available on Dedicated clusters. This is a limited availability (LA) release. [Secret management](https://docs.redpanda.com/cloud-data-platform/develop/connect/configuration/secret-management/) is also available on BYOC, Dedicated, and Serverless clusters so that you can add secrets to your pipelines without exposing them. ### [](#leader-pinning)Leader pinning For a Redpanda cluster deployed across multiple availability zones (AZs), [leader pinning](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/leader-pinning/) ensures that a topic’s partition leaders are geographically closer to clients. Leader pinning can lower networking costs and help guarantee lower latency by routing produce and consume requests to brokers located in certain AZs. ## [](#november-2024)November 2024 ### [](#byovpc-on-aws-beta)BYOVPC on AWS: beta With standard BYOC clusters, Redpanda manages security policies and resources for your VPC, including subnetworks, service accounts, IAM roles, firewall rules, and storage buckets. For the highest level of security, you can manage these resources yourself with a [BYOVPC on AWS](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/vpc-byo-aws/), previously known as _customer-managed VPC_. ### [](#customer-managed-vnet-on-azure-la)Customer-managed VNet on Azure: LA With standard BYOC clusters, Redpanda manages security policies and resources for your virtual network (VNet), including subnetworks, managed identities, IAM roles, security groups, and storage accounts. For the highest level of security, you can manage these resources yourself with a [customer-managed VNet on Azure](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/vnet-azure/). Because Azure functionality is provided in limited availability, to unlock this feature, contact [Redpanda support](https://support.redpanda.com/hc/en-us/requests/new). ## [](#october-2024)October 2024 ### [](#byoc-support-in-the-terraform-provider-0-10)BYOC support in the Terraform provider 0.10 The [Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) now supports BYOC clusters. You can use the provider to create and manage BYOC clusters in Redpanda Cloud. ### [](#azure-marketplace-for-dedicated-clusters)Azure Marketplace for Dedicated clusters You can contact [Redpanda sales](https://redpanda.com/try-redpanda?section=enterprise-trial) to request a private offer for monthly or annual [committed use through the Azure Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/azure-commit/). You can then quickly provision Dedicated clusters in Redpanda Cloud, and you can view your bills and manage your subscription directly in Azure Marketplace. ### [](#support-for-aws-graviton3)Support for AWS Graviton3 Redpanda now supports compute-optimized tiers with AWS Graviton3 processors. This saves over 50% in instance costs in all [BYOC tiers](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/). ### [](#redpanda-terraform-provider-for-redpanda-cloud-beta)Redpanda Terraform Provider for Redpanda Cloud: beta The [Redpanda Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) lets you create and manage resources in Redpanda Cloud, such as clusters, topics, users, ACLs, networks, and resource groups. ## [](#september-2024)September 2024 ### [](#schedule-maintenance-windows)Schedule maintenance windows Redpanda Cloud now offers greater flexibility to schedule upgrades to your cluster. By default, Redpanda Cloud may run maintenance operations on any day at any time. You can override this default and \* [schedule a maintenance window](https://docs.redpanda.com/cloud-data-platform/manage/maintenance/#maintenance-windows), which requires Redpanda Cloud to run operations on your specified day and time. ### [](#redpanda-connect-la-for-byoc-beta-for-serverless)Redpanda Connect: LA for BYOC, beta for Serverless [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/) is now integrated into Redpanda Cloud and available as a fully-managed service. This is a limited availability (LA) release for BYOC and a beta release for Serverless. [Choose from a range of connectors, processors, and other components](https://docs.redpanda.com/cloud-data-platform/develop/connect/components/about/) to quickly build and deploy streaming data pipelines or AI applications from the [Cloud UI](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/) or using the [Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/group/endpoint-redpanda-connect-pipeline). Comprehensive metrics, monitoring, and per pipeline scaling are also available. To start using Redpanda Connect, [try this quickstart](https://docs.redpanda.com/cloud-data-platform/develop/connect/connect-quickstart/). For more detailed information about recent component updates, see [What’s New in Redpanda Connect](https://docs.redpanda.com/connect/get-started/whats-new/). ### [](#dedicated-on-azure-la)Dedicated on Azure: LA Redpanda now supports [Dedicated clusters on Azure](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/create-dedicated-cloud-cluster/). This is a limited availability (LA) release for Dedicated clusters. ### [](#remote-read-replicas-on-customer-managed-vpc)Remote read replicas on customer-managed VPC The beta release of [remote read replicas](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/remote-read-replicas/) has been extended to support customer-managed VPC deployments. ## [](#july-2024)July 2024 ### [](#redpanda-cloud-docs)Redpanda Cloud docs The [Redpanda Docs site](https://docs.redpanda.com/home/) has been redesigned for an easier experience navigating Redpanda Cloud docs. We hope that our docs help and inspire our users. Please share your feedback with the links at the bottom of any doc page. ### [](#byoc-on-azure-la)BYOC on Azure: LA Redpanda now supports [BYOC clusters on Azure](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/azure/create-byoc-cluster-azure/). This is a limited availability (LA) release for BYOC clusters. ### [](#enhancements-to-serverless-la)Enhancements to Serverless: LA - The [Redpanda Cloud API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-serverless-controlplane-api/) now includes support for [Serverless](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/). - The Redpanda Schema Registry API is now exposed for Serverless. - Serverless subscriptions can now see detailed billing activity on the **Billing** page. - Serverless added a 99.5% uptime [SLA](https://www.redpanda.com/legal/redpanda-cloud-service-level-agreement) (service level agreement). ### [](#self-service-sign-up-for-dedicated-on-aws-marketplace)Self service sign up for Dedicated on AWS Marketplace To start using Dedicated, sign up on the [AWS Marketplace](https://docs.redpanda.com/cloud-data-platform/billing/aws-pay-as-you-go/). New subscriptions receive $300 (USD) in free credits to spend in the first 30 days. AWS Marketplace charges for anything beyond $300, unless you cancel the subscription. After your credits have been used, you can continue using your cluster without any commitment, only paying for what you consume. ### [](#support-for-additional-regions-5)Support for additional regions For [BYOC clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/#byoc-supported-regions) and [Dedicated clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/dedicated-tiers/#dedicated-supported-regions), Redpanda added support for the following regions: - GCP: asia-east1 (Taiwan), asia-northeast1 (Tokyo), southamerica-east1 (São Paulo) - AWS: ap-east-1 (Hong Kong), ap-northeast-1 (Tokyo), me-central-1 (UAE) ## [](#june-2024)June 2024 ### [](#remote-read-replica-topics-on-byoc-beta)Remote read replica topics on BYOC: beta You can now create [remote read replica topics](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/remote-read-replicas/) on a BYOC cluster with the Cloud API. A remote read replica topic is a read-only topic that mirrors a topic on a different cluster. It can serve any consumer, without increasing the load on the source cluster. ### [](#higher-connection-limits-in-usage-tiers)Higher connection limits in usage tiers Redpanda has increased the number of client connections in all [tiers](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/). For example, tier 1 now supports up to 9,000 maximum connections, and tier 9 supports up to 450,000 maximum connections. Connections are regulated per broker for best performance. ## [](#may-2024)May 2024 ### [](#cloud-api-beta)Cloud API: beta The Cloud API allows you to programmatically manage clusters and resources in your Redpanda Cloud organization. For more information, see the [Cloud API Quickstart](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-quickstart), the [Cloud API Overview](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview), and the full [Control Plane API](https://docs.redpanda.com/api/doc/cloud-controlplane/) and [Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/) reference documentation. ### [](#mtls-authentication-for-kafka-api-clients)mTLS authentication for Kafka API clients mTLS authentication is now available for Kafka API clients. You can [enable mTLS](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#mtls) for your cluster using the Cloud API. ### [](#manage-private-connectivity-in-the-ui)Manage private connectivity in the UI You can now manage GCP Private Service Connect and AWS PrivateLink connections to your BYOC or Dedicated cluster on the **Dataplane settings** page in Redpanda Cloud. See the steps for [PrivateLink](https://docs.redpanda.com/cloud-data-platform/networking/configure-privatelink-in-cloud-ui/) and [Private Service Connect](https://docs.redpanda.com/cloud-data-platform/networking/configure-private-service-connect-in-cloud-ui/). ### [](#single-message-transforms)Single message transforms Redpanda now provides [single message transforms (SMTs)](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/transforms/) to help you modify data as it passes through a connector, without needing additional stream processors. ### [](#support-for-additional-regions-6)Support for additional regions - For [BYOC clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/#byoc-supported-regions), Redpanda added support for the GPC us-west1 region (Oregon) and the AWS ap-south-1 region (Mumbai). - For [Dedicated clusters](https://docs.redpanda.com/cloud-data-platform/reference/tiers/dedicated-tiers/#dedicated-supported-regions), Redpanda added support for the AWS ap-south-1 region. ### [](#simplified-navigation-and-namespaces-renamed-resource-groups)Simplified navigation and namespaces renamed resource groups Redpanda Cloud has a simplified navigation, with clusters and networks available at the top level. It now has a global view of all resources in your organization. Namespaces are now called [resource groups](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#resource-group), although the functionality remains the same. ## [](#april-2024)April 2024 ### [](#additional-cloud-tiers-for-byoc)Additional cloud tiers for BYOC When you create a BYOC or Dedicated cluster, you select a [cloud tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/) with the expected usage for your cluster, including the maximum ingress, egress, partitions (pre-replication), and connections. Redpanda has added tiers 8 and 9 for BYOC clusters, which provide higher supported configurations. ## [](#march-2024)March 2024 ### [](#serverless-limited-availability)Serverless: limited availability [Redpanda Serverless](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) moved out of beta and into limited availability (LA). This means that it has usage limits. During LA, existing clusters can scale to the usage limits, but new clusters may need to wait for availability. Serverless is the fastest and easiest way to start data streaming. It is a production-ready deployment option with automatically-scaling clusters available instantly. To start using Serverless, [sign up for a free trial](https://redpanda.com/try-redpanda/cloud-trial#serverless). This is no base cost, and with pay-as-you-go billing after the trial, you only pay for what you consume. ### [](#authentication-with-sso)Authentication with SSO Redpanda Cloud now supports OpenID Connect (OIDC) integration, so administrators can leverage existing identity providers for user authentication to your Redpanda organization with [single sign-on](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/#single-sign-on) (SSO). Redpanda uses OIDC to delegate the authentication process to an external IdP, such as Okta. To enable this for your account, contact [Redpanda support](https://support.redpanda.com/hc/en-us/requests/new). ## [](#february-2024)February 2024 ### [](#aws-privatelink)AWS PrivateLink [AWS PrivateLink](https://docs.redpanda.com/cloud-data-platform/networking/aws-privatelink/) is now available as an easy and highly secure way to connect to Redpanda Cloud from your VPC. You can set up the PrivateLink endpoint service for a new cluster or an existing cluster. To enable AWS PrivateLink for your account, contact [Redpanda support](https://support.redpanda.com/hc/en-us/requests/new). ### [](#additional-cloud-tiers)Additional cloud tiers When you create a cluster, you select a [cloud tier](https://docs.redpanda.com/cloud-data-platform/reference/tiers/byoc-tiers/) with the expected throughput for your cluster, including the maximum ingress, egress, partitions, and connections. On February 5, Redpanda added tiers 6 and 7 for BYOC clusters, which provide higher throughput limits. ## [](#january-2024)January 2024 ### [](#usage-based-billing-in-marketplace)Usage-based billing in marketplace Redpanda Cloud now supports [usage-based billing](https://docs.redpanda.com/cloud-data-platform/billing/billing/) for Dedicated clusters. Contact [Redpanda sales](https://redpanda.com/try-redpanda?section=enterprise-trial) to request a private offer for monthly or annual committed use. You can then use existing Google Cloud Marketplace or AWS Marketplace credits to quickly provision Dedicated Cloud clusters, and you can view your bills and manage your subscription directly in the marketplace. ## [](#december-2023)December 2023 ### [](#serverless-clusters-beta)Serverless clusters: beta [Redpanda Serverless](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/serverless/) is a managed streaming service (Kafka API) that completely abstracts users from scaling and operational concerns, and you only pay for what you consume. It’s the fastest and easiest way to start event streaming in the cloud. You can try the beta release of Redpanda Serverless with a free trial. ## [](#november-2023)November 2023 ### [](#aws-byoc-support-for-arm-based-graviton2)AWS BYOC support for ARM-based Graviton2 BYOC clusters on AWS now support ARM-based Graviton2 instances. This lowers VM costs and supports increased partition count. ### [](#iceberg-sink-connector)Iceberg Sink connector With the [managed connector for Apache Iceberg](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/create-iceberg-sink-connector/), you can write data into Iceberg tables. This enables integration with the data lake ecosystem and efficient data management for complex analytics. ### [](#schema-registry-management)Schema Registry management In the Redpanda Console UI, you can [perform Schema Registry operations](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-ui/), such as registering a schema, creating a new version of it, and configuring compatibility. The **Schema Registry** page lists verified schemas, including their serialization format and versions. Select an individual schema to see which topics it applies to. ### [](#maintenance-windows)Maintenance windows With maintenance windows, you have greater flexibility to plan upgrades to your cluster. By default, Redpanda Cloud upgrades take place on Tuesdays. Optionally, on the **Dataplane settings** page, you can select a window of specific off-hours for your business for Redpanda to apply updates. All times are in Coordinated Universal Time (UTC). Updates may start at any time during that window. --- # Page 377: Manage **URL**: https://docs.redpanda.com/cloud-data-platform/manage.md --- # Manage > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/index.adoc description: Manage Redpanda. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-05-07" --- - [Manage Your Organization](manage-organization/) View details about your Redpanda Cloud organization and delete your Serverless organization from the Manage organization page. - [Redpanda CLI](rpk/) The `rpk` tool is a single binary application that provides a way to interact with your Redpanda clusters from the command line. - [Cluster Maintenance](cluster-maintenance/) Learn about cluster maintenance and configuration properties. - [Mountable Topics](mountable-topics/) Safely attach and detach Tiered Storage topics to and from a cluster. - [Integrate Redpanda with Iceberg](iceberg/) Generate Iceberg tables for your Redpanda topics for data lakehouse access. - [Schema Registry](schema-reg/) Redpanda's Schema Registry provides the interface to store and manage event schemas. - [Disaster Recovery](disaster-recovery/) Learn about disaster recovery options for Redpanda Cloud. - [Redpanda Cloud API](api/) Use REST APIs to manage Redpanda Cloud resources. - [Redpanda Terraform Provider](terraform-provider/) Use the Redpanda Terraform provider to create and manage Redpanda Cloud resources. - [Monitor Redpanda Cloud](monitor-cloud/) Learn how to configure monitoring on your BYOC or Dedicated cluster to maintain system health and optimize performance. --- # Page 378: Redpanda Cloud API **URL**: https://docs.redpanda.com/cloud-data-platform/manage/api.md --- # Redpanda Cloud API > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Redpanda Cloud API latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: api/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: api/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/api/index.adoc description: Use REST APIs to manage Redpanda Cloud resources. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-03-20" --- - [Use the Control Plane API](controlplane/) Use the Control Plane API to manage resources in your Redpanda Cloud organization. - [Use the Data Plane APIs](cloud-dataplane-api/) Use the Data Plane APIs to manage your Redpanda Cloud clusters. --- # Page 379: Use the Control Plane API with BYOC **URL**: https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api.md --- # Use the Control Plane API with BYOC > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use the Control Plane API with BYOC latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: api/cloud-byoc-controlplane-api page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: api/cloud-byoc-controlplane-api.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/api/cloud-byoc-controlplane-api.adoc description: Use the Control Plane API to manage resources in your Redpanda Cloud BYOC environment. page-git-created-date: "2024-08-01" page-git-modified-date: "2025-03-20" --- The Redpanda Cloud API is a collection of REST APIs that allow you to interact with different parts of Redpanda Cloud. The Control Plane API enables you to programmatically manage your organization’s Redpanda infrastructure outside of the Cloud UI. You can call the API endpoints directly, or use tools like Terraform or Python scripts to automate cluster management. See [Control Plane API](https://docs.redpanda.com/api/doc/cloud-controlplane/) for the full API reference documentation. ## [](#control-plane-api)Control Plane API The Control Plane API is one central API that allows you to provision clusters, networks, and resource groups. The Control Plane API consists of the following endpoint groups: - [Clusters](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-clusters) - [Networks](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-networks) - [Operations](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-operations) - [Resource Groups](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-resource-groups) - [Control Plane Role Bindings](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-role-bindings) - [Control Plane Users](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-users) - [Control Plane Service Accounts](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-service-accounts) ## [](#lro)Long-running operations Some endpoints do not directly return the resource itself, but instead return an operation. The following is an example response of [`POST /clusters`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_createcluster): ```bash { "operation": { "id": "cqfc6vdmvio001r4vu4", "metadata": { "@type": "type.googleapis.com/redpanda.api.controlplane.v1.CreateClusterMetadata", "cluster_id": "cqg168balf4e4pm8ptu" }, "state": "STATE_IN_PROGRESS", "started_at": "2024-07-23T20:31:29.948Z", "type": "TYPE_CREATE_CLUSTER", "resource_id": "cqg168balf4e4pm8ptu" } } ``` The response object represents the long-running operation of creating a cluster. Cluster creation is an example of an operation that can take a longer period of time to complete. ### [](#check-operation-state)Check operation state To check the progress of an operation, make a request to the [`GET /operations/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-operationservice_getoperation) endpoint using the operation ID as a parameter: ```bash curl -H "Authorization: Bearer " https://api.redpanda.com/v1/operations/ ``` > 💡 **TIP** > > When using a shell substitution variable for the token, use double quotes to wrap the header value. The response contains the current state of the operation: `IN_PROGRESS`, `COMPLETED`, or `FAILED`. ## [](#cluster-tiers)Cluster tiers When you create a BYOC or Dedicated cluster, you select a usage tier. Each tier provides tested and guaranteed workload configurations for throughput, partitions (pre-replication), and connections. Availability depends on the region and the cluster type. See the full list of regions, zones, and tiers available with each provider in the [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-regions-and-usage-tiers). ## [](#create-a-cluster)Create a cluster To create a new cluster, first create a resource group and network, if you have not already done so. ### [](#create-a-resource-group)Create a resource group Create a resource group by making a POST request to the [`/v1/resource-groups`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-resourcegroupservice_createresourcegroup) endpoint. Pass a name for your resource group in the request body. ```bash curl -H 'Content-Type: application/json' \ -H "Authorization: Bearer " \ -d '{ "resource_group": { "name": "" } }' -X POST https://api.redpanda.com/v1/resource-groups ``` A resource group ID is returned. Pass this ID later when you call the Create Cluster endpoint. ### [](#create-a-network)Create a network Create a network by making a request to [`POST /v1/networks`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-networkservice_createnetwork). Choose a [CIDR range](https://docs.redpanda.com/cloud-data-platform/networking/cidr-ranges/) that does not overlap with your existing VPCs or your Redpanda network. ```bash curl -d \ '{ "network": { "cidr_block": "10.0.0.0/20", "cloud_provider": "CLOUD_PROVIDER_GCP", "cluster_type": "TYPE_BYOC", "name": "", "resource_group_id": "", "region": "us-west1" } }' -H "Content-Type: application/json" \ -H "Authorization: Bearer " -X POST https://api.redpanda.com/v1/networks ``` This endpoint returns a [long-running operation](#lro). ### [](#create-a-new-cluster)Create a new cluster After the network is created, make a request to the [`POST /v1/clusters`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_createcluster) with the resource group ID and network ID in the request body. ```bash curl -d \ '{ "cluster": { "cloud_provider": "CLOUD_PROVIDER_GCP", "connection_type": "CONNECTION_TYPE_PUBLIC", "name": "my-new-cluster", "resource_group_id": "", "network_id": "", "region": "us-west1", "throughput_tier": "", "type": "TYPE_BYOC", "zones": [ "us-west1-a", "us-west1-b", "us-west1-c" ], "cluster_configuration": { "custom_properties": { "audit_enabled":true } } } }' -H "Content-Type: application/json" \ -H "Authorization: Bearer " -X POST https://api.redpanda.com/v1/clusters ``` Replace `` with a usage tier that is valid for your region and cluster type. For example, `tier-1-gcp-v2-x86`. See the [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-regions-and-usage-tiers) for the full list of regions, zones, and tiers. The Create Cluster endpoint returns a [long-running operation](#lro). When the operation completes, you can retrieve cluster details by calling [`GET /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_getcluster), and passing the cluster ID as a parameter. #### [](#additional-steps-to-create-a-byoc-cluster)Additional steps to create a BYOC cluster 1. Ensure that you have installed `rpk`. 2. After making a Create Cluster request, run `rpk cloud byoc`. Pass `metadata.cluster_id` from the Create Cluster response: ##### AWS ```bash rpk cloud byoc aws apply --redpanda-id= ``` ##### Azure ```bash rpk cloud byoc azure apply --redpanda-id= --subscription-id= ``` ##### GCP ```bash rpk cloud byoc gcp apply --redpanda-id= --project-id= ``` ## [](#update-cluster-configuration)Update cluster configuration To update your cluster configuration properties, make a request to the [`PATCH /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) endpoint, passing the cluster ID as a parameter. Include the properties to update in the request body. ```bash curl -H "Authorization: Bearer " \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{ "cluster_configuration": { "custom_properties": { "iceberg_enabled":true, "iceberg_catalog_type":"rest" } } }' -X PATCH "https://api.cloud.redpanda.com/v1/clusters/" ``` The Update Cluster endpoint returns a [long-running operation](#lro). [Check the operation state](#check-operation-state) to verify that the update is complete. ## [](#delete-a-cluster)Delete a cluster To delete a cluster, make a request to the [`DELETE /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_deletecluster) endpoint, passing the cluster ID as a parameter. This is a [long-running operation](#lro). ```bash curl -H "Authorization: Bearer " -X DELETE https://api.redpanda.com/v1/clusters/ ``` ### [](#additional-steps-to-delete-a-byoc-cluster)Additional steps to delete a BYOC cluster 1. Make a request to [`GET /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_getcluster) to check the state of the cluster. Wait until the state is `STATE_DELETING_AGENT`. 2. After the state changes to `STATE_DELETING_AGENT`, run `rpk cloud byoc` to destroy the agent. #### AWS ```bash rpk cloud byoc aws destroy --redpanda-id= ``` #### Azure ```bash rpk cloud byoc azure destroy --redpanda-id= ``` #### GCP ```bash rpk cloud byoc gcp destroy --redpanda-id= --project-id= ``` 3. When the cluster is deleted, the delete operation’s state changes to `STATE_COMPLETED`. At this point, you may make a DELETE request to the [`/v1/networks/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-networkservice_deletenetwork) endpoint to delete the network. This is a long running operation. 4. Optional: After the network is deleted, make a request to [`DELETE /v1/resource-groups/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-resourcegroupservice_deleteresourcegroup) to delete the resource group. ## [](#manage-rbac)Manage RBAC You can also use the Control Plane API to manage [RBAC configurations](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/). ### [](#list-role-bindings)List role bindings To see role assignments for IAM user and service accounts, make a GET request to the [`/v1/role-bindings`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_listrolebindings) endpoint. ```bash curl https://api.redpanda.com/v1/role-bindings?filter.role_name=&filter.scope.resource_type=SCOPE_RESOURCE_TYPE_CLUSTER \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#get-role-binding)Get role binding To see roles assignments for a specific IAM account, make a GET request to the [`/v1/role-bindings/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_getrolebinding) endpoint, passing the role binding ID as a parameter. ```bash curl "https://api.redpanda.com/v1/role-bindings/ \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#get-user)Get user To see details of an IAM user account, make a GET request to the [`/v1/users/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-userservice_getuser) endpoint, passing the user account ID as a parameter. ```bash curl "https://api.redpanda.com/v1/users/ \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#create-role-binding)Create role binding To assign a role to an IAM user or service account, make a POST request to the [`/v1/role-bindings`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_createrolebinding) endpoint. Specify the role and scope, which includes the specific resource ID and an optional resource type, in the request body. ```bash curl -X POST "https://api.redpanda.com/v1/role-bindings" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "role_name": "", "account_id": "", "scope": { "resource_type": "SCOPE_RESOURCE_TYPE_CLUSTER", "resource_id": "" } }' ``` For ``, use one of roles listed in [Predefined roles](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/#predefined-roles) (`Reader`, `Writer`, `Admin`). ### [](#create-service-account)Create service account > 📝 **NOTE** > > Service accounts are assigned the Admin role for all resources in the organization. To create a new service account, make a POST request to the [`/v1/service-accounts`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serviceaccountservice_createserviceaccount) endpoint, with a service account name and optional description in the request body. ```bash curl -X POST "https://api.redpanda.com/v1/service-accounts" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "service_account": { "name": "", "description": "" } }' ``` ## [](#next-steps)Next steps - [Use the Data Plane APIs](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dataplane-api/) --- # Page 380: Use the Data Plane APIs **URL**: https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dataplane-api.md --- # Use the Data Plane APIs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use the Data Plane APIs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: api/cloud-dataplane-api page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: api/cloud-dataplane-api.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/api/cloud-dataplane-api.adoc description: Use the Data Plane APIs to manage your Redpanda Cloud clusters. page-git-created-date: "2024-06-06" page-git-modified-date: "2025-08-20" --- The Redpanda Cloud API is a collection of REST APIs that allow you to interact with different parts of Redpanda Cloud. The Data Plane APIs enable you to programmatically manage the resources within your clusters, including topics, users, access control lists (ACLs), and connectors. You can call the API endpoints directly, or use tools like Terraform or Python scripts to automate resource management. See [Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/) for the full Data Plane API reference documentation. The [data plane](https://docs.redpanda.com/api/doc/cloud-dataplane/topic/topic-cloud-api-overview#topic-cloud-api-architecture) contains the actual Redpanda clusters. Every cluster is its own data plane, and so it has its own distinct [Data Plane API URL](https://docs.redpanda.com/api/doc/cloud-dataplane/topic/topic-cloud-api-overview#topic-data-plane-apis-url). ## [](#get-data-plane-api-url)Get Data Plane API URL ### BYOC or Dedicated To retrieve the Data Plane API URL of a cluster, make a request to the [`GET /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_getcluster) endpoint of the Control Plane API. ### Serverless To retrieve the Data Plane API URL of a cluster, make a request to the [`GET /v1/serverless/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serverlessclusterservice_getserverlesscluster) endpoint of the Control Plane API. The response includes a `dataplane_api.url` value: ```bash "id": "....", "name": "my-cluster", .... "dataplane_api": { "url": "https://api-xyz.abc.fmc.ppd.cloud.redpanda.com" }, ... ``` ## [](#data-plane-apis)Data Plane APIs ### [](#create-a-user)Create a user To create a new user in your Redpanda cluster, make a POST request to the [`/v1/users`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-userservice_createuser) endpoint, including the SASL mechanism, username, and password in the request body: ```bash curl -X POST "https:///v1/users" \ -H "Authorization: Bearer " \ -H "accept: application/json" \ -H "content-type: application/json" \ -d '{"mechanism":"SASL_MECHANISM_SCRAM_SHA_256","name":"payment-service","password":"secure-password"}' ``` > 💡 **TIP** > > When using a shell substitution variable for the token, use double quotes to wrap the header value. The success response returns the newly-created username and SASL mechanism: { "user": { "name": "payment-service", "mechanism": "SASL\_MECHANISM\_SCRAM\_SHA\_256" } } ### [](#create-an-acl)Create an ACL To create a new ACL in your Redpanda cluster, make a [`POST /v1/acls`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-aclservice_createacl) request. The following example ACL allows all operations on any Redpanda topic for a user with the name `payment-service`. ```bash curl -X POST "https:///v1/acls" \ -H "Authorization: Bearer " \ -H "accept: application/json" \ -H "content-type: application/json" \ -d '{"host":"*","operation":"OPERATION_ALL","permission_type":"PERMISSION_TYPE_ALLOW","principal":"User:payment-service","resource_name":"*","resource_pattern_type":"RESOURCE_PATTERN_TYPE_LITERAL","resource_type":"RESOURCE_TYPE_TOPIC"}' ``` The success response is empty, with a 201 status code. {} ### [](#create-a-topic)Create a topic To create a new Redpanda topic without specifying any further parameters, such as the desired topic-level configuration or partition count, make a POST request to [`/v1/topics`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-topicservice_createtopic) endpoint: ```bash curl -X POST "/v1/topics" \ -H "Authorization: Bearer " \ -H "accept: application/json" \ -H "content-type: application/json" \ -d '{"name":""}' ``` ### [](#manage-secrets)Manage secrets Secrets are stored externally in your cloud provider’s secret management service. Redpanda fetches the secrets when you reference them in cluster properties. #### [](#create-a-secret)Create a secret Make a request to [`POST /v1/secrets`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-secretservice_createsecret). You must use a Base64-encoded secret. ```bash curl -X POST "https:///v1/secrets" \ -H "accept: application/json" \ -H "authorization: Bearer " \ -H "content-type: application/json" \ -d '{"id":"","scopes":["SCOPE_REDPANDA_CLUSTER"],"secret_data":""}' ``` You must include the following values: - ``: The base URL for the Data Plane API. - ``: The API key you generated during authentication. - ``: The name of the secret you want to add. Use only the following characters: `^[A-Z][A-Z0-9_]*$`. - ``: The Base64-encoded secret. - This scope: `"SCOPE_REDPANDA_CLUSTER"`. The response returns the name and scope of the secret. You can then use the Control Plane API or `rpk` to [set a cluster property value](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/) to reference a secret, using the secret name. For the Control Plane API, you must use the following notation with the secret name in the request body to correctly reference the secret: ```bash "iceberg_rest_catalog_client_secret": "${secrets.}" ``` #### [](#update-a-secret)Update a secret Make a request to [`PUT /v1/secrets/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-secretservice_updatesecret). You can only update the secret value, not its name. You must use a Base64-encoded secret. ```bash curl -X PUT "https:///v1/secrets/" \ -H "accept: application/json" \ -H "authorization: Bearer " \ -H "content-type: application/json" \ -d '{"scopes":["SCOPE_REDPANDA_CLUSTER"],"secret_data":""}' ``` You must include the following values: - ``: The base URL for the Data Plane API. - ``: The name of the secret you want to update. The secret’s name is also its ID. - ``: The API key you generated during authentication. - This scope: `"SCOPE_REDPANDA_CLUSTER"`. - ``: Your new Base64-encoded secret. The response returns the name and scope of the secret. It might take several minutes for the new secret value to propagate to any cluster properties that reference it. #### [](#delete-a-secret)Delete a secret Before you delete a secret, make sure that you remove references to it from your cluster configuration. Make a request to [`DELETE /v1/secrets/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-secretservice_deletesecret). ```bash curl -X DELETE "https:///v1/secrets/" \ -H "accept: application/json" \ -H "authorization: Bearer " \ ``` You must include the following values: - ``: The base URL for the Data Plane API. - ``: The name of the secret you want to delete. - ``: The API key you generated during authentication. ### [](#use-redpanda-connect)Use Redpanda Connect Use the API to manage [Redpanda Connect pipelines](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/) in Redpanda Cloud. > 📝 **NOTE** > > The Pipeline APIs for Redpanda Connect are supported in BYOC and Serverless clusters only. #### [](#get-redpanda-connect-pipeline)Get Redpanda Connect pipeline To get details of a specific pipeline, make a [`GET /v1/redpanda-connect/pipelines/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_getpipeline) request. ```bash curl "https:///v1/redpanda-connect/pipelines/" ``` #### [](#stop-a-redpanda-connect-pipeline)Stop a Redpanda Connect pipeline To stop a running pipeline, make a [`PUT /v1/redpanda-connect/pipelines/{id}/stop`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_stoppipeline) request. ```bash curl -X PUT "https:///v1/redpanda-connect/pipelines//stop" ``` #### [](#start-a-redpanda-connect-pipeline)Start a Redpanda Connect pipeline To start a previously stopped pipeline, make a [`PUT /v1/redpanda-connect/pipelines/{id}/start`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_startpipeline) request. ```bash curl -X PUT "https:///v1/redpanda-connect/pipelines//start" ``` #### [](#update-a-redpanda-connect-pipeline)Update a Redpanda Connect pipeline To update a pipeline, make a [`PUT /v1/redpanda-connect/pipelines/{id}`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-redpandaconnectservice_updatepipeline) request. You update a pipeline configuration to scale resources, for example the number of CPU cores and amount of memory allocated. ```bash curl -X PUT "https://api.redpanda.com/v1/redpanda-connect/pipelines/" \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{"resources":{"cpu_shares":"8","memory_shares":"8G"}}' ``` ### [](#manage-kafka-connect)Manage Kafka Connect Use the API to configure your [Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/) clusters. > ❗ **IMPORTANT** > > - To enable this feature, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). To disable this feature, see [Disable Kafka Connect](https://docs.redpanda.com/cloud-data-platform/develop/managed-connectors/disable-kc/). > > - Redpanda Support does not manage or monitor Kafka Connect. For fully-supported connectors, consider [Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/). > > - When Kafka Connect is enabled, there is a dedicated node running even when no connectors are deployed. > 📝 **NOTE** > > Kafka Connect is supported in BYOC and Dedicated clusters only. #### [](#create-a-kafka-connect-cluster-secret)Create a Kafka Connect cluster secret Kafka Connect cluster secret data must first be in JSON format, and then Base64-encoded. 1. Prepare the secret data in JSON format: ```none {"secret.access.key": ""} ``` 2. Encode the secret data in Base64: ```none echo '{"secret.access.key": ""}' | base64 ``` 3. Use the [Secrets API](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-kafkaconnectservice_createsecret) to create a secret that stores the Base64-encoded secret data: ```bash curl -X POST "https:///v1/kafka-connect/clusters/redpanda/secrets" \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{"name":"","secret_data":""}' ``` The response returns an `id` that you can use to [create the Kafka Connect connector](#create-a-kafka-connect-connector). #### [](#create-a-kafka-connect-connector)Create a Kafka Connect connector To create a connector, make a POST request to [`/v1/kafka-connect/clusters/{cluster_name}/connectors`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-kafkaconnectservice_createconnector). The following example shows how to create an S3 sink connector with the name `my-connector`: ```bash curl -X POST "/v1/kafka-connect/clusters/redpanda/connectors" \ -H "Authorization: Bearer " \ -H "accept: application/json" \ -H "content-type: application/json" \ -d '{"config":{"connector.class":"com.redpanda.kafka.connect.s3.S3SinkConnector","topics":"test-topic","aws.secret.access.key":"${secretsManager::secret.access.key}","aws.s3.bucket.name":"bucket-name","aws.access.key.id":"access-key","aws.s3.bucket.check":"false","region":"us-east-1"},"name":"my-connector"}' ``` > ⚠️ **CAUTION** > > The field `aws.secret.access.key` in this example contains sensitive information that usually shouldn’t be added to a configuration directly. Redpanda recommends that you first create a secret and then use the secret ID to inject the secret in your Create Connector request. > > If you had created a secret following the example from the previous section [Create a Kafka Connect cluster secret](#create-a-kafka-connect-cluster-secret), use the `id` returned in the Create Secret response to replace the placeholder `` in this Create Connector example. The syntax `${secretsManager::secret.access.key}` tells the Kafka Connect cluster to load ``, specifying the key `secret.access.key` from the secret JSON. Example success response: { "name": "my-connector", "config": { "aws.access.key.id": "access-key", "aws.s3.bucket.check": "false", "aws.s3.bucket.name": "bucket-name", "aws.secret.access.key": "secret-key", "connector.class": "com.redpanda.kafka.connect.s3.S3SinkConnector", "name": "my-connector", "region": "us-east-1", "topics": "test-topic" }, "tasks": \[\], "type": "sink" } #### [](#restart-a-kafka-connect-connector)Restart a Kafka Connect connector To restart a connector, make a POST request to the [`/v1/kafka-connect/clusters/{cluster_name}/connectors/{name}/restart`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-kafkaconnectservice_restartconnector) endpoint: ```bash curl -X POST "/v1/kafka-connect/clusters/redpanda/connectors/my-connector/restart" \ -H "Authorization: Bearer " \ -H "accept: application/json"\ -H "content-type: application/json" \ -d '{"include_tasks":false,"only_failed":false}' ``` ## [](#limitations)Limitations - Client SDKs are not available. --- # Page 381: Use the Control Plane API with Dedicated Cloud **URL**: https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dedicated-controlplane-api.md --- # Use the Control Plane API with Dedicated Cloud > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use the Control Plane API with Dedicated Cloud latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: api/cloud-dedicated-controlplane-api page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: api/cloud-dedicated-controlplane-api.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/api/cloud-dedicated-controlplane-api.adoc description: Use the Control Plane API to manage resources in your Redpanda Cloud Dedicated environment. page-git-created-date: "2024-08-01" page-git-modified-date: "2025-03-20" --- The Redpanda Cloud API is a collection of REST APIs that allow you to interact with different parts of Redpanda Cloud. The Control Plane API enables you to programmatically manage your organization’s Redpanda infrastructure outside of the Cloud UI. You can call the API endpoints directly, or use tools like Terraform or Python scripts to automate cluster management. See [Control Plane API](https://docs.redpanda.com/api/doc/cloud-controlplane/) for the full API reference documentation. ## [](#control-plane-api)Control Plane API The Control Plane API is one central API that allows you to provision clusters, networks, and resource groups. The Control Plane API consists of the following endpoint groups: - [Clusters](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-clusters) - [Networks](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-networks) - [Operations](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-operations) - [Resource Groups](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-resource-groups) - [Control Plane Role Bindings](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-role-bindings) - [Control Plane Users](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-users) - [Control Plane Service Accounts](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-service-accounts) ## [](#lro)Long-running operations Some endpoints do not directly return the resource itself, but instead return an operation. The following is an example response of [`POST /clusters`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_createcluster): ```bash { "operation": { "id": "cqfc6vdmvio001r4vu4", "metadata": { "@type": "type.googleapis.com/redpanda.api.controlplane.v1.CreateClusterMetadata", "cluster_id": "cqg168balf4e4pm8ptu" }, "state": "STATE_IN_PROGRESS", "started_at": "2024-07-23T20:31:29.948Z", "type": "TYPE_CREATE_CLUSTER", "resource_id": "cqg168balf4e4pm8ptu" } } ``` The response object represents the long-running operation of creating a cluster. Cluster creation is an example of an operation that can take a longer period of time to complete. ### [](#check-operation-state)Check operation state To check the progress of an operation, make a request to the [`GET /operations/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-operationservice_getoperation) endpoint using the operation ID as a parameter: ```bash curl -H "Authorization: Bearer " https://api.redpanda.com/v1/operations/ ``` > 💡 **TIP** > > When using a shell substitution variable for the token, use double quotes to wrap the header value. The response contains the current state of the operation: `IN_PROGRESS`, `COMPLETED`, or `FAILED`. ## [](#cluster-tiers)Cluster tiers When you create a BYOC or Dedicated cluster, you select a usage tier. Each tier provides tested and guaranteed workload configurations for throughput, partitions (pre-replication), and connections. Availability depends on the region and the cluster type. See the full list of regions, zones, and tiers available with each provider in the [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-regions-and-usage-tiers). ## [](#create-a-cluster)Create a cluster To create a new cluster, first create a resource group and network, if you have not already done so. ### [](#create-a-resource-group)Create a resource group Create a resource group by making a POST request to the [`/v1/resource-groups`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-resourcegroupservice_createresourcegroup) endpoint. Pass a name for your resource group in the request body. ```bash curl -H 'Content-Type: application/json' \ -H "Authorization: Bearer " \ -d '{ "resource_group": { "name": "" } }' -X POST https://api.redpanda.com/v1/resource-groups ``` A resource group ID is returned. Pass this ID later when you call the Create Cluster endpoint. ### [](#create-a-network)Create a network Create a network by making a request to [`POST /v1/networks`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-networkservice_createnetwork). Choose a [CIDR range](https://docs.redpanda.com/cloud-data-platform/networking/cidr-ranges/) that does not overlap with your existing VPCs or your Redpanda network. ```bash curl -d \ '{ "network": { "cidr_block": "10.0.0.0/20", "cloud_provider": "CLOUD_PROVIDER_GCP", "cluster_type": "TYPE_DEDICATED", "name": "", "resource_group_id": "", "region": "us-central1" } }' -H "Content-Type: application/json" \ -H "Authorization: Bearer " -X POST https://api.redpanda.com/v1/networks ``` This endpoint returns a [long-running operation](#lro). ### [](#create-a-new-cluster)Create a new cluster After the network is created, make a request to the [`POST /v1/clusters`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_createcluster) with the resource group ID and network ID in the request body. ```bash curl -d \ '{ "cluster": { "cloud_provider": "CLOUD_PROVIDER_GCP", "connection_type": "CONNECTION_TYPE_PUBLIC", "name": "my-new-cluster", "resource_group_id": "", "network_id": "", "region": "us-central1", "throughput_tier": "", "type": "TYPE_DEDICATED", "zones": [ "us-central1-a", "us-central1-b", "us-central1-c" ], "cluster_configuration": { "custom_properties": { "audit_enabled":true } } } }' -H "Content-Type: application/json" \ -H "Authorization: Bearer " -X POST https://api.redpanda.com/v1/clusters ``` Replace `` with a usage tier that is valid for your region and cluster type. For example, `tier-1-gcp-v2-x86`. See the [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-regions-and-usage-tiers) for the full list of regions, zones, and tiers. The Create Cluster endpoint returns a [long-running operation](#lro). When the operation completes, you can retrieve cluster details by calling [`GET /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_getcluster), and passing the cluster ID as a parameter. ## [](#update-cluster-configuration)Update cluster configuration To update your cluster configuration properties, make a request to the [`PATCH /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) endpoint, passing the cluster ID as a parameter. Include the properties to update in the request body. ```bash curl -H "Authorization: Bearer " \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{ "cluster_configuration": { "custom_properties": { "audit_enabled":true } } }' -X PATCH "https://api.cloud.redpanda.com/v1/clusters/" ``` The Update Cluster endpoint returns a [long-running operation](#lro). [Check the operation state](#check-operation-state) to verify that the update is complete. ## [](#delete-a-cluster)Delete a cluster To delete a cluster, make a request to the [`DELETE /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_deletecluster) endpoint, passing the cluster ID as a parameter. This is a [long-running operation](#lro). ```bash curl -H "Authorization: Bearer " -X DELETE https://api.redpanda.com/v1/clusters/ ``` ## [](#manage-rbac)Manage RBAC You can also use the Control Plane API to manage [RBAC configurations](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/). ### [](#list-role-bindings)List role bindings To see role assignments for IAM user and service accounts, make a GET request to the [`/v1/role-bindings`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_listrolebindings) endpoint. ```bash curl https://api.redpanda.com/v1/role-bindings?filter.role_name=&filter.scope.resource_type=SCOPE_RESOURCE_TYPE_CLUSTER \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#get-role-binding)Get role binding To see roles assignments for a specific IAM account, make a GET request to the [`/v1/role-bindings/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_getrolebinding) endpoint, passing the role binding ID as a parameter. ```bash curl "https://api.redpanda.com/v1/role-bindings/ \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#get-user)Get user To see details of an IAM user account, make a GET request to the [`/v1/users/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-userservice_getuser) endpoint, passing the user account ID as a parameter. ```bash curl "https://api.redpanda.com/v1/users/ \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#create-role-binding)Create role binding To assign a role to an IAM user or service account, make a POST request to the [`/v1/role-bindings`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_createrolebinding) endpoint. Specify the role and scope, which includes the specific resource ID and an optional resource type, in the request body. ```bash curl -X POST "https://api.redpanda.com/v1/role-bindings" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "role_name": "", "account_id": "", "scope": { "resource_type": "SCOPE_RESOURCE_TYPE_CLUSTER", "resource_id": "" } }' ``` For ``, use one of roles listed in [Predefined roles](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/#predefined-roles) (`Reader`, `Writer`, `Admin`). ### [](#create-service-account)Create service account > 📝 **NOTE** > > Service accounts are assigned the Admin role for all resources in the organization. To create a new service account, make a POST request to the [`/v1/service-accounts`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serviceaccountservice_createserviceaccount) endpoint, with a service account name and optional description in the request body. ```bash curl -X POST "https://api.redpanda.com/v1/service-accounts" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "service_account": { "name": "", "description": "" } }' ``` ## [](#next-steps)Next steps - [Use the Data Plane APIs](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dataplane-api/) --- # Page 382: Use the Control Plane API with Serverless **URL**: https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-serverless-controlplane-api.md --- # Use the Control Plane API with Serverless > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use the Control Plane API with Serverless latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: api/cloud-serverless-controlplane-api page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: api/cloud-serverless-controlplane-api.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/api/cloud-serverless-controlplane-api.adoc description: Use the Control Plane API to manage resources in your Redpanda Serverless environment. page-git-created-date: "2024-08-01" page-git-modified-date: "2025-03-20" --- The Redpanda Cloud API is a collection of REST APIs that allow you to interact with different parts of Redpanda Cloud. The Control Plane API enables you to programmatically manage your organization’s Redpanda infrastructure outside of the Cloud UI. You can call the API endpoints directly, or use tools like Terraform or Python scripts to automate cluster management. See [Control Plane API](https://docs.redpanda.com/api/doc/cloud-controlplane/) for the full API reference documentation. ## [](#control-plane-api)Control Plane API The Control Plane API is one central API that allows you to provision clusters, networks, and resource groups. The Control Plane API consists of the following endpoint groups: - [Operations](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-operations) - [Resource Groups](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-resource-groups) - [Serverless Clusters](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-serverless-clusters) - [Serverless Regions](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-serverless-regions) - [Control Plane Role Bindings](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-role-bindings) - [Control Plane Users](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-users) - [Control Plane Service Accounts](https://docs.redpanda.com/api/doc/cloud-controlplane/group/endpoint-control-plane-service-accounts) ## [](#create-a-cluster)Create a cluster To create a new serverless cluster, you can use the default resource group, or create a new resource group if you like. You need to choose a region where your cluster is hosted. ### [](#create-a-resource-group)Create a resource group > 📝 **NOTE** > > This step is optional. Serverless includes a default resource group. To retrieve the default resource group ID, make a GET request to the [`/v1/resource-groups`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-resourcegroupservice_listresourcegroups) endpoint: > > ```bash > curl -H "Authorization: Bearer " https://api.redpanda.com/v1/resource-groups > ``` Create a resource group by making a POST request to the [`/v1/resource-groups`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-resourcegroupservice_createresourcegroup) endpoint. Pass a name for your resource group in the request body. ```bash curl -H 'Content-Type: application/json' \ -H "Authorization: Bearer " \ -d '{ "name": "" }' -X POST https://api.redpanda.com/v1/resource-groups ``` A resource group ID is returned. Pass this ID later when you call the Create Serverless Cluster endpoint. ### [](#choose-a-region)Choose a region To see the available regions for Redpanda Serverless, make a GET request to the [`/v1/serverless/regions`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serverlessregionservice_listserverlessregions) endpoint. You can specify a cloud provider in your request. Serverless currently only supports AWS. ```bash curl -H "Authorization: Bearer " 'https://api.redpanda.com/v1/serverless/regions?cloud_provider=CLOUD_PROVIDER_AWS' ``` > 💡 **TIP** > > When using a shell substitution variable for the token, use double quotes to wrap the header value. ```json { "serverless_regions": [ { "name": "eu-central-1", "display_name": "eu-central-1", "default_timezone": { "id": "Europe/Berlin", "version": "" }, "cloud_provider": "CLOUD_PROVIDER_AWS", "available": true }, ... ], "next_page_token": "" } ``` You can also see a list of supported regions in [Serverless regions](https://docs.redpanda.com/cloud-data-platform/reference/tiers/serverless-regions/). ### [](#create-a-new-serverless-cluster)Create a new serverless cluster Create a Serverless cluster by making a request to [`POST /v1/serverless/clusters`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serverlessclusterservice_createserverlesscluster) with the resource group ID and serverless region name in the request body. ```bash curl -H 'Content-Type: application/json' \ -H "Authorization: Bearer " \ -d '{ "serverless_cluster": { "name": "", "resource_group_id": "", "serverless_region": "us-east-1" } }' -X POST https://api.redpanda.com/v1/serverless/clusters ``` The Create Serverless Cluster endpoint returns a [long-running operation](#lro-serverless). When the operation completes, you can retrieve cluster details by calling [`GET /v1/serverless/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serverlessclusterservice_getserverlesscluster), and passing the cluster ID as a parameter. ## [](#update-cluster-configuration)Update cluster configuration To update your cluster configuration properties, make a request to the [`PATCH /v1/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) endpoint, passing the cluster ID as a parameter. Include the properties to update in the request body. ```bash curl -H "Authorization: Bearer " \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{ "cluster_configuration": { "custom_properties": { "audit_enabled":true } } }' -X PATCH "https://api.cloud.redpanda.com/v1/clusters/" ``` The Update Cluster endpoint returns a [long-running operation](#lro). [Check the operation state](#check-operation-state) to verify that the update is complete. ## [](#delete-a-cluster)Delete a cluster To delete a cluster, make a request to the [`DELETE /v1/serverless/clusters/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serverlessclusterservice_getserverlesscluster) endpoint, passing the cluster ID as a parameter. This is a [long-running operation](#lro-serverless). ```bash curl -H "Authorization: Bearer " -X DELETE https://api.redpanda.com/v1/serverless/clusters/ ``` Optional: When the cluster is deleted, the delete operation’s state changes to `STATE_COMPLETED`. At this point, you may make a DELETE request to the [`/v1/resource-groups/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-resourcegroupservice_deleteresourcegroup) endpoint to delete the resource group. ## [](#lro-serverless)Long-running operations Some endpoints do not directly return the resource itself, but instead return an operation. The following is an example response of [`POST /serverless/clusters`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serverlessclusterservice_createserverlesscluster): ```bash { "operation": { "id": "cqaramrndjr40k3qei50", "metadata": null, "state": "STATE_IN_PROGRESS", "started_at": { "seconds": "1721087323", "nanos": 888601218 }, "finished_at": null, "type": "TYPE_CREATE_SERVERLESS_CLUSTER" } } ``` The response object represents the long-running operation of creating a cluster. Cluster creation is an example of an operation that can take a longer period of time to complete. ### [](#check-operation-state)Check operation state To check the progress of an operation, make a request to the [`GET /operations/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-operationservice_getoperation) endpoint using the operation ID as a parameter: ```bash curl -H "Authorization: Bearer " https://api.redpanda.com/v1/operations/ ``` The response contains the current state of the operation: `IN_PROGRESS`, `COMPLETED`, or `FAILED`. ## [](#manage-rbac)Manage RBAC You can also use the Control Plane API to manage [RBAC configurations](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/). ### [](#list-role-bindings)List role bindings To see role assignments for IAM user and service accounts, make a GET request to the [`/v1/role-bindings`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_listrolebindings) endpoint. ```bash curl https://api.redpanda.com/v1/role-bindings?filter.role_name=&filter.scope.resource_type=SCOPE_RESOURCE_TYPE_CLUSTER \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#get-role-binding)Get role binding To see roles assignments for a specific IAM account, make a GET request to the [`/v1/role-bindings/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_getrolebinding) endpoint, passing the role binding ID as a parameter. ```bash curl "https://api.redpanda.com/v1/role-bindings/ \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#get-user)Get user To see details of an IAM user account, make a GET request to the [`/v1/users/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-userservice_getuser) endpoint, passing the user account ID as a parameter. ```bash curl "https://api.redpanda.com/v1/users/ \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" ``` ### [](#create-role-binding)Create role binding To assign a role to an IAM user or service account, make a POST request to the [`/v1/role-bindings`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-rolebindingservice_createrolebinding) endpoint. Specify the role and scope, which includes the specific resource ID and an optional resource type, in the request body. ```bash curl -X POST "https://api.redpanda.com/v1/role-bindings" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "role_name": "", "account_id": "", "scope": { "resource_type": "SCOPE_RESOURCE_TYPE_CLUSTER", "resource_id": "" } }' ``` For ``, use one of roles listed in [Predefined roles](https://docs.redpanda.com/cloud-data-platform/security/authorization/rbac/rbac/#predefined-roles) (`Reader`, `Writer`, `Admin`). ### [](#create-service-account)Create service account > 📝 **NOTE** > > Service accounts are assigned the Admin role for all resources in the organization. To create a new service account, make a POST request to the [`/v1/service-accounts`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-serviceaccountservice_createserviceaccount) endpoint, with a service account name and optional description in the request body. ```bash curl -X POST "https://api.redpanda.com/v1/service-accounts" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "service_account": { "name": "", "description": "" } }' ``` ## [](#next-steps)Next steps - [Use the Data Plane APIs](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dataplane-api/) --- # Page 383: Use the Control Plane API **URL**: https://docs.redpanda.com/cloud-data-platform/manage/api/controlplane.md --- # Use the Control Plane API > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use the Control Plane API latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: api/controlplane/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: api/controlplane/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/api/controlplane/index.adoc description: Use the Control Plane API to manage resources in your Redpanda Cloud organization. page-git-created-date: "2024-08-01" page-git-modified-date: "2025-03-20" --- - [Use the Control Plane API with BYOC](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/) Use the Control Plane API to manage resources in your Redpanda Cloud BYOC environment. - [Use the Control Plane API with Dedicated Cloud](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dedicated-controlplane-api/) Use the Control Plane API to manage resources in your Redpanda Cloud Dedicated environment. - [Use the Control Plane API with Serverless](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-serverless-controlplane-api/) Use the Control Plane API to manage resources in your Redpanda Serverless environment. --- # Page 384: Audit Logging **URL**: https://docs.redpanda.com/cloud-data-platform/manage/audit-logging.md --- # Audit Logging > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Audit Logging latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: audit-logging page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: audit-logging.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/audit-logging.adoc description: Learn how to use Redpanda's audit logging capabilities. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- > 📝 **NOTE** > > Audit logging is supported on BYOC and Dedicated clusters running Redpanda version 24.3 and later. To configure audit logging, see [Configure Cluster Properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/). Many scenarios for streaming data include the need for fine-grained auditing of user activity related to the system. This is especially true for regulated industries such as finance, healthcare, and the public sector. Complying with [PCI DSS v4](https://www.pcisecuritystandards.org/document_library/?document=pci_dss) standards, for example, requires verbose and detailed activity auditing, alerting, and analysis capabilities. Redpanda’s auditing capabilities support recording both administrative and operational interactions with topics and with users. Redpanda complies with the Open Cybersecurity Schema Framework (OCSF), providing a predictable and extensible solution that works seamlessly with industry standard tools. With audit logging enabled, there should be no noticeable changes in performance other than slightly elevated CPU usage. ## [](#audit-log-flow)Audit log flow The Redpanda audit log mechanism functions similar to the Kafka flow. When a user interacts with another user or with a topic, Redpanda writes an event to a specialized audit topic. The audit topic is immutable. Only Redpanda can write to it. Users are prevented from writing to the audit topic directly and the Kafka API cannot create or delete it. ![Audit log flow](https://docs.redpanda.com/cloud-data-platform/shared/_images/audit-logging-flow.png) By default, any management and authentication actions performed on the cluster yield messages written to the audit log topic that are retained for seven days. Interactions with all topics by all principals are audited. Actions performed using the Kafka API and Admin API are all audited, as are actions performed directly through `rpk`. Messages recorded to the audit log topic comply with the [open cybersecurity schema framework](https://schema.ocsf.io/). Any number of analytics frameworks, such as Splunk or Sumo Logic, can receive and process these messages. Using an open standard ensures Redpanda’s audit logs coexist with those produced by other IT assets, powering holistic monitoring and analysis of your assets. ## [](#audit-log-configuration-options)Audit log configuration options Redpanda’s audit logging mechanism supports several options to control the volume and availability of audit records. Configuration is applied at the cluster level. To configure audit logging, see [Configure Cluster Properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/). - [`audit_enabled`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#audit_enabled): Boolean value to enable audit logging. When you set this to `true`, Redpanda checks for an existing topic named `_redpanda.audit_log`. If none is found, Redpanda automatically creates one for you. Default: `true`. - [`audit_enabled_event_types`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#audit_enabled_event_types): List of strings in JSON style identifying the event types to include in the audit log. This may include any of the following: `management, produce, consume, describe, heartbeat, authenticate, schema_registry, admin`. Default: `'["management","authenticate","admin"]'`. - [`audit_excluded_principals`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#audit_excluded_principals): List of strings in JSON style identifying the principals the audit logging system should ignore. Principals can be listed as `User:name` or `name`, both are accepted. Default: `null`. ## [](#enable-audit-logging)Enable audit logging Audit logging is enabled by default. Cluster administrators can configure the audited topics and principals. However, only the Redpanda team can configure the type of audited events. For more information or support, contact your Redpanda account team. ## [](#configure-retention-for-audit-logs)Configure retention for audit logs You can export audit events to your SIEM for long-term retention to support audit and compliance needs. Redpanda Data recommends that you retain audit logs for at least one year in a separate system like your SIEM, so if there is an issue with the Redpanda cluster you have access to the audit logs. If you need to change the default seven-day retention period, update the retention settings using the `retention.ms` property for the `_redpanda.audit_log` topic: ```bash # Set 1-year retention (in milliseconds) on the audit log topic rpk topic alter-config _redpanda.audit_log --set retention.ms=31536000000 ``` > 📝 **NOTE** > > In Redpanda Cloud, both `retention.ms` (time-based) and `retention.bytes` (size-based) retention policies are applied simultaneously. Data becomes eligible for deletion when either limit is reached, depending on whichever occurs first. This means neither setting strictly takes precedence; the earliest limit (by time or size) triggers data cleanup. When updating audit log retention, check to make sure you do not already have a size-based retention policy that might remove logs before the period you specify. ## [](#next-steps)Next steps [See samples of audit log messages](audit-log-samples/) --- # Page 385: Sample Audit Log Messages **URL**: https://docs.redpanda.com/cloud-data-platform/manage/audit-logging/audit-log-samples.md --- # Sample Audit Log Messages > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Sample Audit Log Messages latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: audit-logging/audit-log-samples page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: audit-logging/audit-log-samples.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/audit-logging/audit-log-samples.adoc description: Sample Redpanda audit log messages. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-05-26" --- Redpanda’s audit logs comply with version 1.0.0 of the [Open Cybersecurity Schema Framework (OCSF)](https://github.com/ocsf). This provides a predictable and extensible solution that works seamlessly with industry standard tools. This page aggregates several sample log files covering a range of scenarios. ## [](#standard-ocsf-messages)Standard OCSF messages Redpanda produces the following standard OCSF class messages: - Authentication (3002) for all authentication events - Application Lifecycle (6002) for when the audit system is enabled or disabled or when Redpanda starts or stops (if auditing is enabled when Redpanda starts or stops) - API Activity (6003) for any access to the Kafka API, Admin API, or Schema Registry Refer to the [OCSF Schema Definition](https://schema.ocsf.io/) for the field definitions for each event class. ## [](#authentication-events)Authentication events These messages illustrate various scenarios around successful and unsuccessful authentication events. Authentication successful This scenario shows the message resulting from an admin using rpk with successful authentication. This is an authentication type event. ```json { "category_uid": 3, "class_uid": 3002, "metadata": { "product": { "name": "Redpanda", // This is the Node ID of the broker that produced this audit event "uid": "2", "vendor_name": "Redpanda Data, Inc.", "version": "v23.3.0-dev-2457-g76dc896f8c" }, "version": "1.0.0" }, "severity_id": 1, "time": 1700533469078, "type_uid": 300201, "activity_id": 1, "auth_protocol": "SASL-SCRAM", "auth_protocol_id": 99, // This is the IP address of the Kafka broker that received the authorization request "dst_endpoint": { "ip": "127.0.0.1", "port": 19092, // Name of the Redpanda kafka server "svc_name": "kafka rpc protocol" }, // Indicates that credentials were not encrypted using TLS "is_cleartext": true, "is_mfa": false, "service": { "name": "kafka rpc protocol" }, // This is the IP address of the client that generated the authorization request "src_endpoint": { "ip": "127.0.0.1", // This is the client ID of the kafka client "name": "rpk", "port": 42906 }, "status_id": 1, "user": { "name": "user", "type_id": 1 } } ``` Authentication successful (OIDC with group claims) This scenario shows a successful OIDC authentication event that includes the user’s IdP group memberships in the `user.groups` field. Group memberships are extracted from the OIDC token and included in all authentication events for OIDC users. ```json { "category_uid": 3, "class_uid": 3002, "metadata": { "product": { "name": "Redpanda", "uid": "0", "vendor_name": "Redpanda Data, Inc.", "version": "v26.1.1" }, "version": "1.0.0" }, "severity_id": 1, "time": 1700533469078, "type_uid": 300201, "activity_id": 1, "auth_protocol": "SASL-OAUTHBEARER", "auth_protocol_id": 99, "dst_endpoint": { "ip": "127.0.0.1", "port": 9092, "svc_name": "kafka rpc protocol" }, "is_cleartext": false, "is_mfa": false, "service": { "name": "kafka rpc protocol" }, "src_endpoint": { "ip": "10.0.1.50", "name": "kafka-client", "port": 48210 }, "status_id": 1, // IdP group memberships extracted from the OIDC token "user": { "name": "alice@example.com", "type_id": 1, "groups": [ {"type": "idp_group", "name": "engineering"}, {"type": "idp_group", "name": "analytics"} ] } } ``` Authentication failed This scenario illustrates a common failure where a user entered the wrong credentials. This is an authentication type event. ```json { "category_uid": 3, "class_uid": 3002, "metadata": { "product": { "name": "Redpanda", "uid": "1", "vendor_name": "Redpanda Data, Inc.", "version": "v23.3.0-dev-2457-g76dc896f8c" }, "version": "1.0.0" }, "severity_id": 1, "time": 1700534756350, "type_uid": 300201, "activity_id": 1, "auth_protocol": "SASL-SCRAM", "auth_protocol_id": 99, "dst_endpoint": { "ip": "127.0.0.1", "port": 19092, "svc_name": "kafka rpc protocol" }, "is_cleartext": true, "is_mfa": false, "service": { "name": "kafka rpc protocol" }, "src_endpoint": { "ip": "127.0.0.1", "name": "rpk", "port": 45236 }, "status_id": 2, "status_detail": "SASL authentication failed: security: Invalid credentials", "user": { "name": "admin", "type_id": 1 } } ``` ## [](#kafka-api-events)Kafka API events The Redpanda Kafka API offers a wide array of options for interacting with your Redpanda clusters. Following are examples of messages from common interactions with the API. Create ACL entry This example illustrates an ACL update that also requires a superuser authentication. It lists the edited ACL and the updated permissions. This is a management type event. ```json { "category_uid": 6, "class_uid": 6003, "metadata": { "product": { "name": "Redpanda", "vendor_name": "Redpanda Data, Inc.", "version": "v23.3.0-dev-2457-g76dc896f8c" }, "profiles": [ "cloud" ], "version": "1.0.0" }, "severity_id": 1, "time": 1700533393776, "type_uid": 600303, "activity_id": 3, "actor": { "authorizations": [ { "decision": "authorized", // This shows a superuser level authorization "policy": { "desc": "superuser", "name": "aclAuthorization" } } ], "user": { "name": "admin", "type_id": 2 } }, "api": { // The API operation performed "operation": "create_acls", "service": { "name": "kafka rpc protocol" } }, "cloud": { "provider": "" }, "dst_endpoint": { "ip": "127.0.0.1", "port": 19092, "svc_name": "kafka rpc protocol" }, // List of resources accessed "resources": [ // The created ACL { "name": "create acl", "type": "acl_binding", "data": { "resource_type": "topic", "resource_name": "*", "pattern_type": "literal", "acl_principal": "{type user name user}", "acl_host": "{{any_host}}", "acl_operation": "all", "acl_permission": "allow" } }, // Below indicates that the user had cluster level authorization { "name": "kafka-cluster", "type": "cluster" } ], "src_endpoint": { "ip": "127.0.0.1", "name": "rpk", "port": 50276 }, "status_id": 1, "unmapped": { // Provides a more parsable output of how the // authorization decision was made "authorization_metadata": { "acl_authorization": { "host": "", "op": "", "permission_type": "AUTHORIZED", "principal": "" }, "resource": { "name": "", "pattern": "", "type": "" } } } } ``` Authorization matched on a group ACL This example shows an API Activity (6003) where the authorization decision matched an ALLOW ACL on a `Group:` principal. The `actor.user.groups` field includes the matched group with type `idp_group`, and the `authorization_metadata` shows the group ACL that granted access. See [Group-Based Access Control](https://docs.redpanda.com/cloud-data-platform/security/authorization/gbac/). ```json { "category_uid": 6, "class_uid": 6003, "metadata": { "product": { "name": "Redpanda", "uid": "0", "vendor_name": "Redpanda Data, Inc.", "version": "v26.1.0" }, "version": "1.0.0" }, "severity_id": 1, "time": 1774544504327, "type_uid": 600303, "activity_id": 3, "actor": { "authorizations": [ { "decision": "authorized", "policy": { "desc": "acl: {principal type {group} name {/sales} host {{any_host}} op all perm allow}, resource: type {topic} name {sales-topic} pattern {literal}", "name": "aclAuthorization" } } ], // The matched group appears in the user's groups field "user": { "name": "alice", "type_id": 1, "groups": [ { "type": "idp_group", "name": "/sales" } ] } }, "api": { "operation": "produce", "service": { "name": "kafka rpc protocol" } }, "dst_endpoint": { "ip": "127.0.1.1", "port": 9092, "svc_name": "kafka rpc protocol" }, "resources": [ { "name": "sales-topic", "type": "topic" } ], "src_endpoint": { "ip": "127.0.0.1", "name": "rdkafka", "port": 42728 }, "status_id": 1, "unmapped": { "authorization_metadata": { "acl_authorization": { "host": "{{any_host}}", "op": "all", "permission_type": "allow", "principal": "type {group} name {/sales}" }, "resource": { "name": "sales-topic", "pattern": "literal", "type": "topic" } } } } ``` Metadata request (with counts) This shows a message for a scenario where a user requests a set of metadata using rpk. It provides detailed information on the type of request and the information sent to the user. This is a describe type event. ```json { "category_uid": 6, "class_uid": 6003, // If present, indicates that >1 of the same authz check was performed // within the period of the audit log collecting entries // This provides start and end time (the time period these events were // observed) "count": 2, "end_time": 1700533480725, "metadata": { "product": { "name": "Redpanda", "uid": "0", "vendor_name": "Redpanda Data, Inc.", "version": "v23.3.0-dev-2457-g76dc896f8c" }, "profiles": [ "cloud" ], "version": "1.0.0" }, "severity_id": 1, "start_time": 1700533480724, "time": 1700533480724, "type_uid": 600303, "activity_id": 3, "actor": { "authorizations": [ { "decision": "authorized", // Represents a policy for a non-super user "policy": { "desc": "acl: {principal {type user name user} host {{any_host}} op all perm allow}, resource: type {topic} name {*} pattern {literal}", "name": "aclAuthorization" } } ], "user": { "name": "user", "type_id": 1 } }, "api": { "operation": "metadata", "service": { "name": "kafka rpc protocol" } }, "cloud": { "provider": "" }, "dst_endpoint": { "ip": "127.0.0.1", "port": 19092, "svc_name": "kafka rpc protocol" }, "resources": [ // The topics accessed { "name": "test", "type": "topic" } ], "src_endpoint": { "ip": "127.0.0.1", "name": "rpk", "port": 53602 }, "status_id": 1, "unmapped": { "authorization_metadata": { "acl_authorization": { "host": "{{any_host}}", "op": "all", "permission_type": "allow", "principal": "{type user name user}" }, "resource": { "name": "*", "pattern": "literal", "type": "topic" } } } } ``` --- # Page 386: Cluster Maintenance **URL**: https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance.md --- # Cluster Maintenance > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Cluster Maintenance latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-maintenance/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-maintenance/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/cluster-maintenance/index.adoc description: Learn about cluster maintenance and configuration properties. page-git-created-date: "2025-04-08" page-git-modified-date: "2025-05-07" --- - [Cluster State](cluster-state/) Learn about the current status of a cluster. - [Upgrades and Maintenance](https://docs.redpanda.com/cloud-data-platform/manage/maintenance/) Learn how Redpanda Cloud manages maintenance operations. - [Configure Cluster Properties](config-cluster/) Learn how to configure cluster properties to enable and manage features. - [Audit Logging](https://docs.redpanda.com/cloud-data-platform/manage/audit-logging/) Learn how to use Redpanda's audit logging capabilities. - [About Client Throughput Quotas](about-throughput-quotas/) Understand how Redpanda's user-based and client ID-based throughput quotas work, including entity hierarchy, precedence rules, and quota tracking behavior. - [Manage Throughput](manage-throughput/) Configure broker-wide and client-specific throughput quotas to prevent resource exhaustion and noisy-neighbor issues. - [Fetch Read Coalescing](fetch-read-coalescing/) Reduce redundant read CPU and fetch-response memory under high consumer fan-out by sharing one read result across concurrent fetches of the same data. - [Configure Client Connections](configure-client-connections/) Learn about guidelines for configuring client connections in Redpanda clusters for optimal availability. --- # Page 387: About Client Throughput Quotas **URL**: https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/about-throughput-quotas.md --- # About Client Throughput Quotas > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: About Client Throughput Quotas latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-maintenance/about-throughput-quotas page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-maintenance/about-throughput-quotas.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/cluster-maintenance/about-throughput-quotas.adoc description: Understand how Redpanda's user-based and client ID-based throughput quotas work, including entity hierarchy, precedence rules, and quota tracking behavior. learning-objective-1: Describe the difference between user-based and client ID-based quotas learning-objective-2: Determine which quota type to use for your use case learning-objective-3: Explain quota precedence rules and how Redpanda tracks quota usage page-git-created-date: "2026-03-31" page-git-modified-date: "2026-05-26" --- Redpanda uses throughput quotas to limit the rate of produce and consume requests from clients. Understanding how quotas work helps you prevent individual clients from disproportionately consuming resources and causing performance degradation for other clients (also known as the "noisy-neighbor" problem), and ensure fair resource sharing across users and applications. After reading this page, you will be able to: - Describe the difference between user-based and client ID-based quotas - Determine which quota type to use for your use case - Explain quota precedence rules and how Redpanda tracks quota usage To configure and manage throughput quotas, see [Manage Throughput](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/). ## [](#throughput-control-overview)Throughput control overview Redpanda provides two ways to control throughput: - Broker-wide limits: Configured using cluster properties. For details, see [Broker-wide throughput limits](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/#broker-wide-throughput-limits). - Client throughput quotas: Configured using the Kafka API. Client quotas enable per-user and per-client rate limiting with fine-grained control through entity hierarchy and precedence rules. This page focuses on client quotas. ## [](#supported-quota-types)Supported quota types Redpanda supports three Kafka API-based quota types: | Quota type | Description | | --- | --- | | producer_byte_rate | Limit throughput of produce requests (bytes per second) | | consumer_byte_rate | Limit throughput of fetch requests (bytes per second) | | controller_mutation_rate | Limit rate of topic mutation requests (partitions created or deleted per second) | All quota types can be applied to groups of client connections based on user principals, client IDs, or combinations of both. ## [](#quota-entities)Quota entities Redpanda uses two pieces of identifying information from each client connection to determine which quota applies: - Client ID: An ID that clients self-declare. Quotas can target an exact client ID (`client-id`) or a prefix (`client-id-prefix`). Multiple client connections that share a client ID or ID prefix are grouped into a single quota entity. - User [principal](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#principal): An authenticated identity verified through SASL, mTLS, or OIDC. Connections that share the same user are considered one entity. You can configure quotas that target either entity type, or combine both for fine-grained control. ### [](#client-id-based-quotas)Client ID-based quotas Client ID-based quotas apply to clients identified by their `client-id` field, which is set by the client application. The client ID is typically a configurable property when you create a client with Kafka libraries. When using client ID-based quotas, multiple clients using the same client ID share the same quota tracking. Client ID-based quotas rely on clients honestly reporting their identity and correctly setting the `client-id` property. This makes client ID-based quotas unsuitable for guaranteeing isolation between tenants. Use client ID-based quotas when: - Authentication is not enabled. - Grouping by application or service name is sufficient. - You operate a single-tenant environment where all clients are trusted. - You need simple rate limiting without user-level isolation. ### [](#user-based-quotas)User-based quotas > ❗ **IMPORTANT** > > User-based quotas require [authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/) to be enabled on your cluster. User-based quotas apply to authenticated user principals. Each user has a separate quota, providing a way to limit the impact of individual users on the cluster. User-based quotas rely on Redpanda’s authentication system to verify user identity. The user principal is extracted from SASL credentials, mTLS certificates, or OIDC tokens and cannot be forged by clients. Use user-based quotas when: - You operate a multi-tenant environment, such as SaaS platforms or enterprises with departments. - You require isolation between users or tenants, to avoid noisy-neighbor issues. - You need per-user billing or metering. ### [](#combined-user-and-client-quotas)Combined user and client quotas You can combine user and client identities for fine-grained control over specific (user, client) combinations. Use combined quotas when: - You need fine-grained control, for example: user `alice` using a specific application. - Different rate limits apply to different apps used by the same user. For example, `alice`'s `payment-processor` gets 10 MB/s, but `alice`'s `analytics-consumer` gets 50 MB/s. See [Quota precedence and tracking](#quota-precedence-and-tracking) for examples. ## [](#quota-precedence-and-tracking)Quota precedence and tracking When a request arrives, Redpanda resolves which quota to apply by matching the request’s authenticated user principal and client ID against configured quotas. Redpanda applies the most specific match, using the precedence order in the following table (highest priority first). The precedence level that matches also determines how quota usage is tracked. Redpanda tracks quota usage using a tracker key that determines which connections share the same quota bucket. How connections are grouped into buckets depends on the type of entity the quota targets. To get independent quota tracking per user and client ID combination, configure quotas that include both dimensions, such as `/config/users//clients/` or `/config/users//clients/`. | Level | Match type | Config path | Tracker key | Isolation behavior | | --- | --- | --- | --- | --- | | 1 | Exact user + exact client | /config/users//clients/ | (user, client-id) | Each unique (user, client-id) pair tracked independently | | 2 | Exact user + client prefix | /config/users//client-id-prefix/ | (user, client-id-prefix) | Clients matching the prefix share tracking within that user | | 3 | Exact user + default client | /config/users//clients/ | (user, client-id) | Each unique (user, client-id) pair tracked independently | | 4 | Exact user only | /config/users/ | user | All clients for that user share a single tracking bucket | | 5 | Default user + exact client | /config/users//clients/ | (user, client-id) | Each unique (user, client-id) pair tracked independently | | 6 | Default user + client prefix | /config/users//client-id-prefix/ | (user, client-id-prefix) | Clients matching the prefix share tracking within each user | | 7 | Default user + default client | /config/users//clients/ | (user, client-id) | Each unique (user, client-id) pair tracked independently | | 8 | Default user only | /config/users/ | user | All clients for each user share a single tracking bucket (per user) | | 9 | Exact client only | /config/clients/ | client-id | All users with that client ID share a single tracking bucket | | 10 | Client prefix only | /config/client-id-prefix/ | client-id-prefix | All clients matching the prefix share a single bucket across all users | | 11 | Default client only | /config/clients/ | client-id | Each unique client ID tracked independently | | 12 | No quota configured | N/A | N/A | No tracking / unlimited throughput | > ❗ **IMPORTANT** > > The `` entity matches any user or client that doesn’t have a more specific quota configured. This is different from an empty/unauthenticated user (`user=""`), or undeclared client ID (`client-id=""`), which are treated as specific entities. ### [](#unauthenticated-connections)Unauthenticated connections Unauthenticated connections have an empty user principal (`user=""`) and are not treated as `user=`. Unauthenticated connections: - Fall back to client-only quotas. - Have unlimited throughput only if no client-only quota matches. ### [](#example-precedence-resolution)Example: Precedence resolution Given these configured quotas: ```bash rpk cluster quotas alter --add consumer_byte_rate=5000000 --name user=alice --name client-id=app-1 rpk cluster quotas alter --add consumer_byte_rate=10000000 --name user=alice rpk cluster quotas alter --add consumer_byte_rate=20000000 --name client-id=app-1 ``` | User + Client ID | Precedence match | | --- | --- | | user=alice, client-id=app-1 | Level 1: Exact user + exact client | | user=alice, client-id=app-2 | Level 4: Exact user only | | user=bob, client-id=app-1 | Level 9: Exact client only | | user=bob, client-id=app-2 | Level 12: No quota configured | When no quota matches (level 12), the connection is not throttled. ### [](#example-user-only-quota)Example: User-only quota If you configure a 10 MB/s produce quota for user `alice`: ```bash rpk cluster quotas alter --add producer_byte_rate=10000000 --name user=alice ``` Then `alice` connecting with client ID `app-1` and `alice` connecting with client ID `app-2` share the same 10 MB/s produce limit. To give each of `alice`'s clients an independent 10 MB/s limit, configure: ```bash rpk cluster quotas alter --add producer_byte_rate=10000000 --name user=alice --default client-id ``` ### [](#example-user-default-quota)Example: User default quota If you configure a default 10 MB/s produce quota for all users: ```bash rpk cluster quotas alter --add producer_byte_rate=10000000 --default user ``` This quota applies to all users who don’t have a more specific quota configured. Each user is tracked independently: `alice` gets her own 10 MB/s bucket, `bob` gets his own 10 MB/s bucket, and so on. Within each user, all client ID values share that user’s bucket. `alice` connecting with client ID `app-1` and `alice` connecting with client ID `app-2` share the same 10 MB/s produce limit, while `bob`'s connections have a separate 10 MB/s limit. ## [](#throttling-enforcement)Throughput throttling enforcement > 📝 **NOTE** > > As of v24.2, Redpanda enforces all throughput limits per broker, including client throughput. Redpanda enforces throughput limits by applying backpressure to clients. When a connection exceeds its throughput limit, Redpanda throttles the connection to bring the rate back within the allowed level: 1. Redpanda adds a `throttle_time_ms` field to responses, indicating how long the client should wait. 2. If the client doesn’t honor the throttle time, Redpanda inserts delays on the connection’s next read operation. In Redpanda Cloud, the throttling delay is set to 30 seconds. ## [](#default-behavior)Default behavior Quotas are opt-in restrictions and not enforced by default. When no quotas are configured, clients have unlimited throughput. ## [](#next-steps)Next steps - [Configure throughput quotas](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/) - [Enable authentication for user-based quotas](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/) --- # Page 388: Cluster State **URL**: https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/cluster-state.md --- # Cluster State > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Cluster State latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-maintenance/cluster-state page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-maintenance/cluster-state.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/cluster-maintenance/cluster-state.adoc description: Learn about the current status of a cluster. page-git-created-date: "2025-07-23" page-git-modified-date: "2025-07-24" --- The cluster state shows the current status of a cluster. Redpanda Cloud updates the state automatically, allowing you to monitor a cluster’s health and availability. ## Serverless | State | Description | | --- | --- | | Creating | Cluster is in the process of having its control plane state created. | | Placing | Cluster is in the process of being placed on a cell with sufficient resources in the data plane. | | Ready | Cluster is running and accepting external requests. | | Deleting | Cluster is in the process of having its control plane state removed. Resources dedicated to the cluster in the data plane are released. | | Failed | Cluster is unable to enter the Ready state from either the Creating or Placing states.Try re-creating the cluster. | | Suspended | Cluster is running but blocks all external requests.This can happen when credits run out. Enter a credit card to return to the Ready state. | ## BYOC/Dedicated | State | Description | | --- | --- | | Creating agent | Cluster is in the process of having its control plane state created, and the Redpanda Cloud agent is being deployed. | | Creating | Cluster is in the process of having its control plane state created. | | Ready | Cluster is running and accepting external requests. | | Deleting | Cluster is in the process of having its control plane state removed. Resources dedicated to the cluster in the data plane are released. | | Deleting agent | Cluster is in the process of having its control plane state and Redpanda Cloud agent removed. | | Upgrading | Cluster is undergoing a rolling upgrade or a scaling operation. | | Failed | Cluster is unable to enter the Ready state from either the Creating or the Creating agent states.Try re-creating the cluster. | | Suspended | Cluster is running but blocks all external requests.This can happen when credits run out. Enter a credit card to return to the Ready state. | --- # Page 389: Configure Cluster Properties **URL**: https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster.md --- # Configure Cluster Properties > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Cluster Properties latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-maintenance/config-cluster page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-maintenance/config-cluster.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/cluster-maintenance/config-cluster.adoc description: Learn how to configure cluster properties to enable and manage features. page-git-created-date: "2025-04-08" page-git-modified-date: "2026-07-02" --- Cluster configuration properties are set to their default values and are automatically replicated across all brokers. You can use cluster properties to enable and manage features such as [Iceberg topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/), [data transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/), and [audit logging](https://docs.redpanda.com/cloud-data-platform/manage/audit-logging/). For a complete list of the cluster properties available in Redpanda Cloud, see [Cluster Configuration Properties](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/) and [Object Storage Properties](https://docs.redpanda.com/cloud-data-platform/reference/properties/object-storage-properties/). > 📝 **NOTE** > > Some properties are read-only and cannot be changed. For example, `cluster_id` is a read-only property that is automatically set when the cluster is created. ## [](#prerequisites)Prerequisites - **`rpk` version 25.1.2+**: To check your current version, see [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). - **Redpanda version 25.1.2+**: You can find the version on your cluster’s Overview page in the Redpanda Cloud UI. To verify that you’re logged into the Redpanda control plane and have the correct `rpk` profile configured for your target cluster, run `rpk cloud login` and select your cluster. ## [](#limitations)Limitations Cluster properties are supported on BYOC and Dedicated clusters running on AWS and GCP. - They are not available on BYOC and Dedicated clusters running on Azure. - They are not available on Serverless clusters. Only the cluster properties listed in [Cluster Configuration Properties](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/) and [Object Storage Properties](https://docs.redpanda.com/cloud-data-platform/reference/properties/object-storage-properties/) are configurable in Redpanda Cloud. Setting an unsupported property through the Cloud API or the [Redpanda Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/) returns a `REASON_INVALID_INPUT` error indicating that the property is not allowed. To control the maximum message size, see [Configure the maximum message size](#max-message-size). ## [](#set-cluster-configuration-properties)Set cluster configuration properties You can set cluster configuration properties using the `rpk` command-line tool or the Cloud API. ### rpk Use `rpk cluster config` to set cluster properties. For example, to enable audit logging, set [`audit_enabled`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#audit_enabled) to `true`: ```bash rpk cluster config set audit_enabled true ``` To set a cluster property with a secret, you must use the following notation: ```bash rpk cluster config set iceberg_rest_catalog_client_secret '${secrets.}' ``` > 📝 **NOTE** > > Some properties require a rolling restart, and it can take several minutes for the update to complete. The `rpk cluster config set` command returns the operation ID. ### Cloud API Use the Cloud API to set cluster properties: - Create a cluster by making a [`POST /v1/clusters`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_createcluster) request. Edit `cluster_configuration` in the request body with a key-value pair for `custom_properties`. - Update a cluster by making a [`PATCH /v1/clusters/{cluster.id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) request, passing the cluster ID as a parameter. Include the properties to update in the request body. For example, to set [`audit_enabled`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#audit_enabled) to `true`: ```bash # Store your cluster ID in a variable. export RP_CLUSTER_ID= # Retrieve a Redpanda Cloud access token. export RP_CLOUD_TOKEN=`curl -X POST "https://auth.prd.cloud.redpanda.com/oauth/token" \ -H "content-type: application/x-www-form-urlencoded" \ -d "grant_type=client_credentials" \ -d "client_id=" \ -d "client_secret="` # Update your cluster configuration to enable audit logging. curl -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" -X PATCH \ "https://api.cloud.redpanda.com/v1/clusters/${RP_CLUSTER_ID}" \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{"cluster_configuration":{"custom_properties": {"audit_enabled":true}}}' ``` The [`PATCH /clusters/{cluster.id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) request returns the ID of a long-running operation. You can check the status of the operation by polling the [`GET /operations/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-operationservice_getoperation) endpoint. To set a cluster property with a secret, you must use the following notation with the secret name: ```bash curl -H "Authorization: Bearer " -X PATCH \ "https://api.cloud.redpanda.com/v1/clusters/" \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{"cluster_configuration": { "custom_properties": { "iceberg_rest_catalog_client_secret": "${secrets.}" } } }' ``` > 📝 **NOTE** > > Some properties require a rolling restart for the update to take effect. This triggers a [long-running operation](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/#lro) that can take several minutes to complete. ## [](#set-default-topic-properties)Set default topic properties You can set cluster-wide defaults that apply to new topics: | Property | Description | | --- | --- | | default_topic_partitions | Default number of partitions for new topics. | | log_retention_ms | Default length of time to retain topic data before it becomes eligible for deletion. This sets the default topic retention period, which applies when a topic doesn’t set its own retention.ms. | | retention_bytes | Default maximum size per partition before the oldest data becomes eligible for deletion. Applies when a topic doesn’t set its own retention.bytes. | Set these properties using `rpk` or the Cloud API. For example, to set the default topic retention period to 24 hours: ### rpk ```bash rpk cluster config set log_retention_ms 86400000 ``` ### Cloud API ```bash # Store your cluster ID in a variable. export RP_CLUSTER_ID= # Retrieve a Redpanda Cloud access token. export RP_CLOUD_TOKEN=`curl -X POST "https://auth.prd.cloud.redpanda.com/oauth/token" \ -H "content-type: application/x-www-form-urlencoded" \ -d "grant_type=client_credentials" \ -d "client_id=" \ -d "client_secret="` # Set the default topic retention period to 24 hours. curl -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" -X PATCH \ "https://api.cloud.redpanda.com/v1/clusters/${RP_CLUSTER_ID}" \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{"cluster_configuration":{"custom_properties": {"log_retention_ms":86400000}}}' ``` ## [](#max-message-size)Configure the maximum message size To control the maximum message size in Redpanda Cloud, set the `max.message.bytes` property on each topic. The default is 20 MiB for BYOC and Dedicated clusters and 8 MiB for Serverless clusters. You can increase it up to 32 MiB for BYOC and Dedicated clusters and 20 MiB for Serverless clusters. To set this property, see [Topics Overview](https://docs.redpanda.com/cloud-data-platform/develop/topics/create-topic/). The cluster property `kafka_batch_max_bytes` cannot be set through the Cloud API or the [Redpanda Terraform provider](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/). Attempting to set it returns `REASON_INVALID_INPUT` with the message `The properties [kafka_batch_max_bytes] are not allowed`. ## [](#view-cluster-property-values)View cluster property values You can see the value of a cluster configuration property using `rpk` or the Cloud API. ### rpk Use `rpk cluster config get` to view the current cluster property value. For example, to view the current value of [`audit_enabled`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#audit_enabled), run: ```bash rpk cluster config get audit_enabled ``` ### Cloud API Use the Cloud API to get the current configuration property values for a cluster. Make a [`GET /clusters/{cluster.id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_getcluster) request, passing the cluster ID as a parameter. The response body contains the current `computed_properties` values. For example, to get the current value of [`audit_enabled`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#audit_enabled): ```bash # Store your cluster ID in a variable. export RP_CLUSTER_ID= # Retrieve a Redpanda Cloud access token. export RP_CLOUD_TOKEN=`curl -X POST "https://auth.prd.cloud.redpanda.com/oauth/token" \ -H "content-type: application/x-www-form-urlencoded" \ -d "grant_type=client_credentials" \ -d "client_id=" \ -d "client_secret="` # Get your cluster configuration property values. curl -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" -X GET \ "https://api.cloud.redpanda.com/v1/clusters/${RP_CLUSTER_ID}" \ -H 'accept: application/json'\ -H 'content-type: application/json' \ ``` ## [](#suggested-reading)Suggested reading - [Introduction to rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/intro-to-rpk/) - [Redpanda Cloud API Overview](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview) - [Redpanda Cloud API Quickstart](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-quickstart) --- # Page 390: Configure Client Connections **URL**: https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/configure-client-connections.md --- # Configure Client Connections > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Client Connections latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-maintenance/configure-client-connections page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-maintenance/configure-client-connections.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/cluster-maintenance/configure-client-connections.adoc description: Learn about guidelines for configuring client connections in Redpanda clusters for optimal availability. page-git-created-date: "2025-11-19" page-git-modified-date: "2026-05-26" --- Optimize the availability of your clusters by configuring and tuning properties. > 💡 **TIP** > > Before you configure connection limits or reconnection settings, start by gathering detailed data about your client connections. > > - Use the [`redpanda_rpc_active_connections` metric](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_rpc_active_connections) to view current Kafka client connections. > > - For clusters on v25.3 and later, use [`rpk cluster connections list`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-connections-list/) or the `GET /v1/monitoring/kafka/connections` endpoint in the Data Plane API to identify: > > - Which clients and applications are connected > > - Long-lived connections and long-running requests > > - Connections with no activity > > - Whether any clients are causing excessive load > > > By reviewing connection details, you can make informed decisions about tuning connection limits and troubleshooting issues. > > > See also: [Data Plane API reference](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-monitoringservice_listkafkaconnections), [Monitor Redpanda Cloud](https://docs.redpanda.com/cloud-data-platform/manage/monitor-cloud/#throughput) ## [](#limit-client-connections)Limit client connections To mitigate the risk of a client creating too many connections and using too many system resources, you can configure a Redpanda cluster to impose limits on the number of client connections that can be created. The following Redpanda cluster properties limit the number of connections: - [`kafka_connections_max_per_ip`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#kafka_connections_max_per_ip): Similar to Kafka’s `max.connections.per.ip`, this sets the maximum number of connections accepted per IP address by a broker. - [`kafka_connections_max_overrides`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#kafka_connections_max_overrides): A list of IP addresses for which `kafka_connections_max_per_ip` is overridden and doesn’t apply. > 📝 **NOTE** > > - These connection limit properties are disabled by default. You must manually enable them. > > - The total number of connections is not equal to the number of clients, because a client can open multiple connections. As a conservative estimate, for a cluster with N brokers, plan for N + 2 connections per client. ### [](#configure-connection-count-limit-by-client-ip)Configure connection count limit by client IP Configure the `kafka_connections_max_per_ip` property to limit the number of connections from each client IP address. > ❗ **IMPORTANT** > > Per-IP connection controls require Redpanda to see individual client IPs. If clients connect through private link endpoints, NAT gateways, or other shared-IP egress, the per-IP limit applies to the shared IP, affecting all clients behind it and preventing isolation of a single offending client. Similarly, multiple clients running on the same host will share the same IP address, and the limit applies collectively to all those clients. See also: [Configure Cluster Properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/) #### [](#configure-the-limit)Configure the limit To configure `kafka_connections_max_per_ip` safely without disrupting legitimate clients, follow these steps: 1. Set up your monitoring stack for your cluster. See [Monitor Redpanda Cloud](https://docs.redpanda.com/cloud-data-platform/manage/monitor-cloud/). 2. Monitor current connection patterns using the `redpanda_rpc_active_connections` metric with the `redpanda_server="kafka"` filter: ```none redpanda_rpc_active_connections{redpanda_id="CLOUD_CLUSTER_ID", redpanda_server="kafka"} ``` 3. Analyze the connection data to identify the normal range of connections for each broker during typical traffic cycles. For example, in the following Grafana screenshot, the normal range is around 200-300 connections: ![Range of active connections over time](https://docs.redpanda.com/cloud-data-platform/shared/_images/monitor_connections.png) 4. Set the `kafka_connections_max_per_ip` value based on your analysis. Use the upper bound of normal connections observed, or use a lower value if you know how many connections per client IP are being opened. 5. Continue monitoring the connection metrics after applying the limit to ensure that legitimate clients are not affected and that the problematic client is properly controlled. > 📝 **NOTE** > > If you find a high load of unexpected connections from multiple IP addresses, `kafka_connections_max_per_ip` alone may be insufficient. If offending IPs outnumber legitimate client IPs, you may need to set `kafka_connections_max_per_ip` so low that it affects legitimate clients. If this is the case, use `kafka_connections_max_overrides` to exempt known legitimate client IPs from the connection limit. #### [](#limitations)Limitations - Decreasing the limit does not terminate any currently open Kafka API connections. - This limit does not apply to Kafka HTTP Proxy connections. - Clients behind NAT gateways or private links share the same IP address as seen by Redpanda brokers. - The limit may negatively affect tail latencies across all client connections. - All clients behind the shared IP are collectively subject to the single `kafka_connections_max_per_ip` limit. - Connection rejections occur randomly among clients when the limit is reached. For example, suppose `kafka_connections_max_per_ip` is set to 100, but clients behind a NAT gateway collectively need 150 connections. When the limit is reached, clients can make only some of the connections while others get rejected, leaving the client in a not-working state. - Redpanda may modify this property during internal operations. - Availability incidents caused by misconfiguring this feature are excluded from the Redpanda Cloud SLA. ## [](#configure-client-reconnections)Configure client reconnections You can configure the Kafka client backoff and retry properties to change the default behavior of the clients to suit your failure requirements. Set the following Kafka client properties on your application’s producer or consumer to manage client reconnections: - `reconnect.backoff.ms`: Amount of time to wait before attempting to reconnect to the broker. The default is 50 milliseconds. - `reconnect.backoff.max.ms`: Maximum amount of time in milliseconds to wait when reconnecting to a broker. The backoff increases exponentially for each consecutive connection failure, up to this maximum. The default is 1000 milliseconds (1 second). Additionally, you can use Kafka properties to control message retry behavior. Delivery fails when either the delivery timeout or the number of retries is met. - `delivery.timeout.ms`: Amount of time for message delivery, so messages are not retried forever. The default is 120000 milliseconds (2 minutes). - `retries`: Number of times a producer can retry sending a message before marking it as failed. The default value is 2147483647 for Kafka >= 2.1, or 0 for Kafka <= 2.0. - `retry.backoff.ms`: Amount of time to wait before attempting to retry a failed request to a given topic partition. The default is 100 milliseconds. ## [](#see-also)See also - [Configure Producers](https://docs.redpanda.com/cloud-data-platform/develop/produce-data/configure-producers/) - [Manage Throughput](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput/) --- # Page 391: Fetch Read Coalescing **URL**: https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/fetch-read-coalescing.md --- # Fetch Read Coalescing > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Fetch Read Coalescing latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-maintenance/fetch-read-coalescing page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-maintenance/fetch-read-coalescing.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/cluster-maintenance/fetch-read-coalescing.adoc description: Reduce redundant read CPU and fetch-response memory under high consumer fan-out by sharing one read result across concurrent fetches of the same data. page-git-created-date: "2026-08-15" page-git-modified-date: "2026-08-15" --- When many consumers fetch the same partition at the same offset concurrently (high read fan-out), the broker does redundant work: each fetch performs its own log read, its own serialization, and allocates its own copy of the response bytes. With a fan-out of N consumers, that is roughly N times the read CPU and N times the fetch-response memory for byte-identical output. Fetch read coalescing removes that redundancy. The broker reads and serializes each unique read once, and shares the single result with every concurrent (and eligible back-to-back) consumer of the same data: one read, one serialization, one buffer, fanned out to all requesters. When all N consumers fetch the same partition at the same offset with the same fetch settings, read CPU and fetch-response memory drop from roughly N times to one. Fetch read coalescing is available on BYOC and Dedicated clusters. Fetch read coalescing is disabled by default, and the cluster property that controls it is managed by Redpanda. See [Enable fetch read coalescing](#enable-fetch-read-coalescing). It benefits workloads where many consumers tail the same partitions with the same fetch settings, for example fan-out delivery of one stream to many downstream applications. Request it only for high fan-out workloads: at low fan-out (for example, one or two consumers per partition), the de-duplication logic saves little to nothing and its overhead can increase reactor utilization. ## [](#how-fetch-read-coalescing-works)How fetch read coalescing works Redpanda coalesces fetch reads only when they request the same data in the same way: the same partition and offset, with the same isolation level and fetch size (`max_bytes`). When matching reads arrive concurrently, or back-to-back while a previous result is still in use, the broker performs the read once and shares the single result with every requester. Reads that differ in any of these properties run independently. If a shared read fails, all coalesced fetches receive the error. Coalesced consumers share one response buffer instead of each holding its own copy, which is where the memory savings come from. A completed result is retained only while at least one fetch still references it, so the coalescer never pins memory on its own. Coalescing is scoped per shard: it collapses only the fan-out that lands on the same shard. It composes with [follower fetching](https://docs.redpanda.com/cloud-data-platform/develop/consume-data/follower-fetching/) rather than replacing it: follower fetching spreads consumers across replicas to distribute the read load, and coalescing removes the redundancy within each shard. ## [](#limitations)Limitations Before requesting fetch read coalescing, familiarize yourself with the following limitations: - **Only identical reads coalesce.** Consumers reading the same data but with a different `max_bytes`, a different offset, or a different isolation level do not share a read. The benefit scales with how closely the concurrent fetches match: the ideal case is many consumers tailing the same partition at the same offset with the same fetch sizing. - **Obligatory and strict reads are grouped apart.** Redpanda distinguishes reads that must return at least one batch regardless of the byte limit (obligatory) from reads that strictly honor `max_bytes` (strict), and neither serves the other. A single partition and offset fetched both ways at the same time incurs one duplicate read by design. - **Retention is best-effort.** Completed results are only reusable by later readers while a fetch still references them. The coalescer never keeps a result alive on its own. The reliable savings are on genuinely concurrent readers of the same data, and back-to-back reuse is opportunistic. - **Per-shard scope.** Coalescing collapses only the fan-out that lands on the same shard. Where consumers are spread across replicas and shards, for example with follower fetching, each shard coalesces the fan-out local to it. - **De-duplication adds per-read overhead.** The de-duplication logic runs on every fetch read, whether or not reads coalesce. At low fan-out (for example, one or two consumers per partition), this overhead can increase reactor utilization while saving little to nothing, so only enable coalescing for high fan-out workloads. ## [](#enable-fetch-read-coalescing)Enable fetch read coalescing The `kafka_fetch_read_coalescing_enabled` cluster property that controls coalescing is not user-configurable in Redpanda Cloud. To enable coalescing on a BYOC or Dedicated cluster, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). Enabling or disabling coalescing is a runtime change: no cluster restart is required, and the coalescing cache (about 0.5 MB per shard) is allocated only while the feature is enabled. ## [](#suggested-reading)Suggested reading - [Follower Fetching](https://docs.redpanda.com/cloud-data-platform/develop/consume-data/follower-fetching/) - [Monitor Redpanda Cloud](https://docs.redpanda.com/cloud-data-platform/manage/monitor-cloud/) --- # Page 392: Manage Throughput **URL**: https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/manage-throughput.md --- # Manage Throughput > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Manage Throughput latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: cluster-maintenance/manage-throughput page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: cluster-maintenance/manage-throughput.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/cluster-maintenance/manage-throughput.adoc description: Configure broker-wide and client-specific throughput quotas to prevent resource exhaustion and noisy-neighbor issues. learning-objective-1: Set user-based throughput quotas learning-objective-2: Set client ID-based quotas learning-objective-3: Monitor quota usage and throttling behavior page-git-created-date: "2025-08-19" page-git-modified-date: "2026-05-26" --- Redpanda throttles throughput on ingress and egress independently, and you can configure limits at the broker and client levels. This prevents clients from causing unbounded network and disk usage on brokers. You can configure limits at two levels: - Broker limits: These apply to all clients connected to the broker and restrict total traffic on the broker. See [Broker-wide throughput limits](#broker-wide-throughput-limits). - Client limits: These apply to authenticated users or clients defined by their client ID. You can manage client quotas with [`rpk cluster quotas`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-quotas/), with the Redpanda Cloud UI, with the [Redpanda Cloud Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-quotaservice_listquotas), or with the Kafka API. When no quotas apply, the client has unlimited throughput. > 📝 **NOTE** > > Throughput throttling is supported for BYOC and Dedicated clusters only. After reading this page, you will be able to: - Set user-based throughput quotas - Set client ID-based quotas - Monitor quota usage and throttling behavior ## [](#view-connected-client-details)View connected client details Before configuring throughput quotas, check the [current produce and consume throughput](https://docs.redpanda.com/cloud-data-platform/manage/monitor-cloud/#throughput) of a client. Use the [`rpk cluster connections list`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-connections-list/) command or the [`GET /v1/monitoring/kafka/connections`](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-monitoringservice_listkafkaconnections) Data Plane API endpoint to view detailed information about active Kafka client connections. For example, to view a cluster’s connected clients in order of highest current produce throughput, run: ### rpk ```bash rpk cluster connections list --order-by="recent_request_statistics.produce_bytes desc" ``` ```bash UID STATE USER CLIENT-ID IP:PORT NODE SHARD OPEN-TIME IDLE PROD-TPUT/SEC FETCH-TPUT/SEC REQS/MIN b20601a3-624c-4a8c-ab88-717643f01d56 OPEN UNAUTHENTICATED perf-producer-client 127.0.0.1:55012 0 0 9s 0s 78.9MB 0B 292 36338ca5-86b7-4478-ad23-32d49cfaef61 OPEN UNAUTHENTICATED rpk 127.0.0.1:49722 0 0 13s 13.694243104s 0B 0B 1 7e277ef6-0176-4007-b100-6581bfde570f OPEN UNAUTHENTICATED rpk 127.0.0.1:49736 0 0 13s 10.093957335s 0B 0B 2 567d9918-d3dc-4c74-ab5d-85f70cd3ee35 OPEN UNAUTHENTICATED rpk 127.0.0.1:49748 0 0 13s 0.591413542s 0B 0B 5 08616f21-08f9-46e7-8f06-964bd8240d9b OPEN UNAUTHENTICATED rpk 127.0.0.1:49764 0 0 13s 10.094602845s 0B 0B 2 e4d5b57e-5c76-4975-ada8-17a88d68a62d OPEN UNAUTHENTICATED rpk 127.0.0.1:54992 0 0 10s 0.302090085s 0B 14.5MB 27 b41584f3-2662-4185-a4b8-0d8510f5c780 OPEN UNAUTHENTICATED perf-producer-client 127.0.0.1:55002 0 0 8s 7.743592270s 0B 0B 1 62fde947-411d-4ea8-9461-3becc2631b46 CLOSED UNAUTHENTICATED rpk 127.0.0.1:48578 0 0 26s 0.000737836s 0B 0B 1 95387e2e-2ec4-4040-aa5e-4257a3efa1a2 CLOSED UNAUTHENTICATED rpk 127.0.0.1:48564 0 0 26s 0.208180826s 0B 0B 1 ``` ### Data Plane API ```bash curl \ --request GET 'https:///v1/monitoring/kafka/connections' \ --header "Authorization: Bearer $ACCESS_TOKEN" \ --data '{ "filter": "", "order_by": "recent_request_statistics.produce_bytes desc" }' ``` Show example API response ```json { "connections": [ { "node_id": 0, "shard_id": 0, "uid": "b20601a3-624c-4a8c-ab88-717643f01d56", "state": "KAFKA_CONNECTION_STATE_OPEN", "open_time": "2025-10-15T14:15:15.755065000Z", "close_time": "1970-01-01T00:00:00.000000000Z", "authentication_info": { "state": "AUTHENTICATION_STATE_UNAUTHENTICATED", "mechanism": "AUTHENTICATION_MECHANISM_UNSPECIFIED", "user_principal": "" }, "listener_name": "", "tls_info": { "enabled": false }, "source": { "ip_address": "127.0.0.1", "port": 55012 }, "client_id": "perf-producer-client", "client_software_name": "apache-kafka-java", "client_software_version": "3.9.0", "transactional_id": "my-tx-id", "group_id": "", "group_instance_id": "", "group_member_id": "", "api_versions": { "18": 4, "22": 3, "3": 12, "24": 3, "0": 7 }, "idle_duration": "0s", "in_flight_requests": { "sampled_in_flight_requests": [ { "api_key": 0, "in_flight_duration": "0.000406892s" } ], "has_more_requests": false }, "total_request_statistics": { "produce_bytes": "78927173", "fetch_bytes": "0", "request_count": "4853", "produce_batch_count": "4849" }, "recent_request_statistics": { "produce_bytes": "78927173", "fetch_bytes": "0", "request_count": "4853", "produce_batch_count": "4849" } }, ... ], "total_size": "9" } ``` To view connections for a specific client, you can use a filter expression: ### rpk ```bash rpk cluster connections list --client-id="perf-producer-client" ``` ```bash UID STATE USER CLIENT-ID IP:PORT NODE SHARD OPEN-TIME IDLE PROD-TPUT/SEC FETCH-TPUT/SEC REQS/MIN b41584f3-2662-4185-a4b8-0d8510f5c780 OPEN UNAUTHENTICATED perf-producer-client 127.0.0.1:55002 0 0 8s 7.743592270s 0B 0B 1 b20601a3-624c-4a8c-ab88-717643f01d56 OPEN UNAUTHENTICATED perf-producer-client 127.0.0.1:55012 0 0 9s 0s 78.9MB 0B 292 ``` The `USER` field in the connection list shows the authenticated principal. Unauthenticated connections show `UNAUTHENTICATED`, which corresponds to an empty user principal (`user=""`) in quota configurations, not `user=`. ### Data Plane API ```bash curl \ --request GET 'https:///v1/monitoring/kafka/connections' \ --header "Authorization: Bearer $ACCESS_TOKEN" \ --data '{ "filter": "client_id = \"perf-producer-client\"" }' ``` Show example API response ```json { "connections": [ { "node_id": 0, "shard_id": 0, "uid": "b41584f3-2662-4185-a4b8-0d8510f5c780", "state": "KAFKA_CONNECTION_STATE_OPEN", "open_time": "2025-10-15T14:15:15.219538000Z", "close_time": "1970-01-01T00:00:00.000000000Z", "authentication_info": { "state": "AUTHENTICATION_STATE_UNAUTHENTICATED", "mechanism": "AUTHENTICATION_MECHANISM_UNSPECIFIED", "user_principal": "" }, "listener_name": "", "tls_info": { "enabled": false }, "source": { "ip_address": "127.0.0.1", "port": 55002 }, "client_id": "perf-producer-client", "client_software_name": "apache-kafka-java", "client_software_version": "3.9.0", "transactional_id": "", "group_id": "", "group_instance_id": "", "group_member_id": "", "api_versions": { "18": 4, "3": 12, "10": 4 }, "idle_duration": "7.743592270s", "in_flight_requests": { "sampled_in_flight_requests": [], "has_more_requests": false }, "total_request_statistics": { "produce_bytes": "0", "fetch_bytes": "0", "request_count": "3", "produce_batch_count": "0" }, "recent_request_statistics": { "produce_bytes": "0", "fetch_bytes": "0", "request_count": "3", "produce_batch_count": "0" } }, ... ], "total_size": "2" } ``` The user principal field in the connection list shows the authenticated principal. Unauthenticated connections show `AUTHENTICATION_STATE_UNAUTHENTICATED`, which corresponds to an empty user principal (`user=""`) in quota configurations, not `user=`. To view connections for a specific authenticated user: ```bash rpk cluster connections list --user alice ``` This shows all connections from user `alice`, useful for monitoring clients that are subject to user-based quotas. ## [](#broker-wide-throughput-limits)Broker-wide throughput limits Broker-wide throughput limits account for all Kafka API traffic going into or out of the broker, as data is produced to or consumed from a topic. The limit values represent the allowed rate of data in bytes per second passing through in each direction. Redpanda also provides administrators the ability to exclude clients from throughput throttling and to fine-tune which Kafka request types are subject to throttling limits. ## [](#client-throughput-limits)Client throughput limits Redpanda provides configurable throughput quotas for individual clients or authenticated users. Quotas are managed through the Kafka-compatible AlterClientQuotas and DescribeClientQuotas APIs, accessible with `rpk`, Redpanda Console, or Kafka client libraries. Redpanda supports two types of client throughput quotas: - Client ID-based quotas: Limit throughput based on the self-declared `client-id` field. - User-based quotas: Limit throughput based on authenticated user [principal](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#principal). Requires [authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/). You can also combine both types for fine-grained control (for example, limiting a specific user when using a specific client application). For conceptual information about quota types, entity hierarchy, precedence rules, and how Redpanda tracks and enforces quotas through throttling, see [About Client Throughput Quotas](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/about-throughput-quotas/). ### [](#set-user-based-quotas)Set user-based quotas > ❗ **IMPORTANT** > > User-based quotas require authentication to be enabled. To set up authentication, see [Authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/). #### [](#quota-for-a-specific-user)Quota for a specific user To limit throughput for a specific authenticated user across all clients: ```bash rpk cluster quotas alter --add producer_byte_rate=2000000 --name user=alice ``` This limits user `alice` to 2 MB/s for produce requests regardless of the client ID used. To view quotas for a user: ```bash rpk cluster quotas describe --name user=alice ``` Expected output: ```bash user=alice producer_byte_rate=2000000 ``` #### [](#default-quota-for-all-users)Default quota for all users To set a fallback quota for any user without a more specific quota: ```bash rpk cluster quotas alter --add consumer_byte_rate=5000000 --default user ``` This applies a 5 MB/s fetch quota to all authenticated users who don’t have a more specific quota configured. ### [](#remove-a-user-quota)Remove a user quota To remove a quota for a specific user: ```bash rpk cluster quotas alter --delete consumer_byte_rate --name user=alice ``` To remove all quotas for a user: ```bash rpk cluster quotas delete --name user=alice ``` ### [](#set-client-id-based-quotas)Set client ID-based quotas Client ID-based quotas apply to all users using a specific client ID. These quotas do not require authentication. Because the client ID is self-declared, client ID-based quotas are not suitable for guaranteeing isolation between tenants. For multi-tenant environments, Redpanda recommends user-based quotas for per-tenant isolation. #### [](#individual-client-id-throughput-limit)Individual client ID throughput limit > 📝 **NOTE** > > The following sections show how to manage throughput with `rpk`. You can also manage throughput with the [Redpanda Cloud Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-quotaservice_listquotas). To view current throughput quotas set through the Kafka API, run [`rpk cluster quotas describe`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-quotas-describe/). For example, to see the quotas for client ID `consumer-1`: ```bash rpk cluster quotas describe --name client-id=consumer-1 ``` ```bash client-id=consumer-1 producer_byte_rate=140000 ``` To set a throughput quota for a single client, use the [`rpk cluster quotas alter`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-quotas-alter/) command. ```bash rpk cluster quotas alter --add consumer_byte_rate=200000 --name client-id=consumer-1 ``` ```bash ENTITY STATUS client-id=consumer-1 OK ``` #### [](#group-of-clients-throughput-limit)Group of clients throughput limit Alternatively, you can view or configure throughput quotas for a group of clients based on a match on client ID prefix. The following example sets the `consumer_byte_rate` quota to client IDs prefixed with `consumer-`: ```bash rpk cluster quotas alter --add consumer_byte_rate=200000 --name client-id-prefix=consumer- ``` > 📝 **NOTE** > > A `client-id-prefix` quota group is not related to Kafka consumer groups. The client ID is an application-defined identifier sent with every request. Client libraries typically default to their own name (such as `kgo`, `rdkafka`, `sarama`, or `perf-producer-client`), but applications can set it using the [`client.id`](https://kafka.apache.org/documentation/#consumerconfigs_client.id) configuration property. This makes prefix-based quotas useful for grouping related applications (for example, `inventory-service-` to match `inventory-service-1`, `inventory-service-2`, etc.). #### [](#default-client-throughput-limit)Default client throughput limit You can apply default throughput limits to clients. Redpanda applies the default limits if no quotas are configured for a specific client ID or prefix. To specify a produce quota of 1 GB/s through the Kafka API (applies across all produce requests to a single broker), run: ```bash rpk cluster quotas alter --default client-id --add producer_byte_rate=1000000000 ``` ### [](#set-combined-user-and-client-quotas)Set combined user and client quotas You can set quotas for specific (user, client ID) combinations for fine-grained control. #### [](#user-with-specific-client)User with specific client To limit a specific user when using a specific client: ```bash rpk cluster quotas alter --add consumer_byte_rate=1000000 --name user=alice --name client-id=consumer-1 ``` User `alice` using `client-id=consumer-1` is limited to a 1 MB/s fetch rate. The same user with a different client ID would use a different quota (or fall back to less specific matches). To view combined quotas: ```bash rpk cluster quotas describe --name user=alice --name client-id=consumer-1 ``` #### [](#user-with-client-prefix)User with client prefix To set a shared quota for a user across multiple clients matching a prefix: ```bash rpk cluster quotas alter --add producer_byte_rate=3000000 --name user=bob --name client-id-prefix=app- ``` All clients used by user `bob` with a client ID starting with `app-` share a combined 3 MB/s produce quota. #### [](#default-user-with-specific-client)Default user with specific client To set a quota for a specific client across all users: ```bash rpk cluster quotas alter --add producer_byte_rate=500000 --default user --name client-id=payment-processor ``` Any user using `client-id=payment-processor` is limited to a 500 KB/s produce rate, unless they have a more specific quota configured. ### [](#bulk-manage-client-throughput-limits)Bulk manage client throughput limits To more easily manage multiple quotas, you can use the `cluster quotas describe` and [`cluster quotas import`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-quotas-import/) commands to do a bulk export and update. For example, to export all client quotas in JSON format: ```bash rpk cluster quotas describe --format json ``` `rpk cluster quotas import` accepts the output string from `rpk cluster quotas describe --format `: ```bash rpk cluster quotas import --from '{"quotas":[{"entity":[{"name":"analytics-consumer","type":"client-id"}],"values":[{"key":"consumer_byte_rate","values":"10000000"}]},{"entity":[{"name":"analytics-","type":"client-id-prefix"}],"values":[{"key":"producer_byte_rate","values":"10000000"},{"key":"consumer_byte_rate","values":"5000000"}]}]}' ``` You can also save the JSON or YAML output to a file and pass the file path in the `--from` flag. ### [](#view-throughput-limits-in-redpanda-cloud)View throughput limits in Redpanda Cloud You can also use Redpanda Cloud to view enforced limits. In the side menu, go to **Quotas**. ### [](#monitor-client-throughput)Monitor client throughput The following metrics provide insights into client throughput quota usage: - Client quota throughput per rule and quota type: - `/public_metrics` - [`redpanda_kafka_quotas_client_quota_throughput`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_kafka_quotas_client_quota_throughput) - Client quota throttling delay per rule and quota type, in seconds: - `/public_metrics` - [`redpanda_kafka_quotas_client_quota_throttle_time`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_kafka_quotas_client_quota_throttle_time) To identify which clients are actively connected and generating traffic, see [View connected client details](#view-connected-client-details). Quota metrics use the `redpanda_quota_rule` label to identify which quota was applied to a request. The label distinguishes between different entity types (user, client, or combinations). See the label values in [`redpanda_kafka_quotas_client_quota_throughput`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_kafka_quotas_client_quota_throughput). #### [](#track-quota-use-per-entity)Track quota use per entity When a workload slows down because a client hits its throughput quota, the aggregate quota metrics can confirm that throttling is happening at a broker level, but they cannot reveal which user, client, or group of clients is utilizing a quota or being throttled. Per-entity quota metrics answer these questions. Redpanda labels each throttle-time and throughput series with the identity of the throttled entity, so you can measure how much of an enforced quota each entity actually uses. Use these metrics to identify throttled users and clients by name and right-size quota values based on observed usage. Per-entity quota metrics are disabled by default because each throttled entity adds metric series. To enable them, set the `kafka_per_entity_quota_metrics` cluster property. See [Configure Cluster Properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/). The change takes effect without a broker restart. When enabled, Redpanda exposes two additional counters: - Total per-entity quota throttling delay, in milliseconds: - `/public_metrics` - `redpanda_kafka_quotas_client_quota_throttle_time_ms_by_entity` - Per-entity quota throughput (bytes for produce and fetch quotas, partition mutations for partition mutation quotas): - `/public_metrics` - `redpanda_kafka_quotas_client_quota_throughput_by_entity` Each series is labeled with the entity identity and the quota type: - `redpanda_quota_type`: Always present. One of `produce_quota`, `fetch_quota`, or `partition_mutation_quota`. - `redpanda_quota_user`: The user principal, for user-based quotas. - `redpanda_quota_client_id`: The client ID, for client ID-based quotas. - `redpanda_quota_group_name`: The client ID prefix, for quotas that apply to a [group of clients](#group-of-clients-throughput-limit). Combined quotas, such as a user with a specific client ID, include each matching label on the same series. To keep metric cardinality bounded, Redpanda registers a per-entity series only while an entity is actively being throttled, and removes the series after the entity has been idle. An entity appears in these metrics only after it has been throttled at least once. From that point, Redpanda records the entity’s throughput on every request, not only on throttled requests, so a Prometheus `rate()` query reflects the entity’s actual throughput. To measure how much of its quota an entity is using, divide the entity’s throughput rate by its configured limit. Quota limits apply per broker, so aggregate by broker as well as by entity to keep the results comparable with the configured limits. For example, the following query returns the produce throughput rate for each throttled client ID on each broker: ```promql sum by (redpanda_quota_client_id, pod) ( rate(redpanda_kafka_quotas_client_quota_throughput_by_entity{redpanda_quota_type="produce_quota"}[5m]) ) ``` Compare the result with the entity’s configured limit from [`rpk cluster quotas describe`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-cluster/rpk-cluster-quotas-describe/). A ratio close to 1 means the client is saturating its quota and its requests are being delayed. ## [](#see-also)See also - [About Client Throughput Quotas](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/about-throughput-quotas/) - [Configure Client Connections](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/configure-client-connections/) - [Authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/) --- # Page 393: Disaster Recovery **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery.md --- # Disaster Recovery > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Disaster Recovery latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/index.adoc description: Learn about disaster recovery options for Redpanda Cloud. page-git-created-date: "2025-12-12" page-git-modified-date: "2025-12-12" --- Shadowing complements Redpanda’s existing availability and recovery capabilities. High availability actively protects your day-to-day operations, handling reads and writes seamlessly during node or availability zone failures within a region. Shadowing is your safety net for catastrophic regional disasters. Shadowing delivers near real-time, cross-region replication for mission-critical applications that require rapid failover with minimal data loss. > 📝 **NOTE** > > Shadowing is supported on BYOC and Dedicated clusters running Redpanda version 25.3 and later. - [Shadowing](shadowing/) Learn about shadowing for disaster recovery in Redpanda Cloud. --- # Page 394: Shadowing **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing.md --- # Shadowing > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Shadowing latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/shadowing/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/shadowing/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/shadowing/index.adoc description: Learn about shadowing for disaster recovery in Redpanda Cloud. page-git-created-date: "2025-12-12" page-git-modified-date: "2025-12-12" --- > 📝 **NOTE** > > Shadowing is supported on BYOC and Dedicated clusters running Redpanda version 25.3 and later. - [Shadowing Overview](overview/) Overview of shadowing for disaster recovery in Redpanda Cloud. - [Configure Shadowing](setup/) Learn how to configure shadowing for disaster recovery. - [Migrate Schemas from Confluent Schema Registry](migrate-schemas-confluent/) Replicate subjects, versions, and compatibility settings from a Confluent Schema Registry into a Redpanda Cloud shadow cluster. - [Monitor Shadowing](monitor/) Learn how to monitor shadowing for disaster recovery. - [Configure Failover](failover/) Learn how to configure failover for disaster recovery. - [Failover Runbook](failover-runbook/) Step-by-step runbook for failover procedures in disaster recovery. --- # Page 395: Failover Runbook **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover-runbook.md --- # Failover Runbook > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Failover Runbook latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/shadowing/failover-runbook page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/shadowing/failover-runbook.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/shadowing/failover-runbook.adoc description: Step-by-step runbook for failover procedures in disaster recovery. page-git-created-date: "2025-12-12" page-git-modified-date: "2026-05-26" --- This guide provides step-by-step procedures for emergency failover when your primary Redpanda cluster becomes unavailable. Follow these procedures only during active disasters when immediate failover is required. > ❗ **IMPORTANT** > > This is an emergency procedure. For planned failover testing or day-to-day shadow link management, see [Configure Failover](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover/). Ensure you have completed the [disaster readiness checklist](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#disaster-readiness-checklist) before an emergency occurs. > 📝 **NOTE** > > Shadowing is supported on BYOC and Dedicated clusters running Redpanda version 25.3 and later. ## [](#emergency-failover-procedure)Emergency failover procedure Follow these steps during an active disaster: 1. [Assess the situation](#assess-situation) 2. [Verify shadow cluster status](#verify-shadow-status) 3. [Document current state](#document-state) 4. [Initiate failover](#initiate-failover) 5. [Monitor failover progress](#monitor-progress) 6. [Update application configuration](#update-applications) 7. [Verify application functionality](#verify-functionality) 8. [Clean up and stabilize](#cleanup-stabilize) ### [](#assess-situation)Assess the situation Confirm that failover is necessary: ```bash # Check if the primary cluster is responding rpk cluster info --brokers prod-cluster-1.example.com:9092,prod-cluster-2.example.com:9092 # If primary cluster is down, check shadow cluster health rpk cluster info --brokers shadow-cluster-1.example.com:9092,shadow-cluster-2.example.com:9092 ``` **Decision point**: If the primary cluster is responsive, consider whether failover is actually needed. Partial outages may not require full disaster recovery. **Examples that require full failover:** - Primary cluster is completely unreachable (network partition, regional outage) - Multiple broker failures preventing writes to critical topics - Data center failure affecting majority of brokers - Persistent authentication or authorization failures across the cluster **Examples that may NOT require failover:** - Single broker failure with sufficient replicas remaining - Temporary network connectivity issues affecting some clients - High latency or performance degradation (but cluster still functional) - Non-critical topic or partition unavailability ### [](#verify-shadow-status)Verify shadow cluster status Check the health of your shadow links: #### Cloud UI 1. From the **Shadow Link** page, select the shadow link you want to view. 2. The **Overview** tab shows the state of the shadow link and its topics. #### rpk ```bash # List all shadow links rpk shadow list # Check the configuration of your shadow link rpk shadow describe # Check the status of your disaster recovery link rpk shadow status ``` For detailed command options, see [`rpk shadow list`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-list/), [`rpk shadow describe`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-describe/), and [`rpk shadow status`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-status/). #### Cloud API ```bash # List all shadow links curl "https://api.redpanda.com/v1/shadow-links" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" # Check the configuration of your shadow link curl "https://api.redpanda.com/v1/shadow-links/" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" # Get Data Plane API URL of shadow cluster export DATAPLANE_API_URL=`curl https://api.cloud.redpanda.com/v1/clusters/ \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" | jq .cluster.dataplane_api` # Check the status of your disaster recovery link curl "https://$DATAPLANE_API_URL/v1/shadowlinks/" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` Verify that the following conditions exist before proceeding with failover: - Shadow link state should be `ACTIVE`. - Topics should be in `ACTIVE` state (not `FAULTED`). - Replication lag should be reasonable for your RPO requirements. #### [](#understanding-replication-lag)Understanding replication lag Use [`rpk shadow status`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-status/) or the [Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-shadowlinkservice_listshadowlinktopics) to check lag, which shows the message count difference between source and shadow partitions: - **Acceptable lag examples**: 0-1000 messages for low-throughput topics, 0-10000 messages for high-throughput topics - **Concerning lag examples**: Growing lag over 50,000 messages, or lag that continuously increases without recovering - **Critical lag examples**: Lag exceeding your data loss tolerance (for example, if you can only afford to lose 1 minute of data, lag should represent less than 1 minute of typical message volume) ### [](#document-state)Document current state Record the current lag and status before proceeding: #### Cloud UI Capture the status from the **Shadow Link** page. #### rpk ```bash # Capture current status for post-mortem analysis rpk shadow status > failover-status-$(date +%Y%m%d-%H%M%S).log ``` Example output showing healthy replication before failover: shadow link: Overview: NAME UID STATE ACTIVE Tasks: Name Broker\_ID State Reason 1 ACTIVE 2 ACTIVE Topics: Name: , State: ACTIVE Partition SRC\_LSO SRC\_HWM DST\_HWM Lag 0 1234 1468 1456 12 1 2345 2579 2568 11 #### Cloud API ```bash # Capture current status for post-mortem analysis curl "https://$DATAPLANE_API_URL/v1/shadowlinks//topic" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" > failover-status-$(date +%Y%m%d-%H%M%S).log ``` The partition information shows the following: | Field | Description | | --- | --- | | source_last_stable_offset | Source partition last stable offset | | source_high_watermark | Source partition high watermark | | high_watermark | Shadow (destination) partition high watermark | | Lag | Message count difference between source and shadow partitions | > ❗ **IMPORTANT** > > Note the replication lag to estimate potential data loss during failover. The `Tasks` section shows the health of shadow link replication tasks. For details about what each task does, see [Shadow link tasks](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#shadow-link-tasks). ### [](#initiate-failover)Initiate failover A complete cluster failover is appropriate If you observe that the source cluster is no longer reachable: #### Cloud UI 1. On your **Shadow Link** page, click **Failover All Topics**. 2. Click to confirm the failover action. The failover process promotes all topics to writable status. #### rpk ```bash # Fail over all topics in the shadow link rpk shadow failover --all ``` For detailed command options, see [`rpk shadow failover`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-failover/). #### Cloud API ```bash # Fail over all topics in the shadow link curl -X POST "$DATAPLANE_API_URL/v1/shadowlink//failover" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` For selective topic failover (when only specific services are affected): #### Cloud UI 1. On your **Shadow Link** page, click the **Failover** button for the topics you want to failover. 2. Click to confirm the failover action. The failover process promotes the selected topics to writable status. #### rpk ```bash # Fail over individual topics rpk shadow failover --topic rpk shadow failover --topic ``` #### Cloud API ```bash # Fail over individual topics curl -X POST "$DATAPLANE_API_URL/v1/shadowlinks//failover" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d '{ "shadowTopicName": "" }' curl -X POST "$DATAPLANE_API_URL/v1/shadowlinks//failover" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d '{ "shadowTopicName": "" }' ``` ### [](#monitor-progress)Monitor failover progress Track the failover process: #### Cloud UI 1. From the **Shadow Link** page, select the shadow link you want to view. 2. Click the **Tasks** tab to view all tasks and their status. #### rpk ```bash # Monitor status until all topics show FAILED_OVER watch -n 5 "rpk shadow status " # Check detailed topic status and lag during emergency rpk shadow status --print-topic ``` Example output during successful failover: shadow link: Overview: NAME UID STATE ACTIVE Tasks: Name Broker\_ID State Reason 1 ACTIVE 2 ACTIVE Topics: Name: , State: FAILED\_OVER Name: , State: FAILED\_OVER Name: , State: FAILING\_OVER #### Cloud API ```bash # Monitor status watch -n 5 'curl "https://$DATAPLANE_API_URL/v1/shadowlinks/" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" | jq .' # Check detailed topic status and lag during emergency curl "https://$DATAPLANE_API_URL/v1/shadowlinks//topic" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` **Wait for**: All critical topics to reach `FAILED_OVER` state before proceeding. ### [](#update-applications)Update application configuration Redirect your applications to the shadow cluster by updating connection strings in your applications to point to shadow cluster brokers. If using DNS-based service discovery, update DNS records accordingly. Restart applications to pick up new connection settings and verify connectivity from application hosts to shadow cluster. ### [](#verify-functionality)Verify application functionality Test critical application workflows: ```bash # Verify applications can produce messages rpk topic produce --brokers :9092 # Verify applications can consume messages rpk topic consume --brokers :9092 --num 1 ``` Test message production and consumption, consumer group functionality, and critical business workflows to ensure everything is working properly. ### [](#cleanup-stabilize)Clean up and stabilize After all applications are running normally: #### Cloud UI 1. On your **Shadow Link** page, click **Delete**. 2. Type "delete" to confirm the action. #### rpk ```bash # Optional: Delete the shadow link (no longer needed) rpk shadow delete ``` For detailed command options, see [`rpk shadow delete`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-delete/). #### Cloud API ```bash # Optional: Delete the shadow link (no longer needed) curl -X DELETE https://api.redpanda.com/v1/shadow-links/ \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` For the full API reference, see [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-shadowlinkservice_deleteshadowlink). > 📝 **NOTE** > > This operation [force deletes](#force-delete-warning) the shadow link. Document the time of failover initiation and completion, applications affected and recovery times, data loss estimates based on replication lag, and issues encountered during failover. ## [](#troubleshoot-common-issues)Troubleshoot common issues > 📝 **NOTE** > > Direct access to shadow cluster logs isn’t available in Redpanda Cloud. To troubleshoot, use: > > - Run [`rpk shadow status`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-status/) to check task states, topic status, and lag. > > - Use the [Data Plane API](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-shadowlinkservice_listshadowlinktopics) for programmatic monitoring. > > - Monitor [Prometheus metrics](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/monitor/#shadow-link-metrics) such as `redpanda_shadow_link_shadow_lag` and `redpanda_shadow_link_client_errors`. > > > If these don’t reveal the root cause, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new) with your cluster ID, shadow link name, and the timestamp of the issue. ### [](#topics-stuck-in-failing_over-state)Topics stuck in FAILING_OVER state **Problem**: Topics remain in `FAILING_OVER` state for extended periods **Solution**: Use `rpk shadow status`, the Data Plane API, or Prometheus metrics to check task and topic states, replication lag, and client-error counts, and ensure sufficient cluster resources (CPU, memory, disk space) are available on the shadow cluster. Verify network connectivity between shadow cluster nodes and confirm that all shadow topic partitions have elected leaders and the controller partition is properly replicated with an active leader. If topics remain stuck after addressing these cluster health issues and you need immediate failover, you can force delete the shadow link to failover all topics: #### Cloud UI All failover actions in the Cloud UI include force delete functionality by default. When you failover a shadow link, all topics are immediately promoted to writable status. #### rpk ```bash # Force delete the shadow link to failover all topics rpk shadow delete ``` `rpk shadow delete` force deletes the shadow link by default in Redpanda Cloud. #### Cloud API ```bash curl -X DELETE https://api.redpanda.com/v1/shadow-links/ \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` The `DELETE /shadow-links/` endpoint of the Control Plane API force deletes the shadow link by default in Redpanda Cloud. > ⚠️ **WARNING** > > Force deleting a shadow link immediately fails over all topics in the link. This action is irreversible and should only be used when topics are stuck and you need immediate access to all replicated data. ### [](#topics-in-faulted-state)Topics in FAULTED state **Problem**: Topics show `FAULTED` state and are not replicating **Solution**: Check for authentication issues, network connectivity problems, or source cluster unavailability. Verify that the shadow link service account still has the required permissions on the source cluster. Use `rpk shadow status`, the Data Plane API, or Prometheus metrics to check task and topic states, replication lag, and client-error counts for the faulted topics. ### [](#application-connection-failures)Application connection failures **Problem**: Applications cannot connect to shadow cluster after failover **Solution**: Verify shadow cluster broker endpoints are correct and check security group and firewall rules. Confirm authentication credentials are valid for the shadow cluster and test network connectivity from application hosts. ### [](#consumer-group-offset-issues)Consumer group offset issues **Problem**: Consumers start from beginning or wrong positions **Solution**: Verify consumer group offsets were replicated (check your filters) and use `rpk group describe ` to check offset positions. If necessary, manually reset offsets to appropriate positions. See [How to manage consumer group offsets in Redpanda](https://support.redpanda.com/hc/en-us/articles/23499121317399-How-to-manage-consumer-group-offsets-in-Redpanda) for detailed reset procedures. ## [](#next-steps)Next steps After successful failover, focus on recovery planning and process improvement. Begin by assessing the source cluster failure and determining whether to restore the original cluster or permanently promote the shadow cluster as your new primary. **Immediate recovery planning:** 1. **Assess source cluster**: Determine root cause of the outage 2. **Plan recovery**: Decide whether to restore source cluster or promote shadow cluster permanently 3. **Data synchronization**: Plan how to synchronize any data produced during failover 4. **Fail forward**: Create a new shadow link with the failed over shadow cluster as source to maintain a DR cluster **Process improvement:** 1. **Document the incident**: Record timeline, impact, and lessons learned 2. **Update runbooks**: Improve procedures based on what you learned 3. **Test regularly**: Schedule regular disaster recovery drills 4. **Review monitoring**: Ensure monitoring caught the issue appropriately --- # Page 396: Configure Failover **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover.md --- # Configure Failover > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Failover latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/shadowing/failover page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/shadowing/failover.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/shadowing/failover.adoc description: Learn how to configure failover for disaster recovery. page-git-created-date: "2025-12-12" page-git-modified-date: "2026-05-26" --- Failover is the process of modifying shadow topics or an entire shadow cluster from read-only replicas to fully writable resources, and ceasing replication from the source cluster. You can fail over individual topics for selective workload migration or fail over the entire cluster for comprehensive disaster recovery. This critical operation transforms your shadow resources into operational production assets, allowing you to redirect application traffic when the source cluster becomes unavailable. You can failover a shadow link using the Redpanda Cloud UI, `rpk`, or the Data Plane API. > ❗ **IMPORTANT: Experiencing an active disaster?** > > See [Failover Runbook](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover-runbook/) for immediate step-by-step disaster procedures. > 📝 **NOTE** > > Shadowing is supported on BYOC and Dedicated clusters running Redpanda version 25.3 and later. ## [](#failover-behavior)Failover behavior When you initiate failover, Redpanda performs the following operations: 1. **Stops replication**: Halts all data fetching from the source cluster for the specified topics or entire shadow link 2. **Failover topics**: Converts read-only shadow topics into regular, writable topics 3. **Updates topic state**: Changes topic status from `ACTIVE` to `FAILING_OVER`, then `FAILED_OVER` Topic failover is irreversible. Once failed over, topics cannot return to shadow mode, and automatic fallback to the original source cluster is not supported. > 📝 **NOTE** > > To avoid a split-brain scenario after failover, ensure that all clients are reconfigured to point to the shadow cluster before resuming write activity. ## [](#failover-commands)Failover commands ### [](#get-data-plane-api-url)Get Data Plane API URL If using the Data Plane API, run the following to get the Data Plane API URL of the shadow cluster: ```bash export DATAPLANE_API_URL=`curl https://api.cloud.redpanda.com/v1/clusters/ \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" | jq .cluster.dataplane_api` ``` You can perform failover at different levels of granularity to match your disaster recovery needs: ### [](#individual-topic-failover)Individual topic failover To fail over a specific shadow topic while leaving other topics in the shadow link still replicating, run: #### Cloud UI 1. On the **Shadow Link** page, select your shadow link. 2. For any of the topics you want to failover, click the corresponding **Failover** button. 3. Click to confirm the failover action. The failover process promotes the selected topics to writable status. #### rpk ```bash rpk shadow failover --topic ``` For detailed command options, see [`rpk shadow failover`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-failover/). #### Data Plane API Send a `POST /shadowlink/{shadow_link_name}/failover` request to the Data Plane API. Specify the name of the shadow topic in the request body: ```bash curl -X POST "$DATAPLANE_API_URL/v1/shadowlink//failover" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d '{ "shadowTopicName": "" }' ``` Use this approach when you need to selectively failover specific workloads or when testing failover procedures. ### [](#complete-shadow-link-failover-cluster-failover)Complete shadow link failover (cluster failover) To fail over all shadow topics associated with the shadow link simultaneously, run: #### Cloud UI 1. On the **Shadow Link** page, select your shadow link. 2. Click **Failover All Topics**. 3. Click to confirm the failover action. The failover process promotes all topics to writable status. #### rpk ```bash rpk shadow failover --all ``` #### Data Plane API Send a `POST /shadowlink/{shadow_link_name}/failover` request to the Data Plane API. If you do not specify a shadow topic in the request body, this command requests a failover of all shadow topics associated with the shadow link: ```bash curl -X POST "$DATAPLANE_API_URL/v1/shadowlink//failover" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` Use this approach during a complete regional disaster when you need to activate the entire shadow cluster as your new production environment. ### [](#force-delete-shadow-link-emergency-failover)Force delete shadow link (emergency failover) #### Cloud UI All failover actions in the Cloud UI include force delete functionality by default. When you failover a shadow link, all topics are immediately promoted to writable status. #### rpk `rpk shadow delete` force deletes the shadow link by default in Redpanda Cloud: ```bash rpk shadow delete ``` #### Control Plane API Use the Control Plane API to force delete a shadow link: ```bash curl -X DELETE 'https://api.redpanda.com/v1/shadow-links/' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` > ⚠️ **WARNING** > > Force deleting a shadow link is irreversible and immediately fails over all topics in the link, bypassing the normal failover state transitions. This action should only be used as a last resort when topics are stuck in transitional states and you need immediate access to all replicated data. ## [](#failover-states)Failover states ### [](#shadow-link-states)Shadow link states The shadow link itself has a simple state model: - **`ACTIVE`**: Shadow link is operating normally, replicating data - **`PAUSED`**: Shadow link replication is temporarily halted by user action Shadow links do not have dedicated failover states. Instead, the link’s operational status is determined by the collective state of its shadow topics. ### [](#shadow-topic-states)Shadow topic states Individual shadow topics progress through specific states during failover: - **`ACTIVE`**: Normal replication state before failover - **`FAULTED`**: Shadow topic has encountered an error and is not replicating - **`FAILING_OVER`**: Failover initiated, replication stopping - **`FAILED_OVER`**: Failover completed successfully, topic fully writable - **`PAUSED`**: Replication temporarily halted by user action ## [](#monitor-failover-progress)Monitor failover progress To monitor failover progress using the status command, run: ### Cloud UI Track the progress of failover operations from the **Shadow Link** page in the Cloud UI. ### rpk ```bash rpk shadow status ``` The output shows individual topic states and any issues encountered during the failover process. For detailed command options, see [`rpk shadow status`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-status/). ### Data Plane API ```bash curl "https://$DATAPLANE_API_URL/v1/shadowlinks/" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` Task states during monitoring: - **`ACTIVE`**: Task is operating normally and replicating data - **`FAULTED`**: Task encountered an error and requires attention - **`NOT_RUNNING`**: Task is not currently executing - **`LINK_UNAVAILABLE`**: Task cannot communicate with the source cluster For detailed information about shadow link tasks and their roles, see [Shadow link tasks](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#shadow-link-tasks). ## [](#post-failover-cluster-behavior)Post-failover cluster behavior After successful failover, your shadow cluster exhibits the following characteristics: **Topic accessibility:** - Failed over topics become fully writable and readable. - Applications can produce and consume messages normally. - All Kafka APIs are available for failedover topics. - Original offsets and timestamps are preserved. **Shadow link status:** - The shadow link remains but stops replicating data. - Link status shows topics in `FAILED_OVER` state. - You can safely delete the shadow link after successful failover. **Operational limitations:** - No automatic fallback mechanism to the original source cluster. - Data transforms remain disabled until you manually re-enable them. - Audit log history from the source cluster is not available (new audit logs begin immediately). ## [](#failover-considerations-and-limitations)Failover considerations and limitations Before implementing failover procedures, understand these key considerations that affect your disaster recovery strategy and operational planning. **Data consistency:** - Some data loss may occur due to replication lag at the time of failover. - Consumer group offsets are preserved, allowing applications to resume from their last committed position. - In-flight transactions at the source cluster are not replicated and will be lost. **Recovery-point-objective (RPO):** The amount of potential data loss depends on replication lag when disaster occurs. Monitor lag metrics to understand your effective RPO. **Network partitions:** If the source cluster becomes accessible again after failover, do not attempt to write to both clusters simultaneously. This creates a scenario with potential data inconsistencies, since metadata starts to diverge. **Testing requirements:** Regularly test failover procedures in non-production environments to validate your disaster recovery processes and measure RTO. ## [](#next-steps)Next steps After completing failover: - Update your application connection strings to point to the shadow cluster - Verify that applications can produce and consume messages normally - Consider deleting the shadow link if failover was successful and permanent For emergency situations, see [Failover Runbook](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover-runbook/). --- # Page 397: Migrate Schemas from Confluent Schema Registry **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent.md --- # Migrate Schemas from Confluent Schema Registry > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Migrate Schemas from Confluent Schema Registry latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/shadowing/migrate-schemas-confluent page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/shadowing/migrate-schemas-confluent.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/shadowing/migrate-schemas-confluent.adoc description: Replicate subjects, versions, and compatibility settings from a Confluent Schema Registry into a Redpanda Cloud shadow cluster. learning-objective-1: Configure a shadow link that continuously replicates schemas from a Confluent Schema Registry learning-objective-2: Filter replication by context or subject and map source contexts to destination contexts learning-objective-3: Monitor schema replication status and resolve validation errors page-git-created-date: "2026-08-25" page-git-modified-date: "2026-08-25" --- When you migrate to Redpanda from a deployment that uses a Confluent Schema Registry, your producers and consumers depend on the [schemas](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#schema) stored in that registry. Shadowing removes this migration obstacle: a [shadow link](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#shadow-link) continuously replicates schemas from the source Confluent Schema Registry into the [Schema Registry](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#schema-registry) built into the Redpanda [shadow cluster](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#shadow-cluster), preserving [subject](https://docs.redpanda.com/cloud-data-platform/reference/glossary/#subject) names, versions, and compatibility settings. Because both registries stay synchronized until cutover, your applications keep working on Redpanda without a separate schema migration step. Use this approach when you migrate from Confluent to Redpanda, or when you maintain a Redpanda disaster recovery cluster for a system that keeps its schemas in a Confluent Schema Registry. After reading this page, you will be able to: - Configure a shadow link that continuously replicates schemas from a Confluent Schema Registry - Filter replication by context or subject and map source contexts to destination contexts - Monitor schema replication status and resolve validation errors ## [](#how-http-api-schema-replication-works)How HTTP API schema replication works When you configure a shadow link with the `shadow_schema_registry_api` option, the shadow cluster polls the source Schema Registry over HTTP and imports changes into its own Schema Registry. Two sync cycles keep the registries in step: - **Tail syncs** run frequently (default: every 10 seconds) to pick up incremental changes. - **Full syncs** scan all selected subjects (default: every 5 minutes) to catch anything a tail sync missed. Replicated schemas keep their original subject names and version IDs, so producers and consumers that reference schemas by ID continue to work after failover. Schemas that reference other schemas import in dependency order. Before importing a schema, Redpanda validates it against the Redpanda Schema Registry implementation. If a schema uses features that Redpanda does not support, the sync either reports an error and skips the schema, or removes the unsupported fields and imports the rest, depending on the [validation policy](#choose-a-validation-policy) you choose. While the link is active, the destination contexts that the link replicates into are read-only: the shadow cluster rejects client writes to those contexts so that replicated schemas remain identical to the source. Contexts outside the link’s filter remain writable. > 📝 **NOTE** > > This API-based mode is an alternative to the byte-for-byte `_schemas` topic replication described in [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/#schema-registry-synchronization). A shadow link uses one Schema Registry sync mode or the other, not both: > > - Use **topic mode** (`shadow_schema_registry_topic`) when the source is another Redpanda cluster and you want an exact, complete replica of its Schema Registry. Topic mode shadows the `_schemas` topic byte for byte, so it does not filter, remap, or validate schemas. > > - Use **API mode** (`shadow_schema_registry_api`) when the source is a Confluent Schema Registry, or when you need to replicate only selected contexts or subjects, map source contexts to different destination contexts, or control how schemas that use unsupported features are handled with a validation policy. > > > Schema replication settings live in the shadow link configuration. Two cluster properties, `schema_registry_sync_memory_bytes` and `schema_registry_sync_parallelism`, tune how much memory and concurrency the shadow cluster uses while importing schemas. The defaults suit most deployments. ## [](#use-cases)Use cases - **Migrate from Confluent to Redpanda**: Replicate schemas continuously while [Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/) replicates your topic data, then cut applications over to Redpanda once both are in sync. No separate schema migration tooling is required. - **Phased migration**: Use context and subject filters to migrate one team, application, or environment at a time. - **Registry reorganization**: Map source contexts to different destination contexts to restructure your Schema Registry as part of the migration. ## [](#prerequisites)Prerequisites - A cluster running Redpanda version 26.2 or later. The schema replication feature activates after all brokers complete the upgrade. - A BYOC or Dedicated cluster. API-mode Schema Registry replication is rolling out to Redpanda Cloud; if the Schema Registry options described on this page are not yet available for your cluster, contact [Redpanda Support](https://support.redpanda.com/hc/en-us). - Network connectivity from the shadow cluster to the source Schema Registry HTTP endpoint. - Credentials for the source registry with permission to read subjects, versions, and configuration. For Confluent Cloud, use a Schema Registry API key and secret. - Basic Shadow link settings. See [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/). - The destination contexts that the link replicates into, as determined by your `source_filter` and `destination` mapping, must be empty on the shadow cluster. The rest of the shadow cluster’s Schema Registry does not need to be empty: contexts outside the link’s mappings are unaffected. - To replicate contexts other than the default context, the [`schema_registry_enable_qualified_subjects`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#schema_registry_enable_qualified_subjects) cluster property must be enabled on the shadow cluster (the default). See [Schema Registry Contexts](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/#prerequisites). ## [](#limitations)Limitations - HTTP basic authentication and mTLS are the supported authentication methods for the source registry. You can also connect to a source registry that requires no authentication. - Replication is one way, from the source registry to the shadow cluster. Destination contexts owned by the link are read-only until failover. - Schemas that use Confluent features not supported by the Redpanda Schema Registry are not replicated as-is. Choose a [validation policy](#choose-a-validation-policy) to control whether these schemas are skipped or imported without the unsupported fields. - Topic data replication and schema replication are coordinated, but not immediately consistent: schema arrival on the shadow cluster can be delayed by the tail refresh interval (`tail_interval`), which defaults to 10 seconds. Records serialized in the Confluent SerDes wire format can therefore arrive before the schema IDs they contain. This does not cause replication errors, because schema IDs are not validated during topic data replication, but consumers that look up those schema IDs on the shadow cluster fail until the schemas arrive. See [Monitor replication status](#monitor-replication-status). - Role synchronization requires a Redpanda source. Leave `role_sync_options` unconfigured when the source is a Confluent cluster: a Roles Migrator task pointed at a non-Redpanda source reports `LINK_UNAVAILABLE` while the link itself stays `ACTIVE`. - Deleting and recreating a subject: A hard delete on the source replicates on the next sync, not instantly. Wait until the subject is removed from the shadow cluster before registering a new schema under the same name, as recreating it too soon prevents the sync job from detecting the subject/subject-version change. ## [](#configure-schema-replication)Configure schema replication Schema replication is one synchronization task on a shadow link. You configure it by adding the `shadow_schema_registry_api` option to the `schema_registry_sync_options` section of the same shadow link that replicates your topic data, consumer offsets, and ACLs. The sections below extend the shadow link configuration described in [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/#create-a-shadow-link); they do not replace it. > ❗ **IMPORTANT** > > A shadow link whose configuration contains only `schema_registry_sync_options` replicates schemas and nothing else. For a full migration, keep your topic, consumer offset, and security synchronization options in the same configuration file. This workflow builds the configuration file that `rpk shadow create` consumes. The same settings map one-to-one onto the Cloud UI fields, the Control Plane API request, and the Terraform resource shown in [Create the shadow link](#create-the-shadow-link). The following collapsible sample shows where the schema replication settings sit in a complete shadow link configuration file. The highlighted lines are the schema replication settings, which the sections that follow explain. Explore a sample configuration file ```yaml # Sample shadow link configuration with API-mode Schema Registry replication name: confluent-migration # Unique name for this shadow link client_options: bootstrap_servers: # Source Kafka cluster brokers - : # Example: "pkc-xxxxx.us-east-1.aws.confluent.cloud:9092" - : # For the TLS and authentication settings that the shadow cluster uses to # connect to the source Kafka cluster, see the complete configuration file # reference in Configure Shadowing. # Keep the rest of your migration in the same file. Without these sections, # the link replicates schemas only. topic_metadata_sync_options: # ...your topic replication settings... consumer_offset_sync_options: # ...your consumer offset settings... security_sync_options: # ...your ACL replication settings... schema_registry_sync_options: shadow_schema_registry_api: # API mode: replicate from a Confluent Schema Registry source_url: https://psrc-xxxxx.us-east-1.aws.confluent.cloud # Source Schema Registry endpoint auth_options: basic: username: # Confluent Schema Registry API key password: # Confluent Schema Registry API secret tls_settings: enabled: true # Use TLS for the connection tail_interval: 10s # Optional: poll for incremental changes (default: 10s) full_sync_interval: 5m # Optional: full source scan interval (default: 5m) max_source_requests_per_second: 30 # Optional: rate limit for source requests (default: 30) source_filter: contexts: - "." # The default context subjects: [] # Empty: all subjects in the selected contexts destination: identity: {} # Keep source context names unsupported_schema_feature_policy: FAIL # FAIL (default) or REMOVE ``` For the complete configuration file, including consumer offset and security synchronization options, see [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/#create-a-shadow-link). ### [](#generate-a-configuration-template)Generate a configuration template Generate a configuration file template that includes all available fields with comments: ```bash rpk shadow config generate --print-template -o shadow-config-template.yaml ``` For detailed command options, see [`rpk shadow config generate`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-config-generate/). ### [](#connect-to-the-source-registry)Connect to the source registry Configure the connection to the source Schema Registry: ```yaml schema_registry_sync_options: shadow_schema_registry_api: source_url: https://psrc-xxxxx.us-east-1.aws.confluent.cloud # Source Schema Registry endpoint auth_options: basic: username: # Confluent Schema Registry API key password: # Confluent Schema Registry API secret tls_settings: enabled: true # Use TLS for the connection tail_interval: 10s # How often to poll for incremental changes full_sync_interval: 5m # How often to run a full scan max_source_requests_per_second: 30 # Rate limit for requests to the source registry ``` The intervals and rate limit are optional. If you omit them, Redpanda uses the defaults shown above. To authenticate to the source registry with mTLS instead of HTTP basic authentication, omit `auth_options` and provide a client certificate and key in `tls_settings`. To connect to a source registry that requires no authentication, omit `auth_options` and do not provide a client certificate. ### [](#select-contexts-and-subjects)Select contexts and subjects By default, the link replicates the entire source registry. To replicate a subset, add a `source_filter` with the contexts or subjects to include: ```yaml schema_registry_sync_options: shadow_schema_registry_api: # ...connection settings... source_filter: contexts: - ".prod" # Replicate this entire context subjects: - orders-value # One subject from the default context ``` The two lists combine as a union: - `contexts` selects entire contexts: every subject in each listed context replicates. - `subjects` selects individual subjects, using qualified subject syntax: `orders-value` is the subject in the default context, and `:.staging:orders-value` is the subject of the same name in the `.staging` context. - When both lists are set, the link replicates everything selected by either list. A subject selected by both lists replicates once. For example, the preceding filter replicates every subject in the `.prod` context, plus the single `orders-value` subject from the default context. The `contexts` and `subjects` lists accept literal names only. Wildcard and prefix patterns are not supported. Schema Registry contexts provide independent namespaces for subjects within one registry. The default context is named `.`. For more information, see [Schema Registry Contexts](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/). ### [](#map-source-contexts-to-destination-contexts)Map source contexts to destination contexts Choose how replicated contexts are named on the shadow cluster: - `identity`: Keep the source context names (default behavior for migrations). - `exact`: Map each source context to a different destination context. ```yaml schema_registry_sync_options: shadow_schema_registry_api: # ...connection settings and filters... destination: identity: {} # Keep source context names ``` To rename contexts during replication: ```yaml schema_registry_sync_options: shadow_schema_registry_api: # ...connection settings and filters... destination: exact: mappings: - source: "." # Source context destination: ".shadow" # Destination context on the shadow cluster ``` > ❗ **IMPORTANT** > > With `exact` mapping, the mappings must cover every context that the link replicates. If the link encounters a source context that has no mapping, the schema replication task fails. If new contexts might be created on the source registry after you set up the link, scope the `source_filter` `contexts` list to the mapped contexts so that an unexpected context cannot stop replication. ### [](#choose-a-validation-policy)Choose a validation policy The `unsupported_schema_feature_policy` setting controls what happens when a source schema uses features that the Redpanda Schema Registry does not support. The unsupported Confluent Schema Registry features are: - In schema definitions: rule sets and metadata tags. - In subject configurations: override metadata, override rule sets, default metadata, default rule sets, and compatibility groups. Compatibility groups are not the same as compatibility levels, which do replicate. The policy determines how the sync handles a schema or configuration that uses these features: | Policy | Behavior | | --- | --- | | FAIL (default) | The schema is not replicated. The sync records an error, reports it in the link status, and continues with the remaining schemas. | | REMOVE | The unsupported fields are removed and the rest of the schema is imported. The sync records each modification in the link status. | ```yaml schema_registry_sync_options: shadow_schema_registry_api: # ...connection settings, filters, and destination... unsupported_schema_feature_policy: FAIL ``` ### [](#create-the-shadow-link)Create the shadow link Create the shadow link with the method that fits your workflow. All four methods configure the same settings. #### Cloud UI The create shadow link wizard includes a Schema Registry section on the **Configuration** step: 1. At the organization level of the Cloud UI, navigate to **Shadow Link** and click **Create shadow link**. 2. Complete the **Connection**, **Shadow**, and **Source** steps as described in [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/#create-a-shadow-link). For a Confluent source, select **Bootstrap URL** on the **Connection** step and enter the Confluent Kafka bootstrap servers. 3. On the **Configuration** step, in the Schema Registry section, select the Schema Registry HTTP API mode and configure: 1. The source Schema Registry URL. 2. Authentication: none, or HTTP basic. For Confluent Cloud, use a Schema Registry API key and secret. 3. TLS, with an optional custom CA certificate or an mTLS client certificate. 4. The scope: every context and subject, or specific ones. 5. Destination contexts: preserve the source context names, or map each source context to a destination context. 6. Optionally, the sync behavior: tail interval, full sync interval, maximum source request rate, and the unsupported schema features policy. 4. The wizard validates the connection details and shows the subjects and versions that match your filters before it creates the link. 5. Click **Create shadow link**. #### rpk Create the shadow link with your completed configuration file: ```bash rpk shadow create --config-file shadow-config.yaml ``` For detailed command options, see [`rpk shadow create`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-create/). To change the schema replication settings on an existing link, see [`rpk shadow update`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-update/). #### Control Plane API To create the link programmatically, include `schema_registry_sync_options` in the `POST /v1/shadow-links` request. The following request configures the same API-mode replication as the configuration file above: ```bash curl -X POST 'https://api.redpanda.com/v1/shadow-links' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d '{ "shadow_link": { "shadow_redpanda_id": "", "name": "", "client_options": { "bootstrap_servers": [":"], "tls_settings": { "enabled": true } }, "schema_registry_sync_options": { "shadow_schema_registry_api": { "source_url": "https://psrc-xxxxx.us-east-1.aws.confluent.cloud", "auth_options": { "basic": { "username": "", "password": "" } }, "tls_settings": { "enabled": true }, "source_filter": { "contexts": ["."], "subjects": [] }, "destination": { "identity": {} }, "unsupported_schema_feature_policy": "UNSUPPORTED_SCHEMA_FEATURE_POLICY_FAIL" } } } }' ``` For the source cluster connection settings (`client_options`), authentication, and secret handling, see the Control Plane API tab in [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/#create-a-shadow-link). For the full request schema, see the [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-shadowlinkservice_createshadowlink). #### Terraform To manage the shadow link declaratively instead of with `rpk`, use the [`redpanda_shadow_link` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/shadow_link) in the [Redpanda Terraform provider](https://registry.terraform.io/providers/redpanda-data/redpanda/latest). API-mode Schema Registry replication requires provider version 2.3.0 or later. The configuration keys described on this page map to the resource’s `schema_registry_sync_options.shadow_schema_registry_api` attribute: ```hcl resource "redpanda_shadow_link" "confluent_migration" { # Connection settings for the source Kafka cluster (client_options, # cluster IDs, and secret handling) are the same as for any shadow link. # See the Terraform tab in Configure Shadowing. schema_registry_sync_options = { shadow_schema_registry_api = { source_url = "https://psrc-xxxxx.us-east-1.aws.confluent.cloud" auth_options = { basic = { username = "" password = "$${secrets.${redpanda_secret.sr_api_secret.name}}" } } tls_settings = { enabled = true } source_filter = { contexts = ["."] } destination = { identity = true } unsupported_schema_feature_policy = "FAIL" } } } ``` Replace the placeholders with your own values: - ``: Confluent Schema Registry API key. - The `password` value references a [`redpanda_secret` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/secret) named `sr_api_secret` that stores the Confluent Schema Registry API secret in the shadow cluster’s secret store. The doubled dollar sign (`$$`) keeps Terraform from interpolating the `${secrets…​.}` reference, which Redpanda resolves at connection time. Initialize the working directory if you have not already done so, review the plan, and apply the configuration: ```bash terraform init terraform plan terraform apply ``` The attribute shapes differ from the YAML configuration file in one place: `destination` takes `identity = true` to preserve source context names, or an `exact` block with explicit source-to-destination `mappings`. For the full attribute schema, see the [`redpanda_shadow_link` reference](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/shadow_link). For a complete working configuration, including the source cluster connection and secret resources, see the [provider example](https://github.com/redpanda-data/terraform-provider-redpanda/blob/main/examples/shadow_link/main.tf) and the Terraform tab in [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/#create-a-shadow-link). ## [](#verify-the-configuration)Verify the configuration Confirm that the link is configured for API-based schema replication: ```bash rpk shadow describe --print-registry ``` The `--print-registry` flag prints the Schema Registry section, which includes the shadowing mode, source URL, sync intervals, validation policy, and your context and subject filters. Without it, `rpk shadow describe` prints only the overview and client sections. For detailed command options, see [`rpk shadow describe`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-describe/). ## [](#monitor-replication-status)Monitor replication status Check schema replication progress and errors for a link: ```bash rpk shadow status ``` The Schema Registry section of the output reports: | Field | Description | | --- | --- | | Inventory | The number of selected subjects and subject versions on the source, compared with the number of subjects and versions on the destination. The registries are synchronized when the destination counts match the selected source counts. A destination that is behind the source indicates that replication is still in progress. | | Current sync | The type of the sync in progress (FULL or TAIL) and the number of subject versions, compatibility configurations, and modes it has changed, including how many unsupported features were removed and how many errors occurred. | | Last full sync | Start time, finish time, and change counts for the most recent completed full sync. | | Totals since task start | Cumulative change and error counts since the schema replication task started. | | Last error | The most recent replication error. With the FAIL validation policy, schemas that fail validation appear here. | The schema replication task runs on the broker and shard that hosts the leader of the `_schemas` topic’s partition. The counters in the status output are not persisted: expect them to reset to zero when that broker restarts or when leadership of the `_schemas` partition moves to another broker. For detailed command options, see [`rpk shadow status`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-status/). For general link monitoring, see [Monitor Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/monitor/). ## [](#fail-over)Fail over For best results, shadow links should be failed over as a single unit, which prevents partial failover scenarios and unexpected results. Selective failover, such as topics only, does not stop schema replication. As part of failover, pause the schema replication task by setting `paused: true` in the `schema_registry_sync_options` section of the link configuration. Pausing the task stops further syncing from the source registry and makes the write-blocked destination contexts writable, so your applications can register new schemas on the promoted cluster. For the complete failover procedure, see [Failover](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover/). ## [](#next-steps)Next steps - [Configure Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/) to replicate topic data, consumer offsets, and ACLs alongside your schemas. - [Monitor Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/monitor/) - [Schema Registry Contexts](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-contexts/) --- # Page 398: Monitor Shadowing **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/monitor.md --- # Monitor Shadowing > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Monitor Shadowing latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/shadowing/monitor page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/shadowing/monitor.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/shadowing/monitor.adoc description: Learn how to monitor shadowing for disaster recovery. page-git-created-date: "2025-12-12" page-git-modified-date: "2026-05-26" --- Monitor your [shadow links](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/) to ensure proper replication performance and understand your disaster recovery readiness. Use `rpk` commands, metrics, and status information to track shadow link health and troubleshoot issues. > ❗ **IMPORTANT: Experiencing an active disaster?** > > See [Failover Runbook](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover-runbook/) for immediate step-by-step disaster procedures. ## [](#status-commands)Status commands To list existing shadow links: ### Cloud UI At the organization level of the Cloud UI, navigate to **Shadow Link**. ### rpk ```bash rpk shadow list ``` ### Control Plane API ```bash curl 'https://api.redpanda.com/v1/shadow-links' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` To view shadow link configuration details: ### Cloud UI 1. From the **Shadow Link** page, select the shadow link you want to view. 2. Click the **Tasks** tab to view all tasks and their status. ### rpk ```bash rpk shadow describe ``` For detailed command options, see [`rpk shadow list`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-list/) and [`rpk shadow describe`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-describe/). This command shows the complete configuration of the shadow link, including connection settings, filters, and synchronization options. ### Control Plane API ```bash curl 'https://api.redpanda.com/v1/shadow-links/' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` To check your shadow link status and ensure proper operation: ### Cloud UI 1. From the **Shadow Link** page, select the shadow link you want to view. 2. Click the **Tasks** tab to view all tasks and their status. ### rpk ```bash rpk shadow status ``` For troubleshooting specific issues, you can use command options to show individual status sections. See [`rpk shadow status`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-status/) for available status options. The status output includes the following: ### Cloud API ```bash # Get Data Plane API URL of shadow cluster export DATAPLANE_API_URL=`curl https://api.cloud.redpanda.com/v1/clusters/ \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" | jq .cluster.dataplane_api` curl "https://$DATAPLANE_API_URL/v1/shadowlinks/" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" # View topic state curl "https://$DATAPLANE_API_URL/v1/shadowlinks//topic" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` The status includes the following: - **Shadow link state**: Overall operational state (`ACTIVE`, `PAUSED`). - **Individual topic states**: Current state of each replicated topic (`ACTIVE`, `FAULTED`, `FAILING_OVER`, `FAILED_OVER`, `PAUSED`). - **Task status**: Health of replication tasks across brokers (`ACTIVE`, `FAULTED`, `NOT_RUNNING`, `LINK_UNAVAILABLE`). For details about shadow link tasks, see [Shadow link tasks](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#shadow-link-tasks). - **Failure reason**: When a shadow link or one of its tasks reports a failed state, the `Reason` field explains why. Check this field first when a link is not replicating as expected. - **Lag information**: Replication lag per partition showing source vs shadow high watermarks (HWM). - **Schema Registry sync status**: For links that replicate schemas through the Schema Registry API, inventory counts for source and destination subjects, sync progress, and the most recent error. See [Monitor replication status](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/#monitor-replication-status). ## [](#shadow-link-metrics)Metrics Shadowing provides comprehensive metrics to track replication performance and health with the [`public_metrics`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/) endpoint. | Metric | Type | Description | | --- | --- | --- | | redpanda_shadow_link_shadow_lag | Gauge | The lag of the shadow partition against the source partition, calculated as source partition LSO (Last Stable Offset) minus shadow partition HWM (High Watermark). Monitor by shadow_link_name, topic, and partition to understand replication lag for each partition. | | redpanda_shadow_link_total_bytes_fetched | Count | The total number of bytes fetched by a sharded replicator (bytes received by the client). Labeled by shadow_link_name and shard to track data transfer volume from the source cluster. | | redpanda_shadow_link_total_bytes_written | Count | The total number of bytes written by a sharded replicator (bytes written to the write_at_offset_stm). Uses shadow_link_name and shard labels to monitor data written to the shadow cluster. | | redpanda_shadow_link_client_errors | Count | The number of errors seen by the client. Track by shadow_link_name and shard to identify connection or protocol issues between clusters. | | redpanda_shadow_link_shadow_topic_state | Gauge | Number of shadow topics in the respective states. Labeled by shadow_link_name and state to monitor topic state distribution across your shadow links. | | redpanda_shadow_link_total_records_fetched | Count | The total number of records fetched by the sharded replicator (records received by the client). Monitor by shadow_link_name and shard to track message throughput from the source. | | redpanda_shadow_link_total_records_written | Count | The total number of records written by a sharded replicator (records written to the write_at_offset_stm). Uses shadow_link_name and shard labels to monitor message throughput to the shadow cluster. | See also: [Metrics Reference](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/) ## [](#monitoring-best-practices)Monitoring best practices ### [](#health-check-procedures)Health check procedures Establish regular monitoring workflows to ensure shadow link health: #### Cloud UI 1. From the **Shadow Link** page, select the shadow link you want to view. 2. Click the **Tasks** tab to view all tasks and their status. #### rpk ```bash # Check all shadow links are active rpk shadow list | grep -v "ACTIVE" || echo "All shadow links healthy" # Monitor lag for critical topics rpk shadow status | grep -E "LAG|Lag" ``` #### Cloud API ```bash # Check all shadow links are active curl 'https://api.redpanda.com/v1/shadow-links' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" | \ jq -r 'if all(.state == "SHADOW_LINK_STATE_ACTIVE") then "All shadow links healthy" else .[] | select(.state != "SHADOW_LINK_STATE_ACTIVE") end' # Monitor lag for critical topics curl "https://$DATAPLANE_API_URL/v1/shadowlinks//topic" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" ``` ### [](#alert-conditions)Alert conditions Configure monitoring alerts for the following conditions, which indicate problems with Shadowing: - **High replication lag**: When `redpanda_shadow_link_shadow_lag` exceeds your RPO requirements - **Connection errors**: When `redpanda_shadow_link_client_errors` increases rapidly - **Topic state changes**: When topics move to `FAULTED` state - **Task failures**: When replication tasks enter `FAULTED` or `NOT_RUNNING` states - **Throughput drops**: When bytes/records fetched drops significantly - **Link unavailability**: When tasks show `LINK_UNAVAILABLE` indicating source cluster connectivity issues For more information about shadow link tasks and their states, see [Shadow link tasks](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#shadow-link-tasks). --- # Page 399: Shadowing Overview **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview.md --- # Shadowing Overview > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Shadowing Overview latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/shadowing/overview page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/shadowing/overview.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/shadowing/overview.adoc description: Overview of shadowing for disaster recovery in Redpanda Cloud. page-git-created-date: "2025-12-12" page-git-modified-date: "2026-05-26" --- > 📝 **NOTE** > > Shadowing is supported on BYOC and Dedicated clusters running Redpanda version 25.3 and later. Shadowing is Redpanda’s enterprise-grade disaster recovery solution that establishes asynchronous, offset-preserving replication between two distinct Redpanda clusters. A cluster is able to create a dedicated client that continuously replicates source cluster data, including offsets, timestamps, and cluster metadata. This creates a read-only shadow cluster that you can quickly failover to handle production traffic during a disaster. Shadowing keeps data flowing, even during regional outages. > ❗ **IMPORTANT: Experiencing an active disaster?** > > See [Failover Runbook](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover-runbook/) for immediate step-by-step disaster procedures. Unlike traditional replication tools that re-produce messages, Shadowing copies data at the byte level, ensuring shadow topics contain identical copies of source topics with preserved offsets and timestamps. Shadowing replicates: - **Topic data**: All records with preserved offsets and timestamps - **Topic configurations**: Partition counts, retention policies, and other topic properties - **Consumer group offsets**: Enables seamless consumer resumption after failover - **Access control lists (ACLs)**: User permissions and security policies - **Schema Registry data**: Schema definitions, versions, and compatibility settings, replicated from another Redpanda cluster or from a Confluent Schema Registry ## [](#how-shadowing-fits-into-disaster-recovery)How Shadowing fits into disaster recovery Shadowing addresses enterprise disaster recovery requirements driven by regulatory compliance and business continuity needs. Organizations typically want to minimize both recovery time objective (RTO) and recovery point objective (RPO), and Shadowing asynchronous replication helps you achieve both goals by reducing data loss during regional outages and enabling rapid application recovery. The architecture follows an active-passive pattern. The source cluster processes all production traffic while the shadow cluster remains in read-only mode, continuously receiving updates. If a disaster occurs, you can failover the shadow topics, making them fully writable. At that point, you can redirect your applications to the shadow cluster, which becomes the new production cluster. > 📝 **NOTE** > > To avoid a split-brain scenario after failover, ensure that all clients are reconfigured to point to the shadow cluster before resuming write activity. Shadowing complements Redpanda’s existing availability and recovery capabilities. High availability actively protects your day-to-day operations, handling reads and writes seamlessly during node or availability zone failures within a region. Shadowing is your safety net for catastrophic regional disasters. Shadowing delivers near real-time, cross-region replication for mission-critical applications that require rapid failover with minimal data loss. ## [](#limitations)Limitations Shadowing for disaster recovery currently has the following limitations: - Shadowing is designed for active-passive disaster recovery scenarios. Each shadow cluster can maintain only one shadow link. - Shadowing operates exclusively in asynchronous mode and doesn’t support active-active configurations. This means there will always be some replication lag. - [Data transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/) are not supported on shadow clusters while Shadowing is active. Writing to shadow topics is blocked. - During a disaster, [audit log](https://docs.redpanda.com/cloud-data-platform/manage/audit-logging/) history from the source cluster is lost, though the shadow cluster begins generating new audit logs immediately after the failover. - After you failover shadow topics, automatic fallback to the original source cluster is not supported. ## [](#shadow-link-tasks)Shadow link tasks Shadow linking operates through specialized tasks that handle different aspects of replication. If you use a `shadow-config.yaml` configuration file to create the shadow link, each task corresponds to a section in the file. Tasks run continuously to maintain synchronization with the source cluster. #### Source Topic Sync The **Source Topic Sync task** manages topic discovery and metadata synchronization. This task periodically queries the source cluster to discover available topics, applies your configured topic filters to determine which topics should become shadow topics, and synchronizes topic properties between clusters. The task is controlled by the `topic_metadata_sync_options` section in the configuration file. It includes: - **Auto-creation filters**: Determines which source topics automatically become shadow topics - **Property synchronization**: Controls which topic properties replicate from source to shadow - **Starting offset**: Sets where new shadow topics begin replication (earliest, latest, or timestamp-based) - **Sync interval**: How frequently to check for new topics and property changes When this task discovers a new topic that matches your filters, it creates the corresponding shadow topic and begins replication from your configured starting offset. #### Consumer Group Shadowing The **Consumer Group Shadowing task** replicates consumer group offsets and membership information from the source cluster. This ensures that consumer applications can resume processing from the correct position after failover. The task is controlled by the `consumer_offset_sync_options` section in the configuration file. It includes: - **Group filters**: Determines which consumer groups have their offsets replicated - **Sync interval**: How frequently to synchronize consumer group offsets - **Offset clamping**: Automatically adjusts replicated offsets to valid ranges on the shadow cluster This task runs on brokers that host the `__consumer_offsets` topic and continuously tracks consumer group coordinators to optimize offset synchronization. #### Security Migrator The **Security Migrator task** replicates security policies, primarily ACLs (access control lists), from the source cluster to maintain consistent authorization across both environments. The task is controlled by the `security_sync_options` section in the configuration file. It includes: - **ACL filters**: Determines which security policies replicate - **Sync interval**: How frequently to synchronize security settings By default, all ACLs replicate to ensure your shadow cluster maintains the same security posture as your source cluster. #### Schema Registry Sync The **Schema Registry Sync task** replicates Schema Registry content so that applications that depend on schemas keep working after failover. The task is controlled by the `schema_registry_sync_options` section in the configuration file. It supports two modes: - **Topic mode** (`shadow_schema_registry_topic`): Shadows the `_schemas` system topic for byte-for-byte replication from another Redpanda cluster. - **API mode** (`shadow_schema_registry_api`): Polls the source Schema Registry over HTTP and imports selected contexts and subjects, with validation. Use this mode to replicate schemas from a Confluent Schema Registry. See [Migrate Schemas from Confluent Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). A shadow link uses one mode or the other, not both. Only API mode runs as a separate task that appears in the shadow link status. Topic mode adds the `_schemas` topic to the set of shadowed topics, so it is monitored like any other shadow topic rather than as a separate task. ### [](#task-status-and-monitoring)Task status and monitoring Each task reports its status through the shadow link status API. Task states include: - **`ACTIVE`**: Task is running normally and performing synchronization - **`PAUSED`**: Task has been manually paused through configuration - **`FAULTED`**: Task encountered an error and requires attention - **`NOT_RUNNING`**: Task is not currently executing - **`LINK_UNAVAILABLE`**: Task cannot communicate with the source cluster You can pause individual tasks by setting the `paused` field to `true` in the corresponding configuration section. This allows you to selectively disable parts of the replication process without affecting the entire shadow link. For monitoring task health and troubleshooting task issues, see [Monitor Shadowing](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/monitor/). ## [](#what-gets-replicated)What gets replicated Shadowing replicates your topic data with complete fidelity, preserving all message records with their original offsets, timestamps, headers, and metadata. The partition structure remains identical between source and shadow clusters, ensuring applications can resume processing from the exact same position after failover. Consumer group data flows according to your group filters, replicating offsets and membership information for matched groups. ACLs replicate based on your security filters. Schema Registry data synchronizes schema definitions, versions, and compatibility settings, either by shadowing the `_schemas` topic or through the [Schema Registry API](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). Partition count is always replicated to ensure the shadow topic matches the source topic’s partition structure. ### [](#topic-properties-replication)Topic properties replication The [Source Topic Sync task](#shadow-link-tasks) handles topic property replication. For topic properties, Redpanda follows these replication rules: **Never replicated** - `redpanda.remote.readreplica` - `redpanda.remote.recovery` - `redpanda.remote.allowgaps` - `redpanda.virtual.cluster.id` - `redpanda.leaders.preference` - `redpanda.cloud_topic.enabled` **Always replicated** - `max.message.bytes` - `cleanup.policy` - `message.timestamp.type` **Always replicated (unless `exclude_default` is `true`)** - `compression.type` - `retention.bytes` - `retention.ms` - `delete.retention.ms` - `replication.factor` - `min.compaction.lag.ms` - `max.compaction.lag.ms` To replicate additional topic properties, explicitly list them in `synced_shadow_topic_properties`. The filtering system you configure determines the precise scope of replication across all components, allowing you to balance comprehensive disaster recovery with operational efficiency. ## [](#best-practices)Best practices To ensure reliable disaster recovery with Shadowing: - **Do not modify shadow topic properties**: Avoid modifying synced topic properties on shadow topics, as these properties automatically revert to source topic values. ## [](#implementation-overview)Implementation overview Choose your implementation approach: - **[Setup and Configuration](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup/)**: Initial shadow configuration, authentication, and topic selection - **[Monitoring and Operations](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/monitor/)**: Health checks, lag monitoring, and operational procedures - **[Planned Failover](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover/)**: Controlled disaster recovery testing and migrations - **[Failover Runbook](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover-runbook/)**: Rapid disaster response procedures > 💡 **TIP** > > You can create and manage shadow links with the Redpanda Cloud UI, the [Cloud API](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview), or `rpk`, giving you flexibility in how you interact with your disaster recovery infrastructure. ## [](#next-steps)Next steps After setting up Shadowing for your Redpanda clusters, consider these additional steps: - **Test your disaster recovery procedures**: Regularly practice failover scenarios in a non-production environment. See [Failover Runbook](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/failover-runbook/) for step-by-step disaster procedures. - **Monitor shadow link health**: Set up alerting on the metrics described above to ensure early detection of replication issues. - **Implement automated failover**: Consider developing automation scripts that can detect outages and initiate failover based on predefined criteria. - **Review security policies**: Ensure your ACL filters replicate the appropriate security settings for your disaster recovery environment. - **Document your configuration**: Maintain up-to-date documentation of your shadow link configuration, including network settings, authentication details, and filter definitions. --- # Page 400: Configure Shadowing **URL**: https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/setup.md --- # Configure Shadowing > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Configure Shadowing latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: disaster-recovery/shadowing/setup page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: disaster-recovery/shadowing/setup.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/disaster-recovery/shadowing/setup.adoc description: Learn how to configure shadowing for disaster recovery. page-git-created-date: "2025-12-12" page-git-modified-date: "2026-05-26" --- You can create and manage shadow links with the Redpanda Cloud UI, the [Cloud API](https://docs.redpanda.com/api/doc/cloud-controlplane/topic/topic-cloud-api-overview), or `rpk`, giving you flexibility in how you interact with your disaster recovery infrastructure. > 💡 **TIP** > > Deploy clusters in different geographic regions to protect against regional disasters. ## [](#prerequisites)Prerequisites ### [](#license-and-cluster-requirements)License and cluster requirements Shadowing is supported on BYOC and Dedicated clusters running Redpanda version 25.3 and later. ### [](#cluster-configuration)Cluster configuration The shadow cluster must have the [`enable_shadow_linking`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#enable_shadow_linking) cluster property set to `true`. > 📝 **NOTE** > > Starting with Redpanda v25.3, this cluster property is enabled by default on new Redpanda Cloud clusters. For existing clusters on versions earlier than v25.3, you must enable this property manually. See [Configure Cluster Properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/). ### [](#replication-service-permissions)Replication service permissions You must configure a service account on the source cluster with the following [ACL](https://docs.redpanda.com/cloud-data-platform/security/authorization/acl/) permissions for shadow link replication: - **Topics**: `read` permission on all topics you want to replicate - **Topic configurations**: `describe_configs` permission on topics for configuration synchronization - **Consumer groups**: `describe` and `read` permission on consumer groups for offset replication - **ACLs**: `describe` permission on ACL resources to replicate security policies - **Cluster**: `describe` permission on the cluster resource to access ACLs This service account authenticates from the shadow cluster to the source cluster and performs the actual data replication. The credentials for this account are provided when you set up the shadow link. ### [](#network-and-authentication)Network and authentication You must configure network connectivity between clusters with appropriate firewall rules to allow the shadow cluster to connect to the source cluster for data replication. Shadowing uses a pull-based architecture where the shadow cluster fetches data from the source cluster. For detailed networking configuration, see [Networking](#networking). If using [authentication](https://docs.redpanda.com/cloud-data-platform/security/cloud-authentication/) for the shadow link connection, configure the source cluster with your chosen authentication method (SASL/SCRAM, SASL/PLAIN, TLS, mTLS) and ensure the shadow cluster has the proper credentials to authenticate to the source cluster. ## [](#set-up-shadowing)Set up Shadowing To set up Shadowing, you need to create a shadow link and configure filters to select which topics, consumer groups, ACLs, and Schema Registry data to replicate. If using the Cloud API to set up Shadowing, you must [authenticate](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication) to the API by including an access token in your requests. ### [](#create-a-shadow-link)Create a shadow link Any BYOC or Dedicated cluster can create a shadow link to a source cluster. > 💡 **TIP** > > You can use `rpk` to generate a sample configuration file with common filter patterns: > > ```bash > # Generate a sample configuration file with placeholder values > rpk shadow config generate --for-cloud -o shadow-config.yaml > ``` > > This creates a complete YAML configuration file that you can customize for your environment. The template includes all available fields with comments explaining their purpose. For detailed command options, see [`rpk shadow config generate --for-cloud`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-config-generate/). Explore the configuration file ```yaml # Sample ShadowLinkConfig YAML with all fields name: # Unique name for this shadow link, example: "production-dr" cloud_options: # Use either source_redpanda_id or bootstrap_servers: only one is required. source_redpanda_id: # Optional: 20 character lowercase ID of the cluster # Example: m7xtv2qq5njbhwruk88f shadow_redpanda_id: # 20 character lowercase ID of the cluster # Example: m7xtv2qq5njbhwruk88f client_options: bootstrap_servers: # Source cluster brokers to connect to - : # Example: "prod-kafka-1.example.com:9092" - : # Example: "prod-kafka-2.example.com:9092" - : # Example: "prod-kafka-3.example.com:9092" source_cluster_id: # Optional: UUID assigned by Redpanda # Example: a882bc98-7aca-40f6-a657-36a0b4daf1fd # This UUID is not available in Redpanda Cloud. # TLS settings using PEM strings tls_settings: enabled: true tls_pem_settings: ca: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- key: ${secrets.} cert: |- -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- # Create SASL credentials in the source cluster. # Then, with this configuration, ensure the shadow cluster uses the credentials # to authenticate to the source cluster. authentication_configuration: # SASL/SCRAM authentication scram_configuration: username: # SASL/SCRAM username, example: "shadow-replication-user" password: ${secrets.} # ID of secret containing SASL/SCRAM password scram_mechanism: SCRAM_SHA_256 # SCRAM mechanism: "SCRAM_SHA_256" or "SCRAM_SHA_512" # Connection tuning - adjust based on network characteristics metadata_max_age_ms: 10000 # How often to refresh cluster metadata (default: 10000ms) connection_timeout_ms: 1000 # Connection timeout (default: 1000ms, increase for high latency) retry_backoff_ms: 100 # Backoff between retries (default: 100ms) fetch_wait_max_ms: 500 # Max time to wait for fetch requests (default: 500ms) fetch_min_bytes: 5242880 # Min bytes per fetch (default: 5MB) fetch_max_bytes: 20971520 # Max bytes per fetch (default: 20MB) fetch_partition_max_bytes: 5242880 # Max bytes per partition fetch (default: 5MB) topic_metadata_sync_options: interval: 30s # How often to sync topic metadata (examples: "30s", "1m", "5m") auto_create_shadow_topic_filters: # Filters for automatic topic creation - pattern_type: LITERAL # Include all topics (wildcard) filter_type: INCLUDE name: '*' - pattern_type: PREFIX # Exclude topics with specific prefix filter_type: EXCLUDE name: # Examples: "temp-", "test-", "debug-" synced_shadow_topic_properties: # Additional topic properties to sync (beyond defaults) - retention.ms # Topic retention time - segment.ms # Segment roll time exclude_default: false # Include default properties (compression, retention, etc.) start_at_earliest: {} # Start from the beginning of source topics (default) paused: false # Enable topic metadata synchronization consumer_offset_sync_options: interval: 30s # How often to sync consumer group offsets paused: false # Enable consumer offset synchronization group_filters: # Filters for consumer groups to sync - pattern_type: LITERAL filter_type: INCLUDE name: '*' # Include all consumer groups security_sync_options: interval: 30s # How often to sync security settings paused: false # Enable security settings synchronization acl_filters: # Filters for ACLs to sync - resource_filter: resource_type: TOPIC # Resource type: "TOPIC", "GROUP", "CLUSTER" pattern_type: PREFIXED # Pattern type: "LITERAL", "PREFIXED" name: # Examples: "prod-", "app-data-" access_filter: principal: User: # Principal name, example: "User:app-service" operation: ANY # Operation: "READ", "WRITE", "CREATE", "DELETE", "ALTER", "DESCRIBE", "ANY" permission_type: ALLOW # Permission: "ALLOW" or "DENY" host: '*' # Host pattern, examples: "*", "10.0.0.0/8", "app-server.example.com" schema_registry_sync_options: # Schema Registry synchronization options shadow_schema_registry_topic: {} # Byte-for-byte _schemas replication (Redpanda source) # For a Confluent source, replace the line above with a # shadow_schema_registry_api block (API mode). See Migrate Schemas # from Confluent Schema Registry. ``` Because the shadow cluster pulls from the source cluster, the shadow cluster requires credentials to connect to the source cluster. And because you cannot store plaintext passwords in Redpanda Cloud, you must create a secret to hold the password for the user on the source cluster. If using mTLS, you must also create a secret to hold the key of the client certificate for the client to authenticate. Reference that secret in `client_options.tls_settings.key_file` in the configuration file. 1. In the shadow cluster, create the secret: #### Cloud UI In the shadow cluster, go to the **Secrets Store** page and create a secret for the source cluster user, scoped to Redpanda Cluster. If necessary, first create the user with all ACLs enabled in the source cluster. #### rpk In the shadow cluster, create a secret to store the authentication credential that the cluster will use (`"scram_configuration": "password"` in the example configuration in the next step). Your secret must be scoped to "Redpanda Cluster". Use [`rpk security secret create`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-security/rpk-security-secret-create/) to create the secret from the command line. #### Data Plane API In the shadow cluster, create a secret to store the authentication credential that the cluster will use (`"scram_configuration": "password"` in the example configuration in the next step). Your secret must be scoped to "Redpanda Cluster". Use the [Data Plane API](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-dataplane-api/) to programmatically create the secret. #### Terraform With Terraform, you define the secret and the shadow link together: the [`redpanda_secret` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/secret) in the next step’s Terraform tab creates this secret as part of the same configuration, so there is nothing to do in this step. 2. In the shadow cluster, create a shadow link to the source cluster. #### Cloud UI 1. At the organization level of the Cloud UI, navigate to **Shadow Link**. 2. Click **Create shadow link**. The Cloud UI guides you through the following steps. 3. On the **Connection** step: 1. Enter a unique name for the shadow link. The name must start and end with lowercase alphanumeric characters, hyphens allowed. 2. Select where the replicated data comes from: **Redpanda Cloud cluster** for an existing Redpanda Cloud cluster, or **Bootstrap URL** to connect to any Kafka-compatible cluster. 3. Configure TLS for the connection. TLS is enabled by default. Optionally, add a custom CA certificate or a mutual TLS (mTLS) client certificate. 4. Enter the authentication details for the source cluster (SASL/SCRAM or SASL/PLAIN), including the username and the name of the secret containing the password created in the previous step. 5. Optionally, expand **Advanced options** to configure client connection settings. 4. On the **Shadow** step, select the shadow cluster that receives the replicated data. 5. On the **Source** step, select the source cluster. This step appears only when the source is an existing Redpanda Cloud cluster. If you entered bootstrap servers, the wizard skips this step. 6. On the **Configuration** step, specify what to shadow from the source cluster: topics, ACLs, consumer offsets, and Schema Registry data. For a Confluent source, configure the Schema Registry section as described in [Migrate Schemas from Confluent Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). 7. Click **Create shadow link**. #### rpk 1. Run `rpk cloud login`. Select your shadow cluster when prompted. 2. To create a shadow link with the source cluster using `rpk`, run the following command from the shadow cluster: ```bash # When logged in, optionally create a new rpk profile to easily # switch to the shadow cluster rpk profile create --from-cloud shadow-cluster # Use the generated configuration file to create the shadow link rpk shadow create --config-file shadow-config.yaml ``` For detailed command options, see [`rpk shadow create`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-create/). > 💡 **TIP** > > Use [`rpk profile`](https://docs.redpanda.com/cloud-data-platform/manage/rpk/config-rpk-profile/) to save your cluster connection details and credentials for both source and shadow clusters. This allows you to easily switch between the two configurations. #### Control Plane API To create a shadow link using the Control Plane API, make a `POST /shadow-links` request from the shadow cluster: ```bash curl -X POST 'https://api.redpanda.com/v1/shadow-links' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d '{ "shadow_link": { "shadow_redpanda_id": "", "name": "", "client_options": { "bootstrap_servers": [":", ":", ":"], "tls_settings": { "enabled": true }, "authentication_configuration": { "scram_configuration": { "username": "", "password": "${secrets.}", "scram_mechanism": "SCRAM_MECHANISM_SCRAM_SHA_256" } } }, "topic_metadata_sync_options": { "interval": "30s", "auto_create_shadow_topic_filters": [ { "name": "*", "filter_type": "FILTER_TYPE_INCLUDE", "pattern_type": "PATTERN_TYPE_LITERAL" }, { "name": "", "filter_type": "FILTER_TYPE_EXCLUDE", "pattern_type": "PATTERN_TYPE_PREFIX" } ], "start_at_earliest": {}, "paused": false }, "consumer_offset_sync_options": { "paused": true }, "security_sync_options": { "paused": true } } }' ``` Replace the placeholders with your own values: - ``: ID of the shadow (destination) cluster. - ``: Unique name for this shadow link, for example, `production-dr`. - `:`, `: …​`: Source cluster brokers to connect to, for example, `prod-kafka-1.example.com:9092`, `prod-kafka-2.example.com:9092`. - ``: SASL/SCRAM username, for example, `shadow-replication-user`. You create this user in the source cluster. - ``: The name of the secret containing the SASL/SCRAM password from the source cluster. - ``: Exclude topics that use this prefix, for example, `temp-`, `test-`, `debug-`. The response object represents the [long-running operation](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/#lro) of creating a shadow link. To include Schema Registry replication in the request, add a `schema_registry_sync_options` object; for the API-mode request used to replicate from a Confluent Schema Registry, see [Migrate Schemas from Confluent Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). For the full API reference, see [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-shadowlinkservice_createshadowlink). #### Terraform Manage shadow links declaratively with the [`redpanda_shadow_link` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/shadow_link) in the [Redpanda Terraform provider](https://registry.terraform.io/providers/redpanda-data/redpanda/latest) (version 2.0.0 or later). The resource supports the full shadow link lifecycle: create, update, import, and destroy. Store the SASL/SCRAM password for the source cluster user in the shadow cluster’s secret store with the [`redpanda_secret` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/secret) (see [Manage cluster secrets](https://docs.redpanda.com/cloud-data-platform/manage/terraform-provider/#manage-cluster-secrets) for naming, write-only behavior, and rotation), then reference the secret from the shadow link configuration: ```hcl resource "redpanda_secret" "source_password" { name = "" secret_data = var.source_user_password secret_data_version = 1 # Increment when you rotate the password. scopes = ["SCOPE_REDPANDA_CLUSTER"] cluster_api_url = redpanda_cluster.shadow.cluster_api_url } resource "redpanda_shadow_link" "production_dr" { name = "" shadow_redpanda_id = "" source_redpanda_id = "" client_options = { tls_settings = { enabled = true } authentication_configuration = { scram_configuration = { scram_mechanism = "SCRAM_SHA_256" username = "" password = "$${secrets.${redpanda_secret.source_password.name}}" } } } } ``` Replace the placeholders with your own values: - ``: Name for the secret that stores the SASL/SCRAM password. Must be uppercase, matching `^[A-Z][A-Z0-9_]*$`, for example, `SOURCE_SCRAM_PASSWORD`. - ``: Unique name for this shadow link, for example, `production-dr`. - ``: ID of the shadow (destination) cluster. - ``: ID of the source cluster. For a source cluster that is not managed in Redpanda Cloud, omit `source_redpanda_id` and set `client_options.bootstrap_servers` instead. - ``: SASL/SCRAM username, for example, `shadow-replication-user`. You create this user in the source cluster, and you can manage it with the [`redpanda_user` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/user). Grant it the permissions listed in [Replication service permissions](#replication-service-permissions), which you can manage with the [`redpanda_acl` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/acl). To configure topic, consumer group, ACL, and Schema Registry synchronization, add the corresponding option blocks (`topic_metadata_sync_options`, `consumer_offset_sync_options`, `security_sync_options`, `schema_registry_sync_options`) to the resource. For the `shadow_schema_registry_api` block used to replicate from a Confluent Schema Registry, see [Migrate Schemas from Confluent Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). For the full schema and a complete working example, see the [`redpanda_shadow_link` reference](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/shadow_link) and the [provider example](https://github.com/redpanda-data/terraform-provider-redpanda/blob/main/examples/shadow_link/main.tf). > 📝 **NOTE** > > If you manage the shadow cluster with the [`redpanda_cluster` resource](https://registry.terraform.io/providers/redpanda-data/redpanda/latest/docs/resources/cluster), you can set the `enable_shadow_linking` cluster property in `cluster_configuration.custom_properties_json`. > > To protect against accidental deletion, Terraform refuses to destroy a shadow link unless the resource sets `allow_deletion = true`. ### [](#set-filters)Set filters Filters determine which resources Shadowing automatically creates when establishing your shadow link. Topic filters select which topics Shadowing automatically creates as shadow topics when they appear on the source cluster. After Shadowing creates a shadow topic, it continues replicating until you failover the topic, delete it, or delete the entire shadow link. Consumer group and ACL filters control which groups and security policies replicate to maintain application functionality. #### [](#filter-types-and-patterns)Filter types and patterns Each filter uses two key settings: - **Pattern type**: Determines how names are matched - `LITERAL`: Matches names exactly (including the special wildcard `*` to match all items) - `PREFIX`: Matches names that start with the specified string - **Filter type**: Specifies whether to INCLUDE or EXCLUDE matching items - `INCLUDE`: Replicate items that match the pattern - `EXCLUDE`: Skip items that match the pattern #### [](#filter-processing-rules)Filter processing rules Redpanda processes filters in the order you define them with EXCLUDE filters taking precedence. Design your filter lists carefully: 1. **Exclude filters win**: If any EXCLUDE filter matches a resource, it is excluded regardless of INCLUDE filters. 2. **Order matters for INCLUDE filters**: Among INCLUDE filters, the first match determines the result. 3. **Default behavior**: Items that don’t match any filter are excluded from replication. #### [](#common-filtering-patterns)Common filtering patterns Replicate all topics except test topics: ```yaml topic_metadata_sync_options: auto_create_shadow_topic_filters: - pattern_type: PREFIX filter_type: EXCLUDE name: test- # Exclude all test topics - pattern_type: LITERAL filter_type: INCLUDE name: '*' # Include all other topics ``` Replicate only production topics: ```yaml topic_metadata_sync_options: auto_create_shadow_topic_filters: - pattern_type: PREFIX filter_type: INCLUDE name: prod- # Include production topics - pattern_type: PREFIX filter_type: INCLUDE name: production- # Alternative production prefix ``` Replicate specific consumer groups: ```yaml consumer_offset_sync_options: group_filters: - pattern_type: LITERAL filter_type: INCLUDE name: critical-app-consumers # Include specific consumer group - pattern_type: PREFIX filter_type: INCLUDE name: prod-consumer- # Include production consumers ``` #### [](#schema-registry-synchronization)Schema Registry synchronization Shadowing can replicate Schema Registry data in one of two modes: - **Topic mode** (`shadow_schema_registry_topic`): Shadows the `_schemas` system topic for byte-for-byte replication of schema definitions, versions, and compatibility settings from another Redpanda cluster. - **API mode** (`shadow_schema_registry_api`): Polls the source Schema Registry over HTTP and imports selected contexts and subjects, with validation. Use this mode to replicate schemas from a Confluent Schema Registry, or to replicate only part of the source registry. See [Migrate Schemas from Confluent Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). A shadow link uses one mode or the other, not both. To enable topic mode, add the following to your shadow link configuration: ```yaml schema_registry_sync_options: shadow_schema_registry_topic: {} ``` To enable API mode instead, add a `shadow_schema_registry_api` block; its connection, filtering, and mapping options are described in [Migrate Schemas from Confluent Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/migrate-schemas-confluent/). Topic mode requirements: - The `_schemas` topic must exist on the source cluster - The `_schemas` topic must not exist on the shadow cluster, or must be empty - Once enabled, the `_schemas` topic will be replicated completely Important: After the `_schemas` topic becomes a shadow topic, it cannot be stopped without either failing over the topic or deleting it entirely. #### [](#system-topic-filtering-rules)System topic filtering rules Redpanda system topics have the following specific filtering restrictions: - Literal filters for `__consumer_offsets` and `_redpanda.audit_log` are rejected. - Prefix filters for topics starting with `_redpanda` or `__redpanda` are rejected. - Wildcard `*` filters will not match topics that start with `_redpanda` or `__redpanda`. - To shadow specific system topics, you must provide explicit literal filters for those individual topics. #### [](#acl-filtering)ACL filtering ACLs are replicated by the [Security Migrator task](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#shadow-link-tasks). This is recommended to ensure that your shadow cluster has the same permissions as your source cluster. To configure ACL filters: ```yaml security_sync_options: acl_filters: # Include read permissions for production topics - resource_filter: resource_type: TOPIC # Filter by topic resource pattern_type: PREFIXED # Match by prefix name: prod- # Production topic prefix access_filter: principal: User:app-user # Application service user operation: READ # Read operation permission_type: ALLOW # Allow permission host: '*' # Any host # Include consumer group permissions - resource_filter: resource_type: GROUP # Filter by consumer group pattern_type: LITERAL # Exact match name: '*' # All consumer groups access_filter: principal: User:app-user # Same application user operation: READ # Read operation permission_type: ALLOW # Allow permission host: '*' # Any host ``` #### [](#consumer-group-filtering-and-behavior)Consumer group filtering and behavior Consumer group filters determine which consumer groups have their offsets replicated to the shadow cluster by the [Consumer Group Shadowing task](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#shadow-link-tasks). Offset replication operates selectively within each consumer group. Only committed offsets for active shadow topics are synchronized, even if the consumer group has offsets for additional topics that aren’t being shadowed. For example, if consumer group "app-consumers" has committed offsets for "orders", "payments", and "inventory" topics, but only "orders" is an active shadow topic, then only the "orders" offsets will be replicated to the shadow cluster. ```yaml consumer_offset_sync_options: interval: 30s # How often to sync consumer group offsets paused: false # Enable consumer offset synchronization group_filters: - pattern_type: PREFIX filter_type: INCLUDE name: prod-consumer- # Include production consumer groups - pattern_type: LITERAL filter_type: EXCLUDE name: test-consumer-group # Exclude specific test groups ``` ##### [](#important-consumer-group-considerations)Important consumer group considerations **Avoid name conflicts:** If you plan to consume data from the shadow cluster, do not use the same consumer group names as those used on the source cluster. While this won’t break shadow linking, it can impact your RPO/RTO because conflicting group names may interfere with offset replication and consumer resumption during disaster recovery. **Offset clamping:** When Redpanda replicates consumer group offsets from the source cluster, offsets are automatically "clamped" during the commit process on the shadow cluster. If a committed offset from the source cluster is above the high watermark (HWM) of the corresponding shadow partition, Redpanda clamps the offset to the shadow partition’s HWM before committing it to the shadow cluster. This ensures offsets remain valid and prevents consumers from seeking beyond available data on the shadow cluster. #### [](#starting-offset-for-new-shadow-topics)Starting offset for new shadow topics When the [Source Topic Sync task](https://docs.redpanda.com/cloud-data-platform/manage/disaster-recovery/shadowing/overview/#shadow-link-tasks) creates a shadow topic for the first time, you can control where replication begins on the source topic. This setting only applies to empty shadow partitions and is crucial for disaster recovery planning. Changing this configuration only affects new shadow topics, existing shadow topics continue replicating from their current position. ```yaml topic_metadata_sync_options: start_at_earliest: {} ``` Alternatively, to start from the most recent offset: ```yaml topic_metadata_sync_options: start_at_latest: {} ``` Or to start from a specific timestamp: ```yaml topic_metadata_sync_options: start_at_timestamp: 2024-01-01T00:00:00Z ``` Starting offset options: - **`earliest`** (default): This replicates all existing data from the source topic. Use this for complete disaster recovery where you need full data history. - **`latest`**: This starts replication from the current end of the source topic, skipping existing data. Use this when you only need new data for disaster recovery and want to minimize initial replication time. - **`timestamp`**: This starts replication from the first record with a timestamp at or after the specified time. Use this for point-in-time disaster recovery scenarios. > ❗ **IMPORTANT** > > The starting offset only affects **new shadow topics**. After a shadow topic exists and has data, changing this setting has no effect on that topic’s replication. #### [](#networking)Networking Configure network connectivity between your source and shadow clusters to enable shadow link replication. The shadow cluster initiates connections to the source cluster using a pull-based architecture. For additional details about networking, see [Network and authentication](#network-and-authentication). ##### [](#connection-requirements)Connection requirements - **Direction**: Shadow cluster connects to source cluster (outbound from shadow, inbound to source) - **Protocol**: Kafka protocol over TCP (default port 9092, or your configured listener ports) - **Persistence**: Connections remain active for continuous replication ##### [](#firewall-configuration)Firewall configuration You must configure firewall rules to allow the shadow cluster to reach the source cluster. **On the source cluster network:** - Allow inbound TCP connections on Kafka listener ports (typically 9092). - Allow connections from the shadow cluster’s IP addresses or subnets. **On the shadow cluster network:** - Allow outbound TCP connections to the source cluster’s Kafka listener ports. - Ensure DNS resolution works for source cluster hostnames. ##### [](#bootstrap-servers)Bootstrap servers Specify multiple bootstrap servers in your shadow link configuration for high availability: ```yaml client_options: bootstrap_servers: # Source cluster brokers to connect to - : # Example: "prod-kafka-1.example.com:9092" - : # Example: "prod-kafka-2.example.com:9092" - : # Example: "prod-kafka-3.example.com:9092" ``` The shadow cluster uses these addresses to discover all brokers in the source cluster. If one bootstrap server is unavailable, the shadow cluster tries the next one in the list. ##### [](#network-security)Network security For production deployments, secure the network connection between clusters: TLS encryption: ```yaml client_options: tls_settings: enabled: true # Enable TLS tls_pem_settings: ca: |- # CA certificate in PEM format -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- key: ${secrets.} # Client private key (can use secrets reference) cert: |- # Optional: Client certificate in PEM format for mutual TLS -----BEGIN CERTIFICATE----- ... -----END CERTIFICATE----- do_not_set_sni_hostname: false # Optional: Skip SNI hostname when using TLS (default: false) ``` Authentication: ```yaml client_options: authentication_configuration: # SASL/SCRAM authentication. # Create SASL credentials in the source cluster. # Then, with this configuration, ensure the shadow cluster uses the credentials # to authenticate to the source cluster. scram_configuration: username: # SASL/SCRAM username, example: "shadow-replication-user" password: ${secrets.} # ID of secret containing SASL/SCRAM password scram_mechanism: SCRAM_SHA_256 # SCRAM mechanism: "SCRAM_SHA_256" or "SCRAM_SHA_512" ``` ##### [](#connection-tuning)Connection tuning Adjust connection parameters based on your network characteristics. For example: ```yaml client_options: # Connection and metadata settings connection_timeout_ms: 1000 # Default 1000ms, increase for high-latency networks retry_backoff_ms: 100 # Default 100ms, backoff between connection retries metadata_max_age_ms: 10000 # Default 10000ms, how often to refresh cluster metadata # Fetch request settings fetch_wait_max_ms: 500 # Default 500ms, max time to wait for fetch requests fetch_min_bytes: 5242880 # Default 5MB, minimum bytes to fetch per request fetch_max_bytes: 20971520 # Default 20MB, maximum bytes to fetch per request fetch_partition_max_bytes: 5242880 # Default 5MB, maximum bytes to fetch per partition ``` ## [](#update-an-existing-shadow-link)Update an existing shadow link To modify a shadow link configuration after creation, run: ### Cloud UI 1. At the organization level of the Cloud UI, navigate to **Shadow Link**. 2. Select the shadow link you want to modify, and click **Edit**. 3. Edit the shadow link settings or the shadowing behavior by specifying which content from the source cluster to shadow (topics, ACLs, consumer groups, Schema Registry). You can also enable additional topic properties to be shadowed or disable optional topic properties from being included in the shadowing. 4. Click **Save** to apply changes. ### rpk ```bash rpk shadow update ``` For detailed command options, see [`rpk shadow update`](https://docs.redpanda.com/cloud-data-platform/reference/rpk/rpk-shadow/rpk-shadow-update/). This opens your default editor to modify the shadow link configuration. Only changed fields are updated on the server. The shadow link name cannot be changed - you must delete and recreate the link to rename it. ### Control Plane API ```bash curl -X PATCH 'https://api.redpanda.com/v1/shadow-links/' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d '{ "security_sync_options": { "paused": false } }' ``` This endpoint returns a [long-running operation](https://docs.redpanda.com/cloud-data-platform/manage/api/cloud-byoc-controlplane-api/#lro). For the full API reference, see [Control Plane API reference](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-shadowlinkservice_updateshadowlink). ### Terraform Edit the `redpanda_shadow_link` resource configuration and run `terraform apply`. The provider sends only the changed fields to the API. The shadow link name and cluster IDs are immutable: if you change them, Terraform plans a replacement (destroy and recreate) instead of an update. The destroy step of a replacement succeeds only when the resource sets `allow_deletion = true`, the same guard described in the create step. --- # Page 401: Integrate Redpanda with Iceberg **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg.md --- # Integrate Redpanda with Iceberg > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Integrate Redpanda with Iceberg latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/index.adoc description: Generate Iceberg tables for your Redpanda topics for data lakehouse access. page-git-created-date: "2025-04-04" page-git-modified-date: "2025-07-30" --- - [About Iceberg Topics](about-iceberg-topics/) Learn how Redpanda can integrate topics with Apache Iceberg. - [Migrate to Iceberg Topics](migrate-to-iceberg-topics/) Migrate existing Iceberg integrations to Redpanda Iceberg topics. - [Specify Iceberg Schema](specify-iceberg-schema/) Learn about supported Iceberg modes and how you can integrate schemas with Iceberg topics. - [Use Iceberg Catalogs](use-iceberg-catalogs/) Learn how to access Redpanda topic data stored in Iceberg tables, using table metadata or a catalog integration. - [Integrate with REST Catalogs](rest-catalog/) Integrate Redpanda topics with managed Iceberg REST Catalogs. - [Query Iceberg Topics](query-iceberg-topics/) Query Redpanda topic data stored in Iceberg tables, based on the topic Iceberg mode and schema. - [Migrate Iceberg Catalogs](migrate-iceberg-catalog/) Switch the Iceberg catalog backend for an existing Redpanda cluster without losing untranslated topic data. - [Tune Performance for Iceberg Topics](iceberg-performance-tuning/) Optimize query performance and translation throughput for Iceberg topics with partitioning, compaction, lag target tuning, and cluster sizing guidance. - [Troubleshoot Iceberg Topics](iceberg-troubleshooting/) Diagnose and resolve errors in Redpanda Iceberg translation, including dead-letter queue (DLQ) inspection and record reprocessing. --- # Page 402: About Iceberg Topics **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics.md --- # About Iceberg Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: About Iceberg Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/about-iceberg-topics page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/about-iceberg-topics.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/about-iceberg-topics.adoc description: Learn how Redpanda can integrate topics with Apache Iceberg. page-git-created-date: "2025-04-04" page-git-modified-date: "2026-05-26" --- The Apache Iceberg integration for Redpanda allows you to store topic data in the cloud in the Iceberg open table format. This makes your streaming data immediately available in downstream analytical systems, including data warehouses like Snowflake, Databricks, ClickHouse, and Redshift, without setting up and maintaining additional ETL pipelines. You can also integrate your data directly into commonly-used big data processing frameworks, such as Apache Spark and Flink, standardizing and simplifying the consumption of streams as tables in a wide variety of data analytics pipelines. Redpanda supports [version 2](https://iceberg.apache.org/spec/#format-versioning) of the Iceberg table format. ## [](#iceberg-concepts)Iceberg concepts [Apache Iceberg](https://iceberg.apache.org) is an open source format specification for defining structured tables in a data lake. The table format lets you quickly and easily manage, query, and process huge amounts of structured and unstructured data. This is similar to the way you would manage and run SQL queries against relational data in a database or data warehouse. The open format lets you use many different languages, tools, and applications to process the same data in a consistent way, so you can avoid vendor lock-in. This data management system is also known as a _data lakehouse_. In the Iceberg specification, tables consist of the following layers: - **Data layer**: Stores the data in data files. The Iceberg integration currently supports the Parquet file format. Parquet files are column-based and suitable for analytical workloads at scale. They come with compression capabilities that optimize files for object storage. - **Metadata layer**: Stores table metadata separately from data files. The metadata layer allows multiple writers to stage metadata changes and apply updates atomically. It also supports database snapshots, and time travel queries that query the database at a previous point in time. - Manifest files: Track data files and contain metadata about these files, such as record count, partition membership, and file paths. - Manifest list: Tracks all the manifest files belonging to a table, including file paths and upper and lower bounds for partition fields. - Metadata file: Stores metadata about the table, including its schema, partition information, and snapshots. Whenever a change is made to the table, a new metadata file is created and becomes the latest version of the metadata in the catalog. For Iceberg-enabled topics, the manifest files are in JSON format. - **Catalog**: Contains the current metadata pointer for the table. Clients reading and writing data to the table see the same version of the current state of the table. The Iceberg integration supports two [catalog integration](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/) types. You can configure Redpanda to catalog files stored in the same object storage bucket or container where the Iceberg data files are located, or you can configure Redpanda to use an [Iceberg REST catalog](https://iceberg.apache.org/terms/#decoupling-using-the-rest-catalog) endpoint to update an externally-managed catalog when there are changes to the Iceberg data and metadata. ![Redpanda’s Iceberg integration](https://docs.redpanda.com/cloud-data-platform/shared/_images/iceberg-integration-optimized.png) When you enable the Iceberg integration for a Redpanda topic, Redpanda brokers store streaming data in the Iceberg-compatible format in Parquet files in object storage, in addition to the log segments uploaded using Tiered Storage. Storing the streaming data in Iceberg tables in the cloud allows you to derive real-time insights through many compatible data lakehouse, data engineering, and business intelligence [tools](https://iceberg.apache.org/vendors/). ## [](#prerequisites)Prerequisites To enable Iceberg for Redpanda topics, you must have the following: - A running [BYOC](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/) or BYOVPC cluster on Redpanda version 25.1 or later. The Iceberg integration is supported only for BYOC and BYOVPC, and the cluster properties to configure Iceberg are available with v25.1. - rpk: See [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). - Familiarity with the Redpanda Cloud API. You must [authenticate](https://docs.redpanda.com/api/doc/cloud-controlplane/authentication) to the Cloud API and use the Control Plane API to update your cluster configuration. ## [](#limitations)Limitations - It is not possible to append topic data to an existing Iceberg table that is not created by Redpanda. - If you enable the Iceberg integration on an existing Redpanda topic, Redpanda does not backfill the generated Iceberg table with topic data. - JSON schemas are supported starting with Redpanda version 25.2. ## [](#enable-iceberg-integration)Enable Iceberg integration To create an Iceberg table for a Redpanda topic, you must set the cluster configuration property `[iceberg_enabled](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_enabled)` to `true`, and also configure the topic property `redpanda.iceberg.mode`. You can choose to provide a schema if you need the Iceberg table to be structured with defined columns. 1. Set the `iceberg_enabled` configuration option on your cluster to `true`. When multiple clusters write to the same catalog, each cluster must use a distinct namespace to avoid table name collisions. This is especially critical for REST catalog providers that offer a single global catalog per account (such as AWS Glue), where there is no other isolation mechanism. By default, Redpanda creates Iceberg tables in a namespace called `redpanda`. To use a unique namespace for your cluster’s REST catalog integration, also set `[iceberg_default_catalog_namespace](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_default_catalog_namespace)` when you set `iceberg_enabled`. You cannot change this property after you enable Iceberg topics on the cluster. #### rpk ```bash rpk cloud login rpk profile create --from-cloud rpk cluster config set iceberg_enabled true # Optional: set a custom namespace (default is "redpanda") # rpk cluster config set iceberg_default_catalog_namespace '[""]' ``` #### Cloud API ```bash # Store your cluster ID in a variable export RP_CLUSTER_ID= # Retrieve a Redpanda Cloud access token export RP_CLOUD_TOKEN=`curl -X POST "https://auth.prd.cloud.redpanda.com/oauth/token" \ -H "content-type: application/x-www-form-urlencoded" \ -d "grant_type=client_credentials" \ -d "client_id=" \ -d "client_secret="` # Update cluster configuration to enable Iceberg topics # Optional: to set a custom namespace (default is "redpanda"), # add "iceberg_default_catalog_namespace":[""] to custom_properties curl -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" -X PATCH \ "https://api.cloud.redpanda.com/v1/clusters/${RP_CLUSTER_ID}" \ -H 'accept: application/json'\ -H 'content-type: application/json' \ -d '{"cluster_configuration":{"custom_properties": {"iceberg_enabled":true}}}' ``` The [`PATCH /clusters/{cluster.id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) request returns the ID of a long-running operation. The operation may take up to ten minutes to complete. You can check the status of the operation by polling the [`GET /operations/{id}`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-operationservice_getoperation) endpoint. 2. (Optional) Create a new topic. ```bash rpk topic create ``` ```bash TOPIC STATUS OK ``` 3. Configure `redpanda.iceberg.mode` for the topic. You can choose one of the following [Iceberg modes](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/): - `key_value`: Creates an Iceberg table using a simple schema, consisting of two columns, one for the record metadata including the key, and another binary column for the record’s value. - `value_schema_id_prefix`: Creates an Iceberg table whose structure matches the Redpanda schema for this topic, with columns corresponding to each field. You must register a schema in the Schema Registry (see next step), and producers must write to the topic using the Schema Registry wire format. - `value_schema_latest`: Creates an Iceberg table whose structure matches the latest schema registered for the subject in the Schema Registry. - `disabled` (default): Disables writing to an Iceberg table for this topic. ```bash rpk topic alter-config --set redpanda.iceberg.mode= ``` ```bash TOPIC STATUS OK ``` 4. Register a schema for the topic. This step is required for the `value_schema_id_prefix` and `value_schema_latest` modes. ```bash rpk registry schema create --schema --type ``` ```bash SUBJECT VERSION ID TYPE 1 1 PROTOBUF ``` ### [](#access-iceberg-data)Access Iceberg data To query the Iceberg table, you need access to the object storage bucket or container where the Iceberg data is stored. For BYOC clusters, the bucket name and table location are as follows: | Cloud provider | Bucket or container name | Iceberg table location | | --- | --- | --- | | AWS | redpanda-cloud-storage- | redpanda-iceberg-catalog/redpanda/ | | Azure | The Redpanda cluster ID is also used as the container name (ID) and the storage account ID. | | GCP | redpanda-cloud-storage- | For BYOVPC clusters, the bucket name is the name you chose when you created the object storage bucket as a customer-managed resource. For Azure clusters, you must add the public IP addresses or ranges from the REST catalog service, or other clients requiring access to the Iceberg data, to your cluster’s allow list. Alternatively, add subnet IDs to the allow list if the requests originate from the same Azure region. For example, to add subnet IDs to the allow list through the Control Plane API [`PATCH /v1/clusters/`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) endpoint, run: ```bash curl -X PATCH https://api.cloud.redpanda.com/v1/clusters/ \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d @- << EOF { "cloud_storage": { "azure": { "allowed_subnet_ids": [ ] } } } EOF ``` As you produce records to the topic, the data also becomes available in object storage for Iceberg-compatible clients to consume. You can use the same analytical tools to [read the Iceberg topic data](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/query-iceberg-topics/) in a data lake as you would for a relational database. See also: [Schema types translation](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/#schema-types-translation). ### [](#iceberg-data-retention)Iceberg data retention Data in an Iceberg-enabled topic is consumable from Kafka based on the configured [topic retention policy](https://docs.redpanda.com/cloud-data-platform/develop/topics/create-topic/). Conversely, data written to Iceberg remains queryable as Iceberg tables indefinitely. The Iceberg table persists unless you: - Delete the Redpanda topic associated with the Iceberg table. This is the default behavior set by the `[iceberg_delete](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_delete)` cluster property and the `redpanda.iceberg.delete` topic property. If you set this property to `false`, the Iceberg table remains even after you delete the topic. - Explicitly delete data from the Iceberg table using a query engine. - Disable the Iceberg integration for the topic and delete the Parquet files in object storage. The DLQ table (`~dlq`) follows the same persistence rules as the main Iceberg table. ## [](#schema-evolution)Schema evolution Redpanda supports schema evolution in accordance with the [Iceberg specification](https://iceberg.apache.org/spec/#schema-evolution). Permitted schema evolutions include reordering fields and promoting field types. When you update the schema in Schema Registry, Redpanda automatically updates the Iceberg table schema to match the new schema. For example, if you produce records to a topic `demo-topic` with the following Avro schema: schema\_1.avsc ```avro { "type": "record", "name": "ClickEvent", "fields": [ { "name": "user_id", "type": "int" }, { "name": "event_type", "type": "string" } ] } ``` ```bash rpk registry schema create demo-topic-value --schema schema_1.avsc echo '{"user_id":23, "event_type":"BUTTON_CLICK"}' | rpk topic produce demo-topic --format='%v\n' --schema-id=topic ``` Then, you update the schema to add a new field `ts`, and produce records with the updated schema: schema\_2.avsc ```avro { "type": "record", "name": "ClickEvent", "fields": [ { "name": "user_id", "type": "int" }, { "name": "event_type", "type": "string" }, { "name": "ts", "type": [ "null", { "type": "long", "logicalType": "timestamp-millis" } ], "default": null # Default value for the new field } ] } ``` The `ts` field can be either null or a long representing epoch milliseconds. The default value is null. ```bash rpk registry schema create demo-topic-value --schema schema_2.avsc echo '{"user_id":858, "event_type":"BUTTON_CLICK", "ts":1737998723230}' | rpk topic produce demo-topic --format='%v\n' --schema-id=topic ``` Querying the Iceberg table for `demo-topic` includes the new column `ts`: ```bash +---------+--------------+--------------------------+ | user_id | event_type | ts | +---------+--------------+--------------------------+ | 858 | BUTTON_CLICK | 2025-02-26T20:05:23.230Z | | 23 | BUTTON_CLICK | NULL | +---------+--------------+--------------------------+ ``` ## [](#next-steps)Next steps - [Use Iceberg Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/) - [Tune Performance for Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-performance-tuning/) - [Troubleshoot Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-troubleshooting/) ## [](#suggested-reading)Suggested reading - [Understanding Apache Kafka Schema Registry](https://www.redpanda.com/blog/schema-registry-kafka-streaming#how-does-serialization-work-with-schema-registry-in-kafka) --- # Page 403: Tune Performance for Iceberg Topics **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-performance-tuning.md --- # Tune Performance for Iceberg Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Tune Performance for Iceberg Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/iceberg-performance-tuning page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/iceberg-performance-tuning.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/iceberg-performance-tuning.adoc description: Optimize query performance and translation throughput for Iceberg topics with partitioning, compaction, lag target tuning, and cluster sizing guidance. page-git-created-date: "2026-05-06" page-git-modified-date: "2026-05-26" --- This guide covers strategies for optimizing the performance of Iceberg topics in Redpanda, including improving downstream query performance, tuning the Iceberg translation pipeline, and monitoring translation throughput. After reading this page, you will be able to: - Apply partitioning and compaction strategies to improve query performance - Choose appropriate lag target values for your workload - Identify translation performance signals using Iceberg metrics ## [](#prerequisites)Prerequisites You must be familiar with how Iceberg topics work in Redpanda. See [About Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/). ## [](#optimize-query-performance)Optimize query performance Query engines read Parquet files from object storage to process Iceberg table data. Partitioning, compaction, and schema design affect how efficiently those reads perform. ### [](#use-custom-partitioning)Use custom partitioning To improve query performance, consider implementing custom [partitioning](https://iceberg.apache.org/docs/nightly/partitioning/) for the Iceberg topic. Use the `redpanda.iceberg.partition.spec` topic property to define the partitioning scheme: ```bash # Create new topic with five topic partitions, replication factor 3, and custom table partitioning for Iceberg rpk topic create -p5 -r3 -c redpanda.iceberg.mode=value_schema_id_prefix -c "redpanda.iceberg.partition.spec=(, , ...)" ``` Valid `` values include a source column name or a transformation of a column. The columns referenced can be Redpanda-defined (such as `redpanda.timestamp`) or user-defined based on a schema that you register for the topic. The Iceberg table stores records that share different partition key values in separate files based on this specification. For example: - To partition the table by a single key, such as a column `col1`, use: `redpanda.iceberg.partition.spec=(col1)`. - To partition by multiple columns, use a comma-separated list: `redpanda.iceberg.partition.spec=(col1, col2)`. - To partition by the year of a timestamp column `ts1`, and a string column `col1`, use: `redpanda.iceberg.partition.spec=(year(ts1), col1)`. To learn more about how partitioning schemes can affect query performance, and for details on the partitioning specification such as allowed transforms, see the [Apache Iceberg documentation](https://iceberg.apache.org/spec/#partitioning). > 💡 **TIP** > > - Partition by columns that you frequently use in queries. Columns with relatively few unique values (low cardinality) are good candidates for partitioning. > > - If you must partition based on columns with high cardinality, for example timestamps, use Iceberg’s available transforms such as extracting the year, month, or day to avoid creating too many partitions. Too many partitions can be detrimental to performance because more files need to be scanned and managed. ### [](#compact-iceberg-tables)Compact Iceberg tables Over time, Iceberg translation can produce many small Parquet files, especially with low-throughput topics or short lag targets. Compaction merges small files into larger ones, reducing the number of metadata operations query engines must perform and improving read performance. - Automatic compaction: Some catalog and data platform services, such as AWS Glue and Databricks, automatically compact Iceberg tables. - Manual or scheduled compaction: Tools like [Apache Spark](https://spark.apache.org/) can run compaction jobs on a schedule. This is useful if your catalog or platform does not compact automatically. If you observe degraded read performance or a high number of small files, investigate whether your catalog or platform supports automatic compaction or schedule periodic compaction jobs. ### [](#avoid-high-column-count)Avoid high column count A high column count or schema field count results in more overhead when translating topics to the Iceberg table format. Small message sizes can also increase CPU utilization. To minimize the performance impact on your cluster, keep to a low column count and large message size for Iceberg topics. ## [](#tune-translation-performance)Tune translation performance Translation is the process in which Redpanda converts topic data into Parquet files for the Iceberg table. Each round of translation processes one topic partition at a time. Under typical conditions, Iceberg translation has the following performance characteristics: - Throughput: Approximately 5 MiB/s per core. - Flush threshold: 32 MiB. Each translation process uploads its on-disk data when accumulated data reaches this threshold. This is the primary control for Parquet file size, and is managed by Redpanda Cloud. - Lag target: Controlled by [`iceberg_target_lag_ms`](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_target_lag_ms) (default: 1 minute). Redpanda tries to commit all data produced to an Iceberg-enabled topic within this window. The flush threshold and lag target together determine the size of the Parquet files written to object storage. Larger Parquet files generally improve downstream query performance by reducing the number of metadata operations query engines must perform. ### [](#tune-the-lag-target)Tune the lag target In Redpanda Cloud, `datalake_translator_flush_bytes` is managed by Redpanda Cloud and is not user-tunable. To adjust the size of Parquet files written to object storage, increase the lag target. A larger lag target gives translators more time to accumulate data before committing, resulting in larger Parquet files with more records per file. You can configure the lag target at the cluster level or per-topic: - Cluster-wide: edit `iceberg_target_lag_ms` in the Redpanda Cloud Console. For instructions, see [Configure Cluster Properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/). - Per-topic: set the `redpanda.iceberg.target.lag.ms` topic property. The topic property overrides the cluster default for that topic. > 📝 **NOTE** > > Increasing the lag target means Iceberg tables receive new data less frequently. Choose a lag value that balances file efficiency against how current your downstream data must be. To check the current cluster-wide value: ```bash rpk cluster config get iceberg_target_lag_ms ``` To check topic-level overrides: ```bash rpk topic describe -c ``` ### [](#optimize-message-size)Optimize message size Redpanda has validated 32 MiB as the maximum recommended message size for Iceberg-enabled topics. With large messages, each Parquet file contains fewer records because the flush threshold is reached sooner. This can reduce the efficiency of analytical queries that need to scan many records. If query latency is a concern and your workload produces large messages, consider: - Reducing individual message sizes if your data model allows it. - Increasing `iceberg_target_lag_ms` to produce Parquet files with more records per file. See [Tune the lag target](#tune-the-lag-target). ### [](#size-clusters-for-iceberg-workloads)Size clusters for Iceberg workloads When you enable Iceberg for any substantial workload and start translating topic data to the Iceberg format, you may see most of your cluster’s CPU utilization increase. If this additional workload overwhelms the brokers and causes the Iceberg table lag to exceed the configured target lag, Redpanda automatically increases the scheduling priority of Iceberg translation to help it catch up with incoming data. However, this does not substitute for adequate cluster resources. You may need to increase the size of your Redpanda cluster to accommodate the additional workload. To ensure that your cluster is sized appropriately, contact the Redpanda Customer Success team. ### [](#monitor-translation-performance)Monitor translation performance Use the following [Iceberg metrics](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#iceberg-metrics) to understand whether translation is keeping pace with incoming data: - [`redpanda_iceberg_translation_raw_bytes_processed`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_translation_raw_bytes_processed): Total raw bytes consumed for translation input. Use this to monitor input throughput and compare against the expected 5 MiB/s per core baseline. - [`redpanda_iceberg_translation_parquet_bytes_added`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_translation_parquet_bytes_added): Total bytes written to Parquet files. Divide by `redpanda_iceberg_translation_files_created` to estimate the average file size produced by your workload. - [`redpanda_iceberg_translation_files_created`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_translation_files_created): Number of Parquet files created. A high file creation rate relative to bytes added indicates many small files. Consider increasing `iceberg_target_lag_ms`. - [`redpanda_iceberg_translation_parquet_rows_added`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_translation_parquet_rows_added): Total rows written to Parquet files. Useful for understanding record-level throughput. - [`redpanda_iceberg_translation_translations_finished`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_translation_translations_finished): Number of completed translator executions. A stalling or zero rate indicates translation has stopped. For metrics related to DLQ files, invalid records, and catalog commit failures, see [Troubleshooting metrics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-troubleshooting/#troubleshooting-metrics). > 💡 **TIP** > > If translation consistently lags despite available CPU headroom, the workload may be partition-bound. Each core translates its assigned partitions independently, so distributing data across more partitions allows more cores to contribute to translation and can improve total throughput. --- # Page 404: Query Iceberg Topics using AWS Glue **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-aws-glue.md --- # Query Iceberg Topics using AWS Glue > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Query Iceberg Topics using AWS Glue page-beta-text: This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/iceberg-topics-aws-glue page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/iceberg-topics-aws-glue.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/iceberg-topics-aws-glue.adoc description: Add Redpanda topics as Iceberg tables that you can access through the AWS Glue Data Catalog. # Beta release status page-beta: "true" page-git-created-date: "2025-08-05" page-git-modified-date: "2026-05-26" release-status: beta - This is a beta feature. Beta features are available for testing and feedback. They are not supported by Redpanda and should not be used in production environments. --- This guide walks you through querying Redpanda topics as Iceberg tables stored in AWS S3, using a catalog integration with [AWS Glue](https://docs.aws.amazon.com/glue/latest/dg/components-overview.html#data-catalog-intro). For general information about Iceberg catalog integrations in Redpanda, see [Use Iceberg Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/). ## [](#prerequisites)Prerequisites - An AWS account with access to [AWS Glue Data Catalog](https://docs.aws.amazon.com/glue/latest/dg/what-is-glue.html). - AWS Glue Data Catalog must be in the same AWS account and region as the cluster. - Redpanda version 25.2 or later. - [`rpk`](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) installed or updated to the latest version. - You can also use the Redpanda Cloud API to [reference secrets in your cluster configuration](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/#set-cluster-configuration-properties). - Admin permissions to create IAM policies and roles in AWS. ## [](#limitations)Limitations ### [](#lowercase-field-names-required)Lowercase field names required Use only lowercase field names. AWS Glue converts all table column names to lowercase, and Redpanda requires exact column name matches to manage schemas. Using uppercase letters prevents Redpanda from finding matching columns, which breaks schema management. ### [](#nested-partition-spec-support)Nested partition spec support AWS Glue does not support partitioning on nested fields. If Redpanda detects that the default partitioning `(hour(redpanda.timestamp))` based on the record metadata is in use, it will instead apply an empty partition spec `()`, which means the table will not be partitioned. To use partitioning, you must implement custom partitioning using your own partition columns (that is, columns that are not nested). > 📝 **NOTE** > > In Redpanda versions 25.2.1 and earlier, an empty partition spec `()` can cause a known issue that prevents certain engines like Amazon Redshift from successfully querying the table. To resolve this issue, specify custom partitioning, or upgrade Redpanda to versions 25.2.2 or later. ### [](#manual-deletion-of-iceberg-tables)Manual deletion of Iceberg tables The AWS Glue catalog integration does not support automatic deletion of Iceberg tables from Redpanda. To manually delete Iceberg tables in AWS Glue, you must either: - Set the cluster property `[iceberg_delete](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_delete)` to `false` when you configure the catalog integration. - Override the cluster property `iceberg_delete` by setting the topic property `redpanda.iceberg.delete` to `false` for the topic you want to delete. When `iceberg_delete` or the topic override `redpanda.iceberg.delete` is set to `false`, you can delete the Redpanda topic, and then delete the table in AWS Glue and the Iceberg data and metadata files in the S3 bucket. If you plan to re-create the topic after deleting it, you must delete the table data entirely before re-creating the topic. ## [](#authorize-access-to-aws-glue)Authorize access to AWS Glue For BYOC clusters created in March 2026 or later, the required AWS Glue IAM policy is automatically provisioned and attached to the cluster’s IAM role when Iceberg is enabled. You don’t need to manually create IAM policies or roles for Glue access. For clusters created before March 2026, you must re-run `rpk cloud byoc aws apply --redpanda-id=` to provision the Glue IAM policy before enabling Iceberg. This is a one-time operation that updates the cluster’s IAM role with the necessary Glue permissions. ## [](#configure-authentication-and-credentials)Configure authentication and credentials You can configure credentials for the AWS Glue Data Catalog integration in either of the following ways: - Allow Redpanda to use the same object storage credential properties already configured for S3. This is the recommended approach, especially in BYOC deployments where the cluster’s existing AWS credentials already include the necessary Glue permissions. For an example cluster configuration that uses the same IAM credentials for both S3 and AWS Glue, see the **Use cluster’s IAM credentials** tab in the [next section](#update-cluster-configuration). - If you want to configure authentication to AWS Glue separately from authentication to S3, there are equivalent credential configuration properties named `iceberg_rest_catalog_aws_*` that override the object storage credentials. These properties only apply to REST catalog authentication, and never to S3 authentication: - `[iceberg_rest_catalog_credentials_source](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_rest_catalog_credentials_source)` - Set the property to `sts` if you want to use the cluster’s default IAM role. - Set to `config_file` if you want to scope Glue access through your own IAM user and policy instead of the cluster’s default IAM role, or if you want to use static credentials. - `[iceberg_rest_catalog_aws_access_key](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_rest_catalog_aws_access_key)` (static credentials only) - `[iceberg_rest_catalog_aws_secret_key](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_rest_catalog_aws_secret_key)` (static credentials only), added as a secret value (see the [next section](#update-cluster-configuration) for details) - `[iceberg_rest_catalog_aws_region](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_rest_catalog_aws_region)` For an example cluster configuration that uses separate access keys for AWS Glue, see the **Use static credentials (override IAM)** tab in the [next section](#update-cluster-configuration). ## [](#update-cluster-configuration)Update cluster configuration To configure your Redpanda cluster to enable Iceberg on a topic and integrate with the AWS Glue Data Catalog: 1. Edit your cluster configuration to set the `iceberg_enabled` property to `true`, and set the catalog integration properties listed in the example below. By default, Redpanda creates Iceberg tables in a namespace called `redpanda`. Because AWS Glue provides a single catalog per account, each Redpanda cluster that writes to the same Glue catalog must use a distinct namespace to avoid table name collisions. To set a unique namespace, also set `[iceberg_default_catalog_namespace](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_default_catalog_namespace)` when you set `iceberg_enabled`. This property cannot be changed after Iceberg is enabled. Use `rpk` as shown in the following examples, or [use the Cloud API](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/#set-cluster-configuration-properties) to update these cluster properties. The update might take several minutes to complete. ### Use cluster’s IAM credentials ```bash # Glue requires Redpanda Iceberg tables to be manually deleted # so iceberg_delete is set to false. rpk cloud login rpk profile create --from-cloud rpk cluster config set \ iceberg_enabled=true \ iceberg_delete=false \ iceberg_default_catalog_namespace='[""]' \ iceberg_catalog_type=rest \ iceberg_rest_catalog_endpoint=https://glue..amazonaws.com/iceberg \ iceberg_rest_catalog_authentication_mode=aws_sigv4 \ iceberg_rest_catalog_credentials_source=sts \ iceberg_rest_catalog_aws_region= \ iceberg_rest_catalog_base_location=s3:/// ``` ### Use static credentials (override IAM) ```bash # Glue requires Redpanda Iceberg tables to be manually deleted # so iceberg_delete is set to false. rpk cluster config set \ iceberg_enabled=true \ iceberg_delete=false \ iceberg_default_catalog_namespace='[""]' \ iceberg_catalog_type=rest \ iceberg_rest_catalog_endpoint=https://glue..amazonaws.com/iceberg \ iceberg_rest_catalog_authentication_mode=aws_sigv4 \ iceberg_rest_catalog_credentials_source=config_file \ iceberg_rest_catalog_aws_region= \ iceberg_rest_catalog_aws_access_key= \ iceberg_rest_catalog_aws_secret_key='${secrets.}' \ iceberg_rest_catalog_base_location=s3:/// ``` Use your own values for the following placeholders: - ``: A unique namespace for this cluster’s Iceberg tables. Each Redpanda cluster that writes to the same Glue catalog must use a distinct namespace to avoid table name collisions. If omitted, the default namespace `redpanda` is used. - ``: The AWS region where your Data Catalog is located. The region in the AWS Glue endpoint must match the region specified in your `[iceberg_rest_catalog_aws_region](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_rest_catalog_aws_region)` property. - `` and ``: AWS Glue requires you to specify the base location where Redpanda stores Iceberg data and metadata files. You must use an S3 URI; for example, `s3:///iceberg`. - Bucket name: For BYOC clusters, the bucket name is `redpanda-cloud-storage-`. For BYOVPC clusters, use the name of the object storage bucket you created as a [customer-managed resource](https://docs.redpanda.com/cloud-data-platform/get-started/cluster-types/byoc/aws/vpc-byo-aws/#configure-the-redpanda-network-and-cluster). This must be the same bucket used for your cluster’s object storage. You cannot specify a different bucket for Iceberg data. - Warehouse: This is a name you choose as the logical name (such as `iceberg`) for the warehouse represented by all Redpanda Iceberg topic data in the cluster. As a security best practice, do not use the bucket root for the base location. Always specify a subfolder to avoid interfering with the rest of your cluster’s data in object storage. - `` (static credentials only): The AWS access key ID for your Glue service account. - `` (static credentials only): The name of the secret that stores the AWS secret access key for your Glue service account. To reference a secret in a cluster property, for example `iceberg_rest_catalog_aws_secret_key`, you must first [store the secret value](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/#store-a-secret-for-rest-catalog-authentication). ```bash Successfully updated configuration. New configuration version is 2. ``` 2. Enable the integration for a topic by configuring the topic property `redpanda.iceberg.mode`. The following examples show how to use [`rpk`](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) to either create a new topic or alter the configuration for an existing topic and set the Iceberg mode to `key_value`. The `key_value` mode creates a two-column Iceberg table for the topic, with one column for the record metadata including the key, and another binary column for the record’s value. See [Specify Iceberg Schema](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/) for more details on Iceberg modes. Create a new topic and set `redpanda.iceberg.mode`: ```bash rpk topic create --topic-config=redpanda.iceberg.mode=key_value ``` Set `redpanda.iceberg.mode` for an existing topic: ```bash rpk topic alter-config --set redpanda.iceberg.mode=key_value ``` 3. Produce to the topic. For example, ```bash echo "hello world\nfoo bar\nbaz qux" | rpk topic produce --format='%k %v\n' ``` You should see the topic as a table with data in AWS Glue Data Catalog. The data may take some time to become visible, depending on your `[iceberg_target_lag_ms](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_target_lag_ms)` setting. 1. In AWS Glue Studio, go to Databases. 2. Select the `redpanda` database. The `redpanda` database and the table within are automatically added for you. The table name is the same as the topic name. ## [](#query-iceberg-table)Query Iceberg table You can query the Iceberg table using different engines, such as Amazon Athena, PyIceberg, or Apache Spark. To query the table or view the table data in AWS Glue, ensure that your account has the necessary permissions to access the catalog, database, and table. To query the table in Amazon Athena: 1. On the list of tables in AWS Glue Studio, click "Table data" under the **View data** column. 2. Click "Proceed" to be redirected to the Athena query editor. 3. In the query editor, select AwsDataCatalog as the data source, and select the `redpanda` database. 4. The SQL query editor should be pre-populated with a query that selects 10 rows from the Iceberg table. Run the query to see a preview of the table data. ```sql SELECT * FROM "AwsDataCatalog"."redpanda"."" limit 10; ``` Your query results should look like the following: ```sql +-----------------------------------------------------+----------------+ | redpanda | value | +-----------------------------------------------------+----------------+ | {partition=0, offset=0, timestamp=2025-07-21 | 77 6f 72 6c 64 | | 18:11:25.070000, headers=null, key=[B@1900af31} | | +-----------------------------------------------------+----------------+ ``` ### [](#manage-access-for-query-engine-users)Manage access for query engine users Redpanda manages the permissions between Redpanda and the AWS Glue Data Catalog. To grant your end users and query engines (such as Amazon Athena or Apache Spark) read access to the Iceberg tables, use [AWS Lake Formation](https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html) to assign table-level and column-level permissions. ## [](#suggested-reading)Suggested reading - [Query Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/query-iceberg-topics/) --- # Page 405: Query Iceberg Topics using Databricks and Unity Catalog **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-databricks-unity.md --- # Query Iceberg Topics using Databricks and Unity Catalog > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Query Iceberg Topics using Databricks and Unity Catalog latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/iceberg-topics-databricks-unity page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/iceberg-topics-databricks-unity.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/iceberg-topics-databricks-unity.adoc description: Add Redpanda topics as Iceberg tables that you can query in Databricks managed by Unity Catalog. page-git-created-date: "2025-06-12" page-git-modified-date: "2026-05-26" --- This guide walks you through querying Redpanda topics as managed Iceberg tables in Databricks, with AWS S3 as object storage and a catalog integration using [Unity Catalog](https://docs.databricks.com/aws/en/data-governance/unity-catalog). For general information about Iceberg catalog integrations in Redpanda, see [Use Iceberg Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/). After reading this page, you will be able to: - Configure a Unity Catalog integration for Redpanda Iceberg topics with AWS S3 - Query Redpanda topic data as Iceberg tables in Databricks SQL ## [](#prerequisites)Prerequisites - A Databricks workspace in the same region as your S3 bucket. See the [list of supported AWS regions](https://docs.databricks.com/aws/en/resources/supported-regions#supported-regions-list). - Unity Catalog enabled in your Databricks workspace. See the [Databricks documentation](https://docs.databricks.com/aws/en/data-governance/unity-catalog/get-started) to set up Unity Catalog for your workspace. - [Predictive optimization](https://docs.databricks.com/aws/en/optimizations/predictive-optimization#enable-predictive-optimization) enabled for Unity Catalog. > 📝 **NOTE** > > When you enable predictive optimization, you must also set the following configurations in your Databricks workspace. These configurations allow predictive optimization to automatically generate column statistics and carry out background compaction for Iceberg tables: > > ```sql > SET spark.databricks.delta.liquid.lazyClustering.backfillStats=true; > SET spark.databricks.delta.computeStats.autoConflictResolution=true; > > /* > After setting these configurations, you can optionally run OPTIMIZE to > immediately trigger compaction and liquid clustering, or let predictive > optimization handle it automatically later. > */ > OPTIMIZE ``.redpanda.``; > ``` - [External data access](https://docs.databricks.com/aws/en/external-access/admin) enabled in your metastore. - Workspace admin privileges to complete the steps to create a Unity Catalog storage credential and external location that connects your cluster’s object storage bucket to Databricks. ## [](#limitations)Limitations The following data types are not currently supported for managed Iceberg tables: | Iceberg type | Equivalent Avro type | | --- | --- | | uuid | uuid | | fixed(L) | fixed | | time | time-millis, time-micros | There are no limitations for Protobuf types. ## [](#create-a-unity-catalog-storage-credential)Create a Unity Catalog storage credential A storage credential is a Databricks object that controls access to external object storage, in this case S3. You associate a storage credential with an AWS IAM role that defines what actions Unity Catalog can perform in the S3 bucket. Follow the steps in the [Databricks documentation](https://docs.databricks.com/aws/en/connect/unity-catalog/cloud-storage/storage-credentials) to create an AWS IAM role that has the required permissions for the bucket. When you have completed these steps, you should have the following configured in AWS and Databricks: - A self-assuming IAM role, meaning you’ve defined the role trust policy so the role trusts itself. - Two IAM policies attached to the IAM role. The first policy grants Unity Catalog read and write access to the bucket. The second policy allows Unity Catalog to configure file events. - A storage credential in Databricks associated with the IAM role, using the role’s ARN. You also use the storage credential’s external ID in the role’s trust relationship policy to make the role self-assuming. ## [](#create-a-unity-catalog-external-location)Create a Unity Catalog external location The external location stores the Unity Catalog-managed Iceberg metadata, and the Iceberg data written by Redpanda. You must use the same bucket configured for object storage for your Redpanda cluster. For BYOC clusters, the bucket name is `redpanda-cloud-storage-`, where `` is the ID of your Redpanda cluster. For BYOVPC clusters, the bucket name is the name you chose when you created the object storage bucket as a customer-managed resource. Follow the steps in the [Databricks documentation](https://docs.databricks.com/aws/en/connect/unity-catalog/cloud-storage/external-locations) to **manually** create an external location. You can create the external location in the Catalog Explorer or with SQL. You must create the external location manually because the location needs to be associated with the existing object storage bucket URL, `s3://`. ## [](#choose-a-catalog-setup)Choose a catalog setup You can either create a new catalog dedicated to Redpanda topics or use an existing catalog. If you create a new catalog, Redpanda automatically creates the required schema for you. If you need to integrate with an existing catalog, you must manually create the schema in that catalog before Redpanda creates any Iceberg tables. After you set up your catalog, the authorization and Redpanda configuration steps are the same for both options. ### [](#option-1-create-a-new-catalog-recommended)Option 1: Create a new catalog (recommended) Follow the steps in the Databricks documentation to [create a standard catalog](https://docs.databricks.com/aws/en/catalogs/create-catalog). When you create the catalog, specify the external location you created in the previous step as the storage location. In this setup, Redpanda creates the default `redpanda` schema for you. You use the catalog name when you set the Iceberg cluster configuration properties in Redpanda in a later step. ### [](#option-2-use-an-existing-catalog-with-a-pre-created-schema)Option 2: Use an existing catalog with a pre-created schema If you need to integrate Redpanda with an existing Unity Catalog catalog object, follow the steps to [create a schema](https://docs.databricks.com/aws/en/schemas/create-schema) in the catalog. - By default, Redpanda creates tables in a schema named `redpanda`. If you want to use a different schema, set `[iceberg_default_catalog_namespace](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_default_catalog_namespace)` before enabling Iceberg, then manually create that schema in the catalog. - Set the schema’s managed storage location to the same S3 bucket used for object storage, using the external location you created in the previous step. Unity Catalog resolves managed storage locations through a hierarchy of metastore > catalog > schema. If you assign the schema its own managed storage location, Redpanda can use the existing catalog while the schema stores its managed Iceberg data in the schema-specific location. For example: - Your existing Unity Catalog catalog stores managed data in `s3://`. - You manually create a `redpanda` schema in that catalog and override its managed storage location, through the external location, to the S3 bucket that Redpanda uses for your cluster’s object storage (`s3://redpanda-cloud-storage-` for BYOC, or your customer-managed bucket for BYOVPC). For more information, see the [Unity Catalog managed storage location hierarchy](https://docs.databricks.com/aws/en/data-governance/unity-catalog/#managed-storage-location-hierarchy) in the Databricks documentation. ## [](#authorize-access-to-unity-catalog)Authorize access to Unity Catalog Redpanda recommends using OAuth for service principals to grant Redpanda access to Unity Catalog. 1. Follow the steps in the [Databricks documentation](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m) to create a service principal, and then generate an OAuth secret. You use the client ID and secret to set Iceberg cluster configuration properties in Redpanda in the next step. 2. Open your catalog in the Catalog Explorer, then click **Permissions**. 3. Click **Grant** to grant the service principal the following permissions on the catalog: - `ALL PRIVILEGES` - `EXTERNAL USE SCHEMA` The Iceberg integration for Redpanda also supports using bearer tokens. ## [](#update-cluster-configuration)Update cluster configuration To configure your Redpanda cluster to enable Iceberg on a topic and integrate with Unity Catalog: 1. Edit your cluster configuration to set the `iceberg_enabled` property to `true`, and set the catalog integration properties listed in the example below. Use `rpk` like in the following example, or use the Cloud API to [update these cluster properties](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/#set-cluster-configuration-properties). The update might take several minutes to complete. To reference a secret in a cluster property, you must first [store the secret value](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/#store-a-secret-for-rest-catalog-authentication). ```bash rpk cloud login rpk profile create --from-cloud rpk cluster config set \ iceberg_enabled=true \ iceberg_catalog_type=rest \ iceberg_rest_catalog_endpoint=https:///api/2.1/unity-catalog/iceberg-rest \ iceberg_rest_catalog_authentication_mode=oauth2 \ iceberg_rest_catalog_oauth2_server_uri=https:///oidc/v1/token \ iceberg_rest_catalog_oauth2_scope=all-apis \ iceberg_rest_catalog_client_id= \ iceberg_rest_catalog_client_secret='${secrets.}' \ iceberg_rest_catalog_warehouse= \ iceberg_disable_snapshot_tagging=true # Optional. Set a custom namespace only if you want to use a schema other than the default `redpanda` # iceberg_default_catalog_namespace='[""]' ``` Use your own values for the following placeholders: - ``: The URL of your [Databricks workspace instance](https://docs.databricks.com/aws/en/workspace/workspace-details#workspace-instance-names-urls-and-ids); for example, `cust-success.cloud.databricks.com`. - ``: The client ID of the service principal you created in an earlier step. - ``: The name of the client secret of the service principal you created in an earlier step. - ``: The name of your catalog in Unity Catalog. ```bash Successfully updated configuration. New configuration version is 2. ``` 2. Enable the integration for a topic by configuring the topic property `redpanda.iceberg.mode`. The following examples show how to use [`rpk`](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) to either create a new topic or alter the configuration for an existing topic and set the Iceberg mode to `key_value`. The `key_value` mode creates an Iceberg table for the topic consisting of two columns, one for the record metadata including the key, and another binary column for the record’s value. See [Specify Iceberg Schema](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/) for more details on Iceberg modes. Create a new topic and set `redpanda.iceberg.mode`: ```bash rpk topic create --topic-config=redpanda.iceberg.mode=key_value ``` Set `redpanda.iceberg.mode` for an existing topic: ```bash rpk topic alter-config --set redpanda.iceberg.mode=key_value ``` 3. Produce to the topic. For example, ```bash echo "hello world\nfoo bar\nbaz qux" | rpk topic produce --format='%k %v\n' ``` You should see the topic as a table with data in Unity Catalog. The data may take some time to become visible, depending on your `[iceberg_target_lag_ms](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_target_lag_ms)` setting. 1. In Catalog Explorer, open your catalog. You should see a `redpanda` schema (or the namespace you configured with `[iceberg_default_catalog_namespace](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_default_catalog_namespace)`), in addition to `default` and `information_schema`. 2. The schema and the table residing within it are automatically added for you. The table name is the same as the topic name. ## [](#query-iceberg-table-using-databricks-sql)Query Iceberg table using Databricks SQL You can query the Iceberg table using different engines, such as Databricks SQL, PyIceberg, or Apache Spark. To query the table or view the table data in Catalog Explorer, ensure that your account has the necessary permissions to read the table. The following example shows how to query the Iceberg table using SQL in Databricks SQL. 1. In the Databricks console, open **SQL Editor**. 2. In the query editor, run: ```sql /* Ensure that the catalog and table name are correctly parsed in case they contain special characters. If you set iceberg_default_catalog_namespace to a custom namespace, replace `redpanda` with that namespace in the query below. */ SELECT * FROM ``.redpanda.`` LIMIT 10; ``` Your query results should look like the following: ```sql -- Example for redpanda.iceberg.mode=key_value with 1 record produced to topic +----------------------------------------------------------------------+------------+ | redpanda | value | +----------------------------------------------------------------------+------------+ | {"partition":0,"offset":"0","timestamp":"2025-04-02T18:57:11.127Z", | 776f726c64 | | "headers":null,"key":"68656c6c6f"} | | +----------------------------------------------------------------------+------------+ ``` ### [](#manage-access-for-query-engine-users)Manage access for query engine users Redpanda manages the permissions between Redpanda and Unity Catalog. To grant your end users or query engines read access to the Iceberg tables, use Unity Catalog to assign the appropriate privileges. Review the Databricks documentation on [granting permissions to objects](https://docs.databricks.com/aws/en/data-governance/unity-catalog/manage-privileges/?language=SQL#grant-permissions-on-objects-in-a-unity-catalog-metastore) and [Unity Catalog privileges](https://docs.databricks.com/aws/en/data-governance/unity-catalog/manage-privileges/privileges) for details. ## [](#suggested-reading)Suggested reading - [Query Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/query-iceberg-topics/) --- # Page 406: Use Iceberg Topics with GCP Lakehouse **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-gcp-biglake.md --- # Use Iceberg Topics with GCP Lakehouse > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Use Iceberg Topics with GCP Lakehouse latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/iceberg-topics-gcp-biglake page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/iceberg-topics-gcp-biglake.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/iceberg-topics-gcp-biglake.adoc description: Add Redpanda topics as Iceberg tables to Google Lakehouse for Apache Iceberg that you can query from Google BigQuery. page-git-created-date: "2026-06-11" page-git-modified-date: "2026-06-11" --- > 💡 **TIP** > > This guide is for integrating Iceberg topics with a managed REST catalog. Integrating with a REST catalog is recommended for production deployments. If it is not possible to use a REST catalog, you can use the [filesystem-based catalog](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/#object-storage). For an example of using the filesystem-based catalog to access Iceberg topics, see the [Getting Started with Iceberg Topics on Redpanda BYOC](https://www.redpanda.com/blog/iceberg-topics-redpanda-cloud-byoc-setup) blog post. This guide walks you through querying Redpanda topics as Iceberg tables stored in Google Cloud Storage, using a REST catalog integration with [Google Lakehouse for Apache Iceberg](https://docs.cloud.google.com/lakehouse/docs/introduction) (formerly BigLake). After completing this guide, you will be able to: - Create a catalog in GCP Lakehouse for Iceberg topic data. - Configure a Redpanda cluster to use GCP Lakehouse as an Iceberg REST catalog. - Query Iceberg topic data from Google BigQuery. For general information about Iceberg catalog integrations in Redpanda, see [Use Iceberg Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/). > 📝 **NOTE** > > Check the [Lakehouse product page](https://docs.cloud.google.com/lakehouse/docs) for the latest status and availability of the REST Catalog API. ## [](#prerequisites)Prerequisites - A Google Cloud Platform (GCP) project. - Lakehouse must be in the same GCP project as the cluster. Cross-project Lakehouse is not supported. If you do not have permissions to manage GCP resources such as VMs, storage buckets, and service accounts in your project, ask your project owner to create or update them for you. - The [`gcloud` CLI](https://docs.cloud.google.com/sdk/docs/install) installed and configured for your GCP project. - [Lakehouse (BigLake) API](https://cloud.google.com/biglake/docs/enable-biglake-api) enabled for your GCP project. - Redpanda version 25.3 or later. - `rpk` [installed or updated](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) to the latest version. - You can also use the Redpanda Cloud API to [reference secrets in your cluster configuration](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/#set-cluster-configuration-properties). > 📝 **NOTE** > > For BYOC clusters created before June 9, 2026, you must re-run `rpk cloud byoc gcp apply --redpanda-id= --project-id=` to enable the required API services before following this guide. This is a one-time operation. ## [](#limitations)Limitations ### [](#multi-region-bucket-support)Multi-region bucket support The Lakehouse runtime catalog does not support multi-region buckets. Use single-region buckets to store your Iceberg topics. ### [](#catalog-deletion)Catalog deletion Currently, it is not possible to delete non-empty Lakehouse Iceberg catalogs through the Lakehouse interface. If you need to reconfigure your setup, create a new bucket or use the REST API to remove the existing catalog. ### [](#topic-names)Topic names Lakehouse does not support Iceberg table names that contain dots (`.`). When creating Iceberg topics in Redpanda that you plan to access through Lakehouse, either: - Use the `iceberg_topic_name_dot_replacement` cluster property to set a replacement string for dots in topic names. Ensure that the replacement value does not cause table name collisions. For example, `current.orders` and `current_orders` would both map to the same table name if you set the replacement to an underscore (`_`). - Ensure that the new topic names do not include dots. You must also set the `iceberg_dlq_table_suffix` property to a value that does not include dots or tildes (`~`). See [Configure Redpanda for Iceberg](#configure-redpanda-for-iceberg) for the list of cluster properties to set when enabling the Lakehouse REST catalog integration. ## [](#set-up-google-cloud-resources)Set up Google Cloud resources For BYOC clusters, the required Lakehouse IAM permissions are automatically provisioned and attached to the cluster’s service account when Iceberg is enabled with a Lakehouse endpoint. You can skip to [Create a Lakehouse catalog](#create-a-lakehouse-catalog). For BYOVPC clusters, you must grant the required permissions to your cluster’s service account and enable the `biglake.googleapis.com` and `bigquery.googleapis.com` APIs in your GCP project. ### [](#grant-required-permissions)Grant required permissions Grant the necessary permissions to your service account. To run the following commands, replace the placeholder values: - ``: The name of your service account. - ``: The name of your storage bucket. 1. Grant the service account the [Storage Object Admin role](https://docs.cloud.google.com/storage/docs/access-control/iam-roles) to access the bucket: ```bash gcloud storage buckets add-iam-policy-binding gs:// \ --member="serviceAccount:@$(gcloud config get-value project).iam.gserviceaccount.com" \ --role="roles/storage.objectAdmin" ``` 2. Grant [Service Usage Consumer](https://docs.cloud.google.com/iam/docs/roles-permissions/serviceusage) and [BigLake Editor](https://docs.cloud.google.com/iam/docs/roles-permissions/biglake#biglake.editor) roles for using the Iceberg REST catalog: ```bash gcloud projects add-iam-policy-binding $(gcloud config get-value project) \ --member="serviceAccount:@$(gcloud config get-value project).iam.gserviceaccount.com" \ --role="roles/serviceusage.serviceUsageConsumer" gcloud projects add-iam-policy-binding $(gcloud config get-value project) \ --member="serviceAccount:@$(gcloud config get-value project).iam.gserviceaccount.com" \ --role="roles/biglake.editor" ``` ### [](#create-a-lakehouse-catalog)Create a Lakehouse catalog Create a Lakehouse Iceberg REST catalog using the [`gcloud biglake`](https://docs.cloud.google.com/sdk/gcloud/reference/biglake/iceberg/catalogs/create) command: ```bash gcloud biglake iceberg catalogs create --catalog-type=gcs-bucket --project= ``` Replace the placeholder values: - ``: Use the name of your storage bucket as the catalog ID. - ``: Your GCP project ID. ## [](#configure-redpanda-for-iceberg)Configure Redpanda for Iceberg 1. Edit your cluster configuration to set the `iceberg_enabled` property to `true`, and set the catalog integration properties listed in the example below. Use `rpk` as shown in the following example, or [use the Cloud API](https://docs.redpanda.com/cloud-data-platform/manage/cluster-maintenance/config-cluster/#set-cluster-configuration-properties) to update these cluster properties. The update might take several minutes to complete. ```bash rpk cloud login rpk profile create --from-cloud rpk cluster config set \ iceberg_enabled=true \ iceberg_catalog_type=rest \ iceberg_rest_catalog_endpoint=https://biglake.googleapis.com/iceberg/v1/restcatalog \ iceberg_rest_catalog_authentication_mode=gcp \ iceberg_rest_catalog_warehouse=gs:/// \ iceberg_rest_catalog_gcp_user_project= \ iceberg_dlq_table_suffix=_dlq ``` - ``: Your Redpanda cluster ID. - ``: For BYOC clusters, the bucket name is `redpanda-cloud-storage-`. For BYOVPC clusters, use the name of the object storage bucket you created as a customer-managed resource. - ``: Your GCP project ID. - You must set the `iceberg_dlq_table_suffix` property to a value that does not include dots or tildes (`~`). The example above uses `_dlq` as the suffix for the [dead-letter queue (DLQ) table](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-troubleshooting/#dead-letter-queue). 2. Enable the REST catalog integration for a topic by configuring the topic property `redpanda.iceberg.mode`. The following examples show how to use [`rpk`](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/) to either create a new topic or alter the configuration for an existing topic and set the Iceberg mode to `key_value`. The `key_value` mode creates a two-column Iceberg table for the topic, with one column for the record metadata including the key, and another binary column for the record’s value. See [Specify Iceberg Schema](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/) for more details on Iceberg modes. Create a new topic and set `redpanda.iceberg.mode`: ```bash rpk topic create --topic-config=redpanda.iceberg.mode=key_value ``` Set `redpanda.iceberg.mode` for an existing topic: ```bash rpk topic alter-config --set redpanda.iceberg.mode=key_value ``` Iceberg data can take a few moments to become available in Lakehouse. ## [](#query-iceberg-topics-in-bigquery)Query Iceberg topics in BigQuery 1. Navigate to the [BigQuery console](https://console.cloud.google.com/bigquery). 2. Query your Iceberg topic using SQL. For example, to query the `transactions` topic in the quickstart cluster: ```sql SELECT * FROM `>redpanda`.transactions ORDER BY redpanda.timestamp DESC LIMIT 10 ``` Replace `` with your bucket name. Your Redpanda topic is now available as Iceberg tables in Lakehouse, allowing you to run analytics queries directly on your streaming data. ### [](#manage-access-for-query-engine-users)Manage access for query engine users Redpanda manages the permissions between Redpanda and the BigLake catalog. To grant your end users and query engines read access to the Iceberg tables in BigQuery, see [Grant permissions for BigLake tables](https://cloud.google.com/bigquery/docs/manage-open-source-metadata#grant_permissions) in the Google Cloud documentation. ## [](#suggested-reading)Suggested reading - [Use Iceberg Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/) - [Query Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/query-iceberg-topics/) - [Google Lakehouse for Apache Iceberg documentation](https://docs.cloud.google.com/lakehouse/docs/introduction) --- # Page 407: Troubleshoot Iceberg Topics **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-troubleshooting.md --- # Troubleshoot Iceberg Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Troubleshoot Iceberg Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/iceberg-troubleshooting page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/iceberg-troubleshooting.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/iceberg-troubleshooting.adoc description: Diagnose and resolve errors in Redpanda Iceberg translation, including dead-letter queue (DLQ) inspection and record reprocessing. page-git-created-date: "2026-05-06" page-git-modified-date: "2026-05-26" --- Diagnose and resolve errors in Redpanda Iceberg translation, including dead-letter queue (DLQ) inspection and record reprocessing. Use this page to: - Diagnose Iceberg translation errors using DLQ tables and metrics - Reprocess or drop invalid records from the DLQ table ## [](#dead-letter-queue)Dead-letter queue If Redpanda encounters an error while writing a record to the Iceberg table, Redpanda by default writes the record to a separate DLQ Iceberg table named `~dlq`. The following can cause errors to occur when translating records in the `value_schema_id_prefix` and `value_schema_latest` modes to the Iceberg table format: - Redpanda cannot find the embedded schema ID in the Schema Registry. - Redpanda fails to translate one or more schema data types to an Iceberg type. - In `value_schema_id_prefix` mode, you do not use the Schema Registry wire format with the magic byte. The DLQ table itself uses the `key_value` schema, consisting of two columns: the record metadata including the key, and a binary column for the record’s value. > 📝 **NOTE** > > Topic property misconfiguration, such as [overriding the default behavior of `value_schema_latest` mode](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/#override-value-schema-latest-default) but not specifying the fully qualified Protobuf message name, does not cause records to be written to the DLQ table. Instead, Redpanda pauses the topic data translation to the Iceberg table until you fix the misconfiguration. ### [](#inspect-dlq-table)Inspect DLQ table You can inspect the DLQ table for records that failed to write to the Iceberg table, and you can take further action on these records, such as transforming and reprocessing them, or debugging issues that occurred upstream. The following example produces a record to a topic named `ClickEvent` and does not use the Schema Registry wire format that includes the magic byte and schema ID: ```bash echo '"key1" {"user_id":2324,"event_type":"BUTTON_CLICK","ts":"2024-11-25T20:23:59.380Z"}' | rpk topic produce ClickEvent --format='%k %v\n' ``` Querying the DLQ table returns the record that was not translated: ```sql SELECT value FROM ."ClickEvent~dlq"; -- Fully qualified table name ``` ```bash +-------------------------------------------------+ | value | +-------------------------------------------------+ | 7b 22 75 73 65 72 5f 69 64 22 3a 32 33 32 34 2c | | 22 65 76 65 6e 74 5f 74 79 70 65 22 3a 22 42 55 | | 54 54 4f 4e 5f 43 4c 49 43 4b 22 2c 22 74 73 22 | | 3a 22 32 30 32 34 2d 31 31 2d 32 35 54 32 30 3a | | 32 33 3a 35 39 2e 33 38 30 5a 22 7d | +-------------------------------------------------+ ``` The data is in binary format, and the first byte is not `0x00`, indicating that it was not produced with a schema. ### [](#reprocess-dlq-records)Reprocess DLQ records You can apply a transformation and reprocess the record in your data lakehouse to the original Iceberg table. In this case, you have a JSON value represented as a UTF-8 binary. Depending on your query engine, you might need to decode the binary value first before extracting the JSON fields. Some query engines decode the binary value automatically: ClickHouse SQL example to reprocess DLQ record ```sql SELECT CAST(jsonExtractString(json, 'user_id') AS Int32) AS user_id, jsonExtractString(json, 'event_type') AS event_type, jsonExtractString(json, 'ts') AS ts FROM ( SELECT CAST(value AS String) AS json FROM .`ClickEvent~dlq` -- Ensure that the table name is properly parsed ); ``` ```bash +---------+--------------+--------------------------+ | user_id | event_type | ts | +---------+--------------+--------------------------+ | 2324 | BUTTON_CLICK | 2024-11-25T20:23:59.380Z | +---------+--------------+--------------------------+ ``` You can now insert the transformed record back into the main Iceberg table. Redpanda recommends using an exactly-once processing strategy to avoid duplicates when reprocessing records. ### [](#drop-invalid-records)Drop invalid records To disable the default behavior and drop an invalid record, set the `redpanda.iceberg.invalid.record.action` topic property to `drop`. You can also configure the default cluster-wide behavior for invalid records by setting the `iceberg_invalid_record_action` property. ## [](#troubleshooting-metrics)Troubleshooting metrics The following [Iceberg metrics](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#iceberg-metrics) help identify translation errors, invalid records, and catalog connectivity issues: - [`redpanda_iceberg_translation_dlq_files_created`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_translation_dlq_files_created): Number of DLQ Parquet files created. A non-zero and increasing value indicates records are failing to translate. See [Inspect DLQ table](#inspect-dlq-table) to examine the failed records. - [`redpanda_iceberg_translation_invalid_records`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_translation_invalid_records): Number of invalid records encountered during translation, labeled by cause. See [Drop invalid records](#drop-invalid-records) to configure how Redpanda handles these records. - [`redpanda_iceberg_rest_client_num_commit_table_update_requests_failed`](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#redpanda_iceberg_rest_client_num_commit_table_update_requests_failed): Failed table commit requests to the REST catalog. Applies only when using a REST catalog (`iceberg_catalog_type: rest`). Persistent failures indicate catalog connectivity or permission issues. --- # Page 408: Migrate Iceberg Catalogs **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/migrate-iceberg-catalog.md --- # Migrate Iceberg Catalogs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Migrate Iceberg Catalogs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/migrate-iceberg-catalog page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/migrate-iceberg-catalog.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/migrate-iceberg-catalog.adoc description: Switch the Iceberg catalog backend for an existing Redpanda cluster without losing untranslated topic data. page-git-created-date: "2026-05-30" page-git-modified-date: "2026-05-30" --- Use the procedures in this topic when moving from the filesystem-based `object_storage` catalog to a managed REST catalog, or when changing between REST catalogs. By doing so, you can switch your cluster from one Iceberg catalog backend to another without losing untranslated topic data. This procedure guides you through pausing Iceberg translation per topic, draining pending commits to the old catalog, applying the new catalog configuration, and restarting the cluster. After completing these steps, you will be able to: - Verify that a target Iceberg catalog supports your existing schemas and partition specs - Pause Iceberg translation and drain pending commits without losing untranslated data - Apply new catalog configuration and resume translation safely > ❗ **IMPORTANT** > > Do not change `[iceberg_catalog_type](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_catalog_type)` or any other catalog cluster property in place without following this procedure. In-flight commits and untranslated data can be lost or stuck if the catalog changes mid-translation. ## [](#prerequisites)Prerequisites - Iceberg topics enabled and running on your Redpanda cluster. - Network connectivity from all brokers to the new catalog endpoint. - Credentials configured for the new catalog (REST endpoint, authentication mode, secret or token). For configuration guidance for each catalog type, see [Use Iceberg Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/). - The new catalog must support the current schema and partition spec of every Iceberg topic. See [Verify catalog compatibility](#verify-catalog-compatibility). ## [](#verify-catalog-compatibility)Verify catalog compatibility Before starting the migration, verify that the new catalog can host every Iceberg topic’s table with its existing schema and partition spec. If the new catalog is incompatible with your existing schemas or partition specs, already-translated Parquet files will fail to commit, and translation will stall in a state that is difficult to recover from. The simplest validation is to manually create a test table in the new catalog with the same schema and partition spec as one of your Iceberg topics. If the create call fails, fix the partition spec or schema before migrating. Delete the test table after validation. > ⚠️ **CAUTION** > > AWS Glue does not support partitioning on a nested field, which is Redpanda’s default partition spec for Iceberg topics. If you migrate to AWS Glue, you must change the partition spec to a Glue-compatible form before starting the migration procedure. ## [](#run-the-migration)Run the migration 1. Save the current `retention.ms` and `retention.bytes` values for every Iceberg topic, then set both to `-1` (infinite retention): ```bash rpk topic alter-config --set retention.ms=-1 --set retention.bytes=-1 ``` While Iceberg translation is paused in the next step, the topic’s retention anchor on the log is released. Without infinite retention, the cluster could delete untranslated data before the migration completes. 2. Pause Iceberg translation on every Iceberg topic by setting `redpanda.iceberg.mode` to `disabled`. Save each topic’s previous mode value so you can restore it later. ```bash rpk topic alter-config --set redpanda.iceberg.mode=disabled ``` Setting the mode to `disabled` stops new translation while letting already-translated data finish committing to the old catalog. For more about Iceberg modes, see [Specify Iceberg Schema](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/). > 📝 **NOTE** > > Do not change `[iceberg_enabled](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_enabled)` at the cluster level. The Iceberg integration must remain enabled at the cluster level so that pending commits can drain to the old catalog. 3. Wait for pending commits to drain. Monitor the `redpanda_iceberg_pending_commit_lag` metric until it reaches `0` for every Iceberg topic-partition. This metric reports the number of offsets pending a commit to the Iceberg catalog. While it is non-zero, Redpanda is still flushing translated data to the old catalog. A non-zero value here is expected while translation is paused, and reflects new records the cluster has not yet translated. > 💡 **TIP** > > If you scrape Prometheus, the following expression returns `0` only when every Iceberg-topic partition has fully drained: > > ```promql > sum(redpanda_iceberg_pending_commit_lag) > ``` 4. Apply the new catalog configuration. For example, to switch from `object_storage` to a REST catalog, update the catalog cluster properties: ```bash rpk cluster config set iceberg_catalog_type rest rpk cluster config set iceberg_rest_catalog_endpoint rpk cluster config set iceberg_rest_catalog_authentication_mode oauth2 # Set additional credential properties for your chosen authentication mode. ``` For full guidance on setting catalog cluster properties, see [Connect to a REST catalog](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/#rest) and the individual [REST catalog integration pages](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/rest-catalog/). 5. Restart Redpanda. The catalog cluster properties require a restart to take effect. Coordinate the restart with Redpanda Support. The restart must occur after `redpanda_iceberg_pending_commit_lag` has reached `0` and before you resume translation. 6. After the cluster comes up, check broker logs for successful catalog requests and the absence of authentication errors to verify the new catalog connection. 7. Resume Iceberg translation by restoring `redpanda.iceberg.mode` on every Iceberg topic to its previous value: ```bash rpk topic alter-config --set redpanda.iceberg.mode= ``` 8. Restore `retention.ms` and `retention.bytes` on every Iceberg topic to the values you saved before starting the migration. ## [](#verify-the-migration)Verify the migration After the migration completes, confirm that new data is reaching the new catalog: - Query an Iceberg table in your query engine using the new catalog and confirm that row counts continue to increase as your topic produces new records. - Check broker logs for any commit failures referencing the new catalog. Repeated failures often indicate a schema or partition spec mismatch. See [Troubleshooting](#troubleshooting) for details. ## [](#troubleshooting)Troubleshooting - Pending commits stuck after restart: A schema or partition spec mismatch between the original tables and the new catalog is the most common cause. See [Verify catalog compatibility](#verify-catalog-compatibility). If you cannot resolve the mismatch, contact [Redpanda Support](https://support.redpanda.com/hc/en-us/requests/new). - Authentication errors against the new REST catalog: Verify that the credential cluster properties (for example, `iceberg_rest_catalog_client_id`, `iceberg_rest_catalog_client_secret`, `iceberg_rest_catalog_token`) match what the new catalog expects. For OAuth, also check `iceberg_rest_catalog_oauth2_server_uri`. - Translation does not resume after restoring `redpanda.iceberg.mode`: Check that `redpanda_iceberg_pending_translation_lag` is increasing as new records are produced. If it remains `0`, the cluster is not translating new records. Verify that your producer is still writing to the topic and that the topic’s mode value is one of `key_value`, `value_schema_id_prefix`, or `value_schema_latest`. --- # Page 409: Migrate to Iceberg Topics **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/migrate-to-iceberg-topics.md --- # Migrate to Iceberg Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Migrate to Iceberg Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/migrate-to-iceberg-topics page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/migrate-to-iceberg-topics.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/migrate-to-iceberg-topics.adoc description: Migrate existing Iceberg integrations to Redpanda Iceberg topics. page-topic-type: how-to learning-objective-1: Compare external Iceberg integrations with Iceberg Topics architectures learning-objective-2: Implement data merge strategies using SQL patterns learning-objective-3: Execute validation checks and perform cutover procedures page-git-created-date: "2026-02-28" page-git-modified-date: "2026-05-26" --- Migrate existing Iceberg pipelines to Redpanda Iceberg topics to simplify your architecture and reduce operational overhead. After reading this page, you will be able to: - Compare external Iceberg integrations with Iceberg Topics architectures - Implement data merge strategies using SQL patterns - Execute validation checks and perform cutover procedures ## [](#why-migrate-to-iceberg-topics)Why migrate to Iceberg Topics Redpanda’s built-in Iceberg-enabled topics offer a simpler alternative to external Iceberg integrations for writing streaming data to Iceberg tables. > 📝 **NOTE** > > This page focuses on migrating from Kafka Connect Iceberg Sink. The migration patterns and SQL examples can be adapted for other Iceberg sources such as Apache Flink or Spark. ### [](#kafka-connect-iceberg-sink-comparison)Kafka Connect Iceberg Sink comparison The following table compares Kafka Connect Iceberg Sink with Redpanda Iceberg Topics: | Aspect | Kafka Connect Iceberg Sink | Iceberg Topics | | --- | --- | --- | | Infrastructure | Requires external Kafka Connect cluster | Built into Redpanda brokers | | Dependencies | Separate service to manage | No external dependencies | | Setup time | Medium (deploy connector) | Fast (enable topic property and post schema) | ## [](#prerequisites)Prerequisites To migrate from an existing Iceberg integration to Iceberg Topics, you must have: - [Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/) enabled on your Redpanda cluster. - Understanding of your current schema format (Avro, Protobuf, or JSON Schema). - For Kafka Connect migrations, knowledge of your Kafka Connect configuration, especially if using `iceberg.tables.route-field` for multi-table routing. - If migrating multi-table fan-out patterns, [data transforms](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/how-transforms-work/) enabled on your cluster. - Access to both source and target (Iceberg Topics) tables in your query engine. - Query engine access (Snowflake, Databricks, ClickHouse, or Spark) for data merging. ## [](#migration-steps)Migration steps Redpanda recommends following a phased approach to ensure data consistency and minimize risk: 1. Enable Iceberg on target topics and verify new data flows. 2. Run both systems concurrently during transition. 3. Choose a strategy to combine historical and new data. 4. Verify data completeness and accuracy. 5. Disable the external Iceberg integration. > ❗ **IMPORTANT** > > Iceberg Topics cannot append to existing Iceberg tables that are not created by Redpanda. You must create new Iceberg tables and merge historical data separately. ### [](#enable-iceberg-topics)Enable Iceberg Topics For simple migrations (one topic mapping to one Iceberg table), enable the Iceberg integration for your Redpanda topics. 1. Set the `iceberg_enabled` configuration property on your cluster to `true`: ###### rpk ```bash rpk cloud login rpk profile create --from-cloud rpk cluster config set iceberg_enabled true ``` ###### Cloud API ```bash # Store your cluster ID in a variable export RP_CLUSTER_ID= # Retrieve a Redpanda Cloud access token export RP_CLOUD_TOKEN=$(curl -X POST "https://auth.prd.cloud.redpanda.com/oauth/token" \ -H "content-type: application/x-www-form-urlencoded" \ -d "grant_type=client_credentials" \ -d "client_id=" \ -d "client_secret=") # Update cluster configuration to enable Iceberg topics curl -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" -X PATCH \ "https://api.cloud.redpanda.com/v1/clusters/${RP_CLUSTER_ID}" \ -H 'accept: application/json' \ -H 'content-type: application/json' \ -d '{"cluster_configuration":{"custom_properties": {"iceberg_enabled":true}}}' ``` 2. Configure the `redpanda.iceberg.mode` property for the topic: ```bash rpk topic alter-config --set redpanda.iceberg.mode= ``` Choose the mode based on your message format and schema configuration. For Kafka Connect migrations, use this mapping: | Kafka Connect Converter | Recommended Iceberg Mode | | --- | --- | | io.confluent.connect.avro.AvroConverter | value_schema_id_prefix (messages already use Schema Registry wire format) | | io.confluent.connect.protobuf.ProtobufConverter | value_schema_id_prefix (messages already use Schema Registry wire format) | | org.apache.kafka.connect.json.JsonConverter with schemas | value_schema_latest (Schema Registry resolves schema automatically) | | org.apache.kafka.connect.json.JsonConverter with embedded schemas | key_value (schema included with each message) | See [Specify Iceberg Schema](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/) to learn more about the different Iceberg modes. 3. If using `value_schema_id_prefix` or `value_schema_latest` modes, register a schema for the topic: ```bash rpk registry schema create -value --schema --type ``` > ❗ **IMPORTANT** > > If using the `value_schema_id_prefix` mode, schema subjects must use the `-value` [naming convention](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-id-validation/#set-subject-name-strategy-per-topic) (TopicNameStrategy). Note the schema ID returned, in case you need it for troubleshooting. 4. Verify that new records are being written to the Iceberg table: - Check that data appears in your query engine. - Validate that the schema translation is correct. - Confirm record counts are increasing. #### [](#multi-table-fan-out-pattern)Multi-table fan-out pattern If your existing integration routes records to multiple Iceberg tables based on a field value (for example, Kafka Connect’s `iceberg.tables.route-field` property), you need to implement equivalent routing logic. You create separate Iceberg-enabled topics for each target table, and Redpanda automatically creates corresponding Iceberg tables. Use either of the following approaches to route records to the correct topic: ##### [](#option-1-data-transforms-with-separate-topics-recommended)Option 1: Data transforms with separate topics (recommended) Use a data transform to read the routing field from each message and write records to separate Iceberg-enabled topics. This approach keeps routing logic within Redpanda and avoids external dependencies. When using Iceberg modes that require schema validation, the transform can register schemas dynamically and encode messages with the appropriate format. 1. Enable data transforms on your cluster: ```bash rpk cluster config set data_transforms_enabled true ``` 2. Create output topics and enable Iceberg with Schema Registry validation: ```bash rpk topic create rpk topic alter-config --set redpanda.iceberg.mode=value_schema_id_prefix rpk topic alter-config --set redpanda.iceberg.mode=value_schema_id_prefix rpk topic alter-config --set redpanda.iceberg.mode=value_schema_id_prefix ``` 3. Implement a transform function that: 1. Reads the routing field from each input message. 2. If using Schema Registry validation, registers schemas dynamically and encodes messages with the appropriate format. 3. Writes to a specific output topic based on the routing field. 4. Deploy the transform, specifying multiple output topics: ```bash rpk transform deploy \ --file transform.wasm \ --name \ --input-topic \ --output-topic \ --output-topic \ --output-topic ``` 5. Validate the fanout by checking that each output topic receives the correct records. For a complete implementation example with dynamic schema registration, see [Multi-topic fan-out with Schema Registry](https://docs.redpanda.com/cloud-data-platform/develop/data-transforms/build/#multi-topic-fanout). The example demonstrates Schema Registry wire format encoding for use with `value_schema_id_prefix` mode. ##### [](#option-2-external-stream-processor)Option 2: External stream processor Use an external stream processor for complex routing logic: 1. Use a stream processor ([Redpanda Connect](https://docs.redpanda.com/cloud-data-platform/develop/connect/about/) or Flink) to split records. 2. Write to separate Iceberg-enabled topics. This approach is more complex but offers more flexibility for advanced routing requirements not supported by data transforms. ### [](#validate-schema-registry-integration)Validate Schema Registry integration If using [`value_schema_id_prefix`](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/#value_schema_id_prefix) mode, verify that messages use the Schema Registry [wire format](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/#wire-format). ```bash rpk topic consume --num=1 --format='%v\n' | xxd | head -n 1 ``` If the first byte is not `00` (magic byte), you must configure your producer to use the wire format. The `value_schema_id_prefix` mode also requires that schema subjects follow the TopicNameStrategy: `-value`. Verify your schemas use the correct naming: ```bash rpk registry schema list ``` #### [](#verify-no-records-in-dlq)Verify no records in DLQ Check that no records failed validation and were written to the dead-letter queue. If records are present, see [Records in DLQ table](#records-in-dlq-table) for resolution steps. ```sql SELECT COUNT(*) FROM ."~dlq"; ``` ### [](#run-systems-in-parallel)Run systems in parallel Keep your existing Iceberg integration running while Iceberg Topics is enabled. This provides a safety net during the transition period: - New data flows to both the source tables and new Iceberg Topics tables. - You can validate data consistency between both systems. - You have a fallback option if issues arise. Run a query to compare record counts between systems: ```sql -- Source table SELECT COUNT(*) AS source_count FROM .; -- Iceberg Topics table SELECT COUNT(*) AS iceberg_topics_count FROM .; ``` Record counts should increase at similar rates, accounting for the time Iceberg Topics was enabled. Check for DLQ records (see [Records in DLQ table](#records-in-dlq-table)). Monitor Iceberg topic metrics to validate that data is flowing at expected rates: - `redpanda_iceberg_translation_parquet_rows_added`: Track rows written to Iceberg tables (compare with source write rate) - `redpanda_iceberg_translation_translations_finished`: Number of completed translation executions - `redpanda_iceberg_translation_invalid_records`: Records that failed validation - `redpanda_iceberg_translation_dlq_files_created`: Dead-letter queue activity - `redpanda_iceberg_rest_client_num_commit_table_update_requests_failed`: Failed table commits to catalog If using data transforms for multi-table fanout, also monitor: - `redpanda_transform_processor_lag`: Records pending processing in transform input topic For a complete list of Iceberg metrics, see the [Iceberg metrics reference](https://docs.redpanda.com/cloud-data-platform/reference/public-metrics-reference/#iceberg-metrics). > 💡 **TIP** > > Run both systems for at least 24-48 hours to ensure stability before proceeding with data merge. ### [](#merge-historical-data)Merge historical data Choose a strategy to combine your historical data with new Iceberg Topics data. #### [](#option-1-insert-into-pattern-recommended)Option 1: INSERT INTO pattern (recommended) Use this approach to create a unified table with all data, taking into consideration the following: - You want a single table for queries. - You can afford the one-time data copy cost. - You need optimal query performance. This SQL pattern uses partition and offset metadata to identify and copy only records not yet in the target table: ```sql -- Step 1: Find the latest offset per partition in the target (Iceberg Topics) table WITH latest_offsets AS ( SELECT partition, MAX(offset) AS max_offset FROM target_iceberg_topics_table GROUP BY partition ) -- Step 2: Insert records from source table that don't exist in target INSERT INTO target_iceberg_topics_table SELECT s.* FROM source_table AS s LEFT JOIN latest_offsets AS t ON s.partition = t.partition WHERE t.max_offset IS NULL -- Partition not seen before in target OR s.offset > t.max_offset; -- Record is newer than target's latest offset ``` - The `latest_offsets` CTE finds the highest offset in the target table for each partition. - The `LEFT JOIN` ensures you include partitions never seen before in the target (`t.max_offset IS NULL`). - The `WHERE` clause filters to only records with offsets greater than the target’s latest. - This avoids duplicates by using Kafka partition and offset as the deduplication key. This approach may take significant time for large datasets. Consider executing this process during low-query periods. You can also execute on an incremental basis to ease the load on your query engine, for example, by date or partition ranges. #### [](#option-2-view-based-query-federation)Option 2: View-based query federation Use this approach to query both tables without copying data if: - You cannot afford data copy time or cost. - You need immediate access to a unified view. - Query complexity and performance are acceptable with federated queries. - You may consolidate data later. Create a view that queries both tables and deduplicates on the fly: ```sql CREATE VIEW unified_iceberg_view AS WITH latest_offsets AS ( SELECT partition, MAX(offset) AS max_offset FROM target_iceberg_topics_table GROUP BY partition ), historical_data AS ( SELECT s.* FROM source_table AS s LEFT JOIN latest_offsets AS t ON s.partition = t.partition WHERE t.max_offset IS NULL OR s.offset <= t.max_offset -- Only historical records not in target ), new_data AS ( SELECT * FROM target_iceberg_topics_table ) SELECT * FROM historical_data UNION ALL SELECT * FROM new_data; ``` Most Iceberg-compatible query engines support views, including Snowflake, Databricks, ClickHouse, and Spark. ### [](#validate-the-migration)Validate the migration After completing the data merge, verify the migration before cutting over: - Record counts match between source and target: ```sql -- Compare record counts SELECT 'Source' AS table_name, COUNT(*) AS record_count FROM . UNION ALL SELECT 'Target', COUNT(*) FROM .; ``` - All partitions are represented in the target: ```sql -- Check for missing partitions SELECT DISTINCT partition FROM . EXCEPT SELECT DISTINCT partition FROM .; -- Should return no rows ``` - Date ranges cover the full historical period. Compare `MIN(timestamp)` and `MAX(timestamp)` between source and target tables to ensure the target covers the same time range. - No gaps in offset sequences: ```sql -- Check for offset gaps (may indicate missing data) WITH offset_check AS ( SELECT partition, offset, LAG(offset) OVER (PARTITION BY partition ORDER BY offset) AS prev_offset FROM . ) SELECT * FROM offset_check WHERE offset - prev_offset > 1; -- Should return no rows ``` - Sample queries return expected results. Spot check specific records by ID to verify data accuracy. - Schema translation is correct. Run `DESCRIBE` on both tables and verify all fields are present with correct data types. - New records are flowing to Iceberg Topics. Check record count for a recent time window (for example, the last hour). - Query performance is acceptable. - Monitoring and alerts are configured. - No records in DLQ (see [Records in DLQ table](#records-in-dlq-table)). ### [](#troubleshoot-common-migration-issues)Troubleshoot common migration issues #### [](#records-in-dlq-table)Records in DLQ table Iceberg Topics write records that fail validation to a dead-letter queue (DLQ) table. Records may appear in the DLQ due to: - Schema Registry issues. For example, using the wrong schema subject name, or Redpanda cannot find the embedded schema ID in Schema Registry. - When using `value_schema_id_prefix` mode: messages not encoded with Schema Registry wire format. - Incompatible schema changes. For example, changing field types or removing required fields. - Data type translation failures. To check for DLQ records during migration: ```sql SELECT COUNT(*) FROM ."~dlq"; ``` If the count is greater than zero, inspect the failed records. See [Troubleshoot Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-troubleshooting/) for steps to inspect and reprocess DLQ records. #### [](#multi-table-fan-out-transform-issues)Multi-table fan-out transform issues If the transform does not process messages, check if: - The specified output topics don’t exist or aren’t enabled with Iceberg. - The routing logic in the transform is incorrect, or the routing field is missing from input messages. - (When using Schema Registry validation) The schema registration failed during initialization, preventing the transform from starting. To check the transform status: ```bash rpk transform list ``` To view logs and check for errors: ```bash rpk transform logs ``` To check for routing errors: ```bash rpk transform logs | grep -i "unknown\|error" ``` If using Schema Registry validation, verify schema registration: ```bash # Check transform logs for schema registration messages rpk transform logs | grep -i "schema" # List registered schemas rpk registry schema list ``` ### [](#plan-for-rollback)Plan for rollback Before cutting over, ensure you have a rollback strategy. See the [Pre-cutover checklist](#pre-cutover-checklist) in the cutover section to verify you’re ready. #### [](#rollback-during-parallel-operation)Rollback during parallel operation If you discover issues while both systems are running: 1. Keep producing to both systems. 2. Point consumers back to source tables. 3. Investigate Iceberg Topics issues using troubleshooting section. 4. Fix issues and re-validate. 5. Attempt cutover again when ready. #### [](#rollback-after-external-integration-disabled)Rollback after external integration disabled > ⚠️ **WARNING** > > Rollback after stopping your external Iceberg integration may result in data loss or gaps. If you must rollback after disabling the external integration: 1. Restart your external Iceberg integration immediately. 2. Identify data written only to Iceberg Topics during the gap. 3. Export that data from Iceberg Topics tables: ```sql SELECT * FROM iceberg_topics_table WHERE timestamp > ''; ``` 4. Write exported data back to the source system (for example, Kafka Connect input topics or directly to source tables). 5. Verify data completeness across both systems. 6. Resume operations on the external integration. Redpanda recommends maintaining the ability to rollback for at least seven days after cutover to allow for issue discovery. ### [](#cut-over-to-iceberg-topics)Cut over to Iceberg Topics #### [](#pre-cutover-checklist)Pre-cutover checklist Before disabling your external Iceberg integration, ensure you have completed all validation steps: - All historical data is successfully merged (see [Merge historical data](#merge-historical-data)). - Parallel operation is complete and stable for at least 24-48 hours. - All validation queries pass (see [Validate the migration](#validate-the-migration)). - No records in DLQ tables, or all DLQ records are investigated and resolved. - Query performance meets requirements. - Downstream consumers are successfully tested with Iceberg Topics tables. - Monitoring and alerts are configured. - Rollback plan is verified and documented. #### [](#cutover-procedure)Cutover procedure 1. Set an appropriate maintenance window, ideally during low-traffic periods. 2. Stop your external Iceberg integration. **For Kafka Connect:** ```bash # Stop connector curl -X PUT http:///kafka-connect/clusters/iceberg-sink-connector/stop # Or delete connector (permanent) curl -X DELETE http:///kafka-connect/clusters/iceberg-sink-connector ``` 3. Monitor Iceberg Topics to ensure data continues flowing. 4. Verify that no new records are being written to source tables: ```sql SELECT MAX(timestamp) FROM .; -- Should not change after integration is stopped ``` 5. Run validation queries from [Validate the migration](#validate-the-migration) after 1-2 hours of operation. 6. Wait for a short period, such as 24-48 hours, to monitor and validate stability. 7. If migrating to a unified table of historical plus new data, optionally delete old source tables after an extended validation period (for example, at least seven days): > 📝 **NOTE** > > Ensure you have backups before deleting historical data. Some organizations keep old tables for compliance or audit purposes. ```sql DROP TABLE .; ``` 8. Decommission external Iceberg infrastructure after an extended safety period (30+ days, for example). If any issues arise during cutover, see [Plan for rollback](#plan-for-rollback). ## [](#next-steps)Next steps - [Query Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/query-iceberg-topics/) - [About Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/) --- # Page 410: Query Iceberg Topics **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/query-iceberg-topics.md --- # Query Iceberg Topics > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Query Iceberg Topics latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/query-iceberg-topics page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/query-iceberg-topics.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/query-iceberg-topics.adoc description: Query Redpanda topic data stored in Iceberg tables, based on the topic Iceberg mode and schema. page-git-created-date: "2025-04-04" page-git-modified-date: "2026-05-26" --- When you access Iceberg topics from a data lakehouse or other Iceberg-compatible tools, how you consume the data depends on the topic [Iceberg mode](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema/) and whether you’ve registered a schema for the topic in the [Redpanda Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/). You do not need to rely on complex ETL jobs or pipelines to access real-time data from Redpanda. After reading this page, you will be able to: - Query Redpanda topic data from an Iceberg-compatible engine - Grant end-user and query engine access to Iceberg data ## [](#access-iceberg-tables)Access Iceberg tables Redpanda generates an Iceberg table with the same name as the topic. Depending on the processing engine and your Iceberg catalog implementation, you may also need to define the table (for example using `CREATE TABLE`) to point the data lakehouse to its location in the catalog. For BYOC clusters, the bucket name and table location are as follows: | Cloud provider | Bucket or container name | Iceberg table location | | --- | --- | --- | | AWS | redpanda-cloud-storage- | redpanda-iceberg-catalog/redpanda/ | | Azure | The Redpanda cluster ID is also used as the container name (ID) and the storage account ID. | | GCP | redpanda-cloud-storage- | For BYOVPC clusters, the bucket name is the name you chose when you created the object storage bucket as a customer-managed resource. For Azure clusters, you must add the public IP addresses or ranges from the REST catalog service, or other clients requiring access to the Iceberg data, to your cluster’s allow list. Alternatively, add subnet IDs to the allow list if the requests originate from the same Azure region. For example, to add subnet IDs to the allow list through the Control Plane API [`PATCH /v1/clusters/`](https://docs.redpanda.com/api/doc/cloud-controlplane/operation/operation-clusterservice_updatecluster) endpoint, run: ```bash curl -X PATCH https://api.cloud.redpanda.com/v1/clusters/ \ -H "Content-Type: application/json" \ -H "Authorization: Bearer ${RP_CLOUD_TOKEN}" \ -d @- << EOF { "cloud_storage": { "azure": { "allowed_subnet_ids": [ ] } } } EOF ``` ### [](#grant-access-to-query-engine-users)Grant access to query engine users Redpanda manages the service-to-service permissions between Redpanda and the catalog (see [Use Iceberg Catalogs](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/)). However, you are responsible for granting your end users and query engines (such as Amazon Athena, Apache Spark, Trino, or Snowflake) read access to the Iceberg data. Use either or both of the following approaches to control access: #### [](#cloud-storage-prefix-level-access)Cloud storage prefix-level access Grant query engine roles or users read access to the Iceberg data prefix in the cluster’s storage bucket. This controls who can read the underlying data and metadata files. Scope permissions to specific prefixes to restrict access to individual tables. - AWS (S3): Use IAM policies to grant `s3:GetObject` and `s3:ListBucket` on the Iceberg prefix (for example, `/redpanda-iceberg-catalog/*`). See [Using IAM policies with Amazon S3](https://docs.aws.amazon.com/AmazonS3/latest/userguide/using-iam-policies.html). - GCP (GCS): Use IAM conditions or bucket-level policies to grant `storage.objects.get` and `storage.objects.list` on the Iceberg prefix. See [GCS IAM permissions](https://cloud.google.com/storage/docs/access-control/iam). - Azure (Blob Storage): Use Azure RBAC roles such as Storage Blob Data Reader scoped to the container or prefix. See [Authorize access to blob data](https://learn.microsoft.com/en-us/azure/storage/blobs/authorize-access-azure-active-directory). #### [](#catalog-level-table-access)Catalog-level table access If you use a REST catalog, you can control access at the table level through the catalog’s own access control layer. Use this approach when query engines access tables through the catalog rather than reading files directly. - AWS Glue: Use [AWS Lake Formation](https://docs.aws.amazon.com/lake-formation/latest/dg/what-is-lake-formation.html) to grant table-level and column-level permissions. - Databricks Unity Catalog: See the [Unity Catalog privileges documentation](https://docs.databricks.com/en/data-governance/unity-catalog/manage-privileges/index.html). - Snowflake Open Catalog: See [Open Catalog access control](https://docs.snowflake.com/en/user-guide/opencatalog/access-control). - GCP BigLake: See [BigLake table permissions](https://cloud.google.com/bigquery/docs/manage-open-source-metadata#grant_permissions). ### [](#refresh-table-data)Refresh table data Some query engines may require you to manually refresh the Iceberg table snapshot (for example, by running a command like `ALTER TABLE REFRESH;`) to see the latest data. If your engine needs the full JSON metadata path, use the following: ```none redpanda-iceberg-catalog/redpanda//metadata/v.metadata.json ``` This provides read access to all snapshots written as of the specified table version (denoted by `version-number`). > 📝 **NOTE** > > Redpanda automatically removes expired snapshots on a periodic basis. Snapshot expiry helps maintain a smaller metadata size and reduces the window available for [time travel](#time-travel-queries). ## [](#query-examples)Query examples To follow along with the examples on this page, suppose you produce the same stream of events to a topic `ClickEvent`, which uses a schema, and another topic `ClickEvent_key_value`, which uses the key-value mode. The topic’s Iceberg data is stored in an AWS S3 bucket. A sample record contains the following data: ```bash {"user_id": 2324, "event_type": "BUTTON_CLICK", "ts": "2024-11-25T20:23:59.380Z"} ``` > 📝 **NOTE** > > The query examples on this page use `redpanda` as the Iceberg namespace, which is the default. If you configured a different namespace using `[iceberg_default_catalog_namespace](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_default_catalog_namespace)`, replace `redpanda` with your configured namespace. ### [](#topic-with-schema-value_schema_id_prefix-mode)Topic with schema (`value_schema_id_prefix` mode) > 📝 **NOTE** > > The steps in this section also apply to the `value_schema_latest` mode, except the produce step. The `value_schema_latest` mode is not compatible with the Schema Registry wire format. The [`rpk topic produce`](#reference:rpk/rpk-topic/rpk-topic-produce) command embeds the wire format header, so you must use your own producer code with `value_schema_latest`. Assume that you have created the `ClickEvent` topic, set `redpanda.iceberg.mode` to `value_schema_id_prefix`, and are connecting to a REST-based Iceberg catalog. The following is an Avro schema for `ClickEvent`: `schema.avsc` ```avro { "type" : "record", "namespace" : "com.redpanda.examples.avro", "name" : "ClickEvent", "fields" : [ { "name": "user_id", "type" : "int" }, { "name": "event_type", "type" : "string" }, { "name": "ts", "type": "string" } ] } ``` 1. Register the schema under the `ClickEvent-value` subject: ```bash rpk registry schema create ClickEvent-value --schema path/to/schema.avsc --type avro ``` 2. Produce to the `ClickEvent` topic using the following format: ```bash echo '"key1" {"user_id":2324,"event_type":"BUTTON_CLICK","ts":"2024-11-25T20:23:59.380Z"}' | rpk topic produce ClickEvent --format='%k %v\n' --schema-id=topic ``` The `value_schema_id_prefix` mode requires that you produce to a topic using the [Schema Registry wire format](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/#wire-format), which includes the magic byte and schema ID in the prefix of the message payload. This allows Redpanda to identify the correct schema version in the Schema Registry for a record. 3. The following Spark SQL query returns values from columns in the `ClickEvent` table, with the table structure derived from the schema, and column names matching the schema fields. If you’ve integrated a catalog, query engines such as Spark SQL provide Iceberg integrations that allow easy discovery and access to existing Iceberg tables in object storage. ```sql SELECT * FROM ``.redpanda.ClickEvent; ``` ```bash +-----------------------------------+---------+--------------+--------------------------+ | redpanda | user_id | event_type | ts | +-----------------------------------+---------+--------------+--------------------------+ | {"partition":0,"offset":0,"timestamp":2025-03-05 15:09:20.436,"headers":null,"key":null} | 2324 | BUTTON_CLICK | 2024-11-25T20:23:59.380Z | +-----------------------------------+---------+--------------+--------------------------+ ``` ### [](#topic-in-key-value-mode)Topic in key-value mode In `key_value` mode, you do not associate the topic with a schema in the Schema Registry, which means using semi-structured data in Iceberg. The record keys and values can have an arbitrary structure, so Redpanda stores them in [binary format](https://apache.github.io/iceberg/spec/?h=spec#primitive-types) in Iceberg. In this example, assume that you have created the `ClickEvent_key_value` topic, and set `redpanda.iceberg.mode` to `key_value`. 1. Produce to the `ClickEvent_key_value` topic using the following format: ```bash echo '"key1" {"user_id":2324,"event_type":"BUTTON_CLICK","ts":"2024-11-25T20:23:59.380Z"}' | rpk topic produce ClickEvent_key_value --format='%k %v\n' ``` 2. The following Spark SQL query returns the semi-structured data in the `ClickEvent_key_value` table. The table consists of two columns: one named `redpanda`, containing the record key and other metadata, and another binary column named `value` for the record’s value: ```sql SELECT * FROM ``.redpanda.ClickEvent_key_value; ``` ```bash +-----------------------------------+------------------------------------------------------------------------------+ | redpanda | value | +-----------------------------------+------------------------------------------------------------------------------+ | {"partition":0,"offset":0,"timestamp":2025-03-05 15:14:30.931,"headers":null,"key":key1} | {"user_id":2324,"event_type":"BUTTON_CLICK","ts":"2024-11-25T20:23:59.380Z"} | +-----------------------------------+------------------------------------------------------------------------------+ ``` Depending on your query engine, you might need to first decode the binary value to display the record key and value using a SQL helper function. For example, see the [`decode` and `unhex`](https://spark.apache.org/docs/latest/api/sql/index.html#unhex) Spark SQL functions, or the [HEX\_DECODE\_STRING](https://docs.snowflake.com/en/sql-reference/functions/hex_decode_string) Snowflake function. Some engines may also automatically decode the binary value for you. ### [](#time-travel-queries)Time travel queries Some query engines, such as Spark, support time travel with Iceberg, allowing you to query the table as it existed at a specific point in the past. You can run a time travel query by specifying a timestamp or version number. Redpanda automatically removes expired snapshots on a periodic basis, which also reduces the window available for time travel queries. By default, Redpanda retains snapshots for five days, so you can query Iceberg tables as of up to five days ago. The following example queries a `ClickEvent` table at a specific timestamp in Spark: ```sql SELECT * FROM ``.redpanda.ClickEvent TIMESTAMP AS OF '2025-03-02 10:00:00'; ``` --- # Page 411: Query Iceberg Topics using Snowflake and Open Catalog **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/redpanda-topics-iceberg-snowflake-catalog.md --- # Query Iceberg Topics using Snowflake and Open Catalog > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Query Iceberg Topics using Snowflake and Open Catalog latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/redpanda-topics-iceberg-snowflake-catalog page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/redpanda-topics-iceberg-snowflake-catalog.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/redpanda-topics-iceberg-snowflake-catalog.adoc description: Add Redpanda topics as Iceberg tables that you can query in Snowflake using an Open Catalog integration. page-git-created-date: "2025-05-21" page-git-modified-date: "2026-05-26" --- This guide walks you through querying Redpanda topics as Iceberg tables in [Snowflake](https://docs.snowflake.com/en/user-guide/tables-iceberg), with Amazon S3 as object storage and a catalog integration using [Open Catalog](https://docs.snowflake.com/en/user-guide/opencatalog/overview). After reading this page, you will be able to: - Configure AWS IAM credentials granting Open Catalog access to your S3 bucket - Integrate Redpanda Iceberg topics with Snowflake using Open Catalog ## [](#prerequisites)Prerequisites - `rpk` or familiarity with the Redpanda Cloud API to use secrets in your cluster configuration. For `rpk`, see [Install or Update rpk](https://docs.redpanda.com/cloud-data-platform/manage/rpk/rpk-install/). For the Cloud API, you must [authenticate](https://docs.redpanda.com/api/cloud-controlplane/authentication) using a service account. - A Snowflake account. - An Open Catalog account. To [create an Open Catalog account](https://other-docs.snowflake.com/en/opencatalog/create-open-catalog-account), you require ORGADMIN access in Snowflake. ## [](#authorize-access-to-open-catalog)Authorize access to Open Catalog You must create an AWS IAM policy and role that Open Catalog and Snowflake use to access the S3 bucket where your Iceberg data is stored. Redpanda writes Iceberg data and metadata files to the bucket using your cluster’s existing object storage credentials, so Redpanda does not need additional IAM configuration for its own S3 access. You finish configuring this role’s trust policy when you create the catalog and the external volume, using values that Open Catalog and Snowflake generate for you. ### [](#create-an-iam-policy)Create an IAM policy Create an IAM policy with the following S3 permissions, scoped to your cluster’s storage bucket: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:PutObject", "s3:GetObject", "s3:GetObjectVersion", "s3:DeleteObject", "s3:DeleteObjectVersion" ], "Resource": "arn:aws:s3:::/*" }, { "Effect": "Allow", "Action": [ "s3:ListBucket", "s3:GetBucketLocation" ], "Resource": "arn:aws:s3:::" } ] } ``` Replace `` with the name of your cluster’s object storage bucket. You use this same IAM policy for both the Open Catalog role and the Snowflake external volume. ### [](#create-an-iam-role)Create an IAM role Create an IAM role and attach the IAM policy you created. Open Catalog and Snowflake each need this role’s ARN before they can generate the IAM user and external ID values you use to finish configuring the trust policy, so create the role with a temporary trust relationship for now: 1. In the AWS IAM console, create a new role. 2. For the trusted entity type, select **AWS account**. Under **An AWS account**, select **This account**. 3. Attach the IAM policy you created. 4. Note the role’s ARN (``). You update this role’s trust policy after Open Catalog and Snowflake generate the values you need, when you create the catalog and the external volume. ## [](#create-a-catalog-in-open-catalog)Create a catalog in Open Catalog Create the catalog that Redpanda’s Iceberg data and metadata are registered to, with your Tiered Storage S3 bucket configured as external storage: 1. In Open Catalog, create a new catalog. 2. For **S3 role ARN**, enter ``, the ARN of the IAM role you created. 3. Configure the catalog’s external storage location to point to your Tiered Storage S3 bucket. > 📝 **NOTE** > > Your Open Catalog account must be in the same AWS region as your S3 bucket. 4. On the Open Catalog home page, in the **Catalogs** pane, select the catalog you created. Under **Storage Details**, copy the **IAM user arn** (``). 5. If you didn’t specify an external ID when you created the IAM role, Open Catalog generates one for you (``). Record this value. For complete steps, see the [Open Catalog documentation](https://docs.snowflake.com/en/user-guide/opencatalog/create-catalog). ### [](#update-the-iam-role-trust-policy-for-open-catalog)Update the IAM role trust policy for Open Catalog Update the trust policy for the IAM role you created, using the IAM user ARN and external ID that Open Catalog generated when you created the catalog: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "AWS": "" }, "Action": "sts:AssumeRole", "Condition": { "StringEquals": { "sts:ExternalId": "" } } } ] } ``` After you update the trust policy, Open Catalog can assume the role to read and write Iceberg data and metadata in your S3 bucket. ## [](#create-a-snowflake-external-volume)Create a Snowflake external volume Create a Snowflake external volume using the same IAM role you created for Open Catalog. Snowflake provisions its own IAM user to assume the role, so you add a second statement to the role’s trust policy rather than replacing the statement that trusts Open Catalog: 1. In Snowflake, run `CREATE EXTERNAL VOLUME`, pointing `STORAGE_AWS_ROLE_ARN` to the IAM role you created: ```sql CREATE OR REPLACE EXTERNAL VOLUME STORAGE_LOCATIONS = ( ( NAME = '' STORAGE_PROVIDER = 'S3' STORAGE_BASE_URL = 's3:///' STORAGE_AWS_ROLE_ARN = '' STORAGE_AWS_EXTERNAL_ID = '' ) ) ALLOW_WRITES = TRUE; ``` Use your own values for the following placeholders: - ``: Provide a name for your external volume in Snowflake. - ``: Provide a name for the storage location. - ``: Choose an external ID for the volume’s trust relationship, distinct from the external ID Open Catalog uses. If you don’t set `STORAGE_AWS_EXTERNAL_ID`, Snowflake generates one for you. 2. Run `DESC EXTERNAL VOLUME` to retrieve the IAM user ARN that Snowflake generated for the volume: ```sql DESC EXTERNAL VOLUME ; ``` Record the `STORAGE_AWS_IAM_USER_ARN` value from the output (``). 3. Add a new statement to the IAM role’s trust policy for the volume’s IAM user ARN and external ID, alongside the existing statement that trusts Open Catalog: > ❗ **IMPORTANT** > > Add this as a new statement in the trust policy’s `Statement` array. If you replace the existing trust policy instead, Open Catalog loses access to the bucket. ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "AWS": "" }, "Action": "sts:AssumeRole", "Condition": { "StringEquals": { "sts:ExternalId": "" } } }, { "Effect": "Allow", "Principal": { "AWS": "" }, "Action": "sts:AssumeRole", "Condition": { "StringEquals": { "sts:ExternalId": "" } } } ] } ``` 4. Run `SYSTEM$VERIFY_EXTERNAL_VOLUME` to confirm Snowflake can access the bucket: ```sql SELECT SYSTEM$VERIFY_EXTERNAL_VOLUME(''); ``` For complete steps, see the [Snowflake documentation](https://docs.snowflake.com/en/user-guide/tables-iceberg-configure-external-volume-s3). ## [](#set-up-catalog-integration-using-open-catalog)Set up catalog integration using Open Catalog To integrate Iceberg-enabled topics with Open Catalog, create a service connection and configure catalog roles. ### [](#create-a-new-open-catalog-service-connection-for-redpanda)Create a new Open Catalog service connection for Redpanda To create a new service connection to integrate the Iceberg-enabled topics into Open Catalog: 1. In Open Catalog, select **Connections**, then **\+ Connection**. 2. In **Configure Service Connection**, provide a name. Open Catalog creates a new principal with this name. 3. Make sure **Create new principal role** is selected. 4. Enter a name for the principal role. Then, click **Create**. After you create the connection, get the client ID and client secret. Save these credentials to add to your cluster configuration in a later step. ### [](#create-a-catalog-role)Create a catalog role Grant privileges to the principal created in the previous step: 1. In Open Catalog, select **Catalogs**, and select your catalog. 2. On the **Roles** tab of your catalog, click **\+ Catalog Role**. 3. Give the catalog role a name. 4. Under **Privileges**, select `CATALOG_MANAGE_CONTENT`. This provides full management [privileges](https://docs.snowflake.com/en/user-guide/opencatalog/access-control#catalog-privileges) for the catalog. Then, click **Create**. 5. On the **Roles** tab of the catalog, click **Grant to Principal Role**. 6. Select the catalog role you just created. 7. Select the principal role you created earlier. Click **Grant**. ### [](#update-cluster-configuration)Update cluster configuration To configure your Redpanda cluster to enable Iceberg on a topic and integrate with Open Catalog: 1. [Store the Open Catalog client secret in your cluster](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/#store-a-secret-for-rest-catalog-authentication) using `rpk` or the Data Plane API. 2. [Edit your cluster configuration](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/#use-a-secret-in-cluster-configuration) to set the `iceberg_enabled` property to `true`, and set the catalog integration properties listed in the example below using `rpk` or the Control Plane API. For example, to use `rpk cluster config set`, run: ```bash rpk cluster config set \ iceberg_enabled=true \ iceberg_catalog_type=rest \ iceberg_rest_catalog_endpoint=https://-.snowflakecomputing.com/polaris/api/catalog \ iceberg_rest_catalog_authentication_mode=oauth2 \ iceberg_rest_catalog_client_id= \ iceberg_rest_catalog_client_secret='${secrets.}' \ iceberg_rest_catalog_warehouse= # Optional properties: # iceberg_translation_interval_ms_default=1000 # iceberg_catalog_commit_interval_ms=1000 ``` Use your own values for the following placeholders: - `` and ``: Your [Open Catalog account URI](https://docs.snowflake.com/en/sql-reference/sql/create-catalog-integration-open-catalog#required-parameters) is composed of these values. > 💡 **TIP** > > In Snowflake, navigate to **Admin**, then **Accounts**. Click the ellipsis near your Open Catalog account name, and select **Manage URLs**. The **Current URL** contains `` and ``. - ``: The client ID of the service connection you created in an earlier step. - ``: The name of the secret you created in the previous step. You must pass the secret name to the `${secrets.}` placeholder, not the secret value itself. - ``: The name of your catalog in Open Catalog. ```bash Successfully updated configuration. New configuration version is 2. ``` 3. Enable the integration for a topic by configuring the topic property `redpanda.iceberg.mode`. This mode creates an Iceberg table for the topic consisting of two columns: one for the record metadata including the key, and another binary column for the record’s value. See [Enable Iceberg integration](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/#enable-iceberg-integration) for more details on Iceberg modes. Use any of the following to set `redpanda.iceberg.mode`: - `rpk`. See the following examples to run `rpk topic` commands. - The Cloud UI. Navigate to **Topics** to create a new topic and specify `redpanda.iceberg.mode` in **Additional Configuration**, or edit an existing topic under the topic’s **Configuration** tab. - The Data Plane API to [create a new topic](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-topicservice_createtopic) or [update a property for an existing topic](https://docs.redpanda.com/api/doc/cloud-dataplane/operation/operation-topicservice_updatetopicconfigurations). Specify the key-value pair for `redpanda.iceberg.mode` in the request body. The following examples show how to use `rpk` to create a new topic or alter the configuration for an existing topic, setting the Iceberg mode to `key_value`. Create a new topic and set `redpanda.iceberg.mode`: ```bash rpk topic create --topic-config=redpanda.iceberg.mode=key_value ``` Set `redpanda.iceberg.mode` for an existing topic: ```bash rpk topic alter-config --set redpanda.iceberg.mode=key_value ``` 4. Produce to the topic. For example, ```bash echo "hello world\nfoo bar\nbaz qux" | rpk topic produce --format='%k %v\n' ``` The topic appears as a table in Open Catalog. 1. In Open Catalog, select **Catalogs**, then open your catalog. 2. Under your catalog, you will see the `redpanda` namespace (or the namespace you configured with `[iceberg_default_catalog_namespace](https://docs.redpanda.com/cloud-data-platform/reference/properties/cluster-properties/#iceberg_default_catalog_namespace)`), and a table with the name of your topic. The namespace and the table are automatically added for you. ## [](#query-iceberg-table-in-snowflake)Query Iceberg table in Snowflake To query the topic in Snowflake, you must create a [catalog integration](https://docs.snowflake.com/en/user-guide/tables-iceberg#catalog-integration) so that Snowflake has access to the table data and metadata. ### [](#configure-catalog-integration-with-snowflake)Configure catalog integration with Snowflake 1. Run the [`CREATE CATALOG INTEGRATION`](https://docs.snowflake.com/sql-reference/sql/create-catalog-integration-open-catalog) command in Snowflake: ```sql CREATE CATALOG INTEGRATION CATALOG_SOURCE = POLARIS TABLE_FORMAT = ICEBERG CATALOG_NAMESPACE = 'redpanda' REST_CONFIG = ( CATALOG_URI = '' WAREHOUSE = '' ) REST_AUTHENTICATION = ( TYPE = OAUTH OAUTH_CLIENT_ID = '' OAUTH_CLIENT_SECRET = '' OAUTH_ALLOWED_SCOPES = ('PRINCIPAL_ROLE:ALL') ) REFRESH_INTERVAL_SECONDS = 30 ENABLED = TRUE; ``` Use your own values for the following placeholders: - ``: Provide a name for your Iceberg catalog integration in Snowflake. - ``: Your [Open Catalog account URI](https://docs.snowflake.com/en/sql-reference/sql/create-catalog-integration-open-catalog#required-parameters) (`[https://-.snowflakecomputing.com/polaris/api/catalog](https://\-\.snowflakecomputing.com/polaris/api/catalog)`). - ``: The name of your catalog in Open Catalog. - ``: The client ID of the service connection you created in an earlier step. - ``: The client secret of the service connection you created in an earlier step. 2. Run the following command to verify that the catalog is integrated correctly: ```sql SELECT SYSTEM$LIST_ICEBERG_TABLES_FROM_CATALOG(''); ``` ```bash # Example result for redpanda.iceberg.mode=key_value +-----------------------------------------------------------------------+ | SYSTEM$LIST_ICEBERG_TABLES_FROM_CATALOG('') | +-----------------------------------------------------------------------+ | [{"namespace":"redpanda","name":""}] | +-----------------------------------------------------------------------+ ``` ### [](#create-iceberg-table-in-snowflake)Create Iceberg table in Snowflake After creating the catalog integration, you must create an externally-managed table in Snowflake. You must run your Snowflake queries against this table. In your Snowflake database, run the [CREATE ICEBERG TABLE](https://docs.snowflake.com/en/sql-reference/sql/create-iceberg-table-rest) command. The following example also specifies that the table should automatically refresh metadata: ```sql CREATE ICEBERG TABLE CATALOG = '' EXTERNAL_VOLUME = '' CATALOG_TABLE_NAME = '' AUTO_REFRESH = TRUE ``` Use your own values for the following placeholders: - ``: Provide a name for your table in Snowflake. - ``: The name of the catalog integration you configured in an earlier step. - ``: The name of the external volume you configured using the Tiered Storage bucket. - ``: The name of the table in your catalog, which is the same as your Redpanda topic name. ### [](#query-the-iceberg-table)Query the Iceberg table To verify that Snowflake has successfully created the table containing the topic data, run the following: ```sql SELECT * FROM ; ``` Your query results should look like the following: ```bash # Example for redpanda.iceberg.mode=key_value with 3 records produced to topic +--------------------------------------------------------------------------------------------------------------+------------+ | REDPANDA | VALUE | +--------------------------------------------------------------------------------------------------------------+------------+ | { "partition": 0, "offset": 0, "timestamp": "2025-02-07 16:29:50.122", "headers": null, "key": "68656C6C6F"} | 776F726C64 | | { "partition": 0, "offset": 1, "timestamp": "2025-02-07 16:29:50.122", "headers": null, "key": "666F6F"} | 626172 | | { "partition": 0, "offset": 2, "timestamp": "2025-02-07 16:29:50.122", "headers": null, "key": "62617A" } | 717578 | +--------------------------------------------------------------------------------------------------------------+------------+ ``` ### [](#manage-access-for-query-engine-users)Manage access for query engine users Redpanda manages the permissions between Redpanda and Open Catalog. To grant your Snowflake users or other query engines read access to the Iceberg tables, use [Open Catalog access control](https://docs.snowflake.com/en/user-guide/opencatalog/access-control) to assign catalog privileges. For example, you can grant `TABLE_READ_DATA` to a read-only role rather than the `CATALOG_MANAGE_CONTENT` privilege used by the Redpanda service principal. ## [](#suggested-reading)Suggested reading - [Query Iceberg Topics](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/query-iceberg-topics/) --- # Page 412: Integrate with REST Catalogs **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/rest-catalog.md --- # Integrate with REST Catalogs > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Integrate with REST Catalogs latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/rest-catalog/index page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/rest-catalog/index.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/rest-catalog/index.adoc description: Integrate Redpanda topics with managed Iceberg REST Catalogs. page-git-created-date: "2025-08-05" page-git-modified-date: "2025-11-27" --- > 💡 **TIP** > > These guides are for integrating Iceberg topics with managed REST catalogs. Integrating with a REST catalog is recommended for production deployments. If it is not possible to use a REST catalog, you can use the [filesystem-based catalog](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/use-iceberg-catalogs/#object-storage). For an example of using the filesystem-based catalog to access Iceberg topics, see the [Getting Started with Iceberg Topics on Redpanda BYOC](https://www.redpanda.com/blog/iceberg-topics-redpanda-cloud-byoc-setup) blog post. - [Query Iceberg Topics using AWS Glue](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-aws-glue/) Add Redpanda topics as Iceberg tables that you can access through the AWS Glue Data Catalog. - [Query Iceberg Topics using Databricks and Unity Catalog](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-databricks-unity/) Add Redpanda topics as Iceberg tables that you can query in Databricks managed by Unity Catalog. - [Use Iceberg Topics with GCP Lakehouse](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-topics-gcp-biglake/) Add Redpanda topics as Iceberg tables to Google Lakehouse for Apache Iceberg that you can query from Google BigQuery. - [Query Iceberg Topics using Snowflake and Open Catalog](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/redpanda-topics-iceberg-snowflake-catalog/) Add Redpanda topics as Iceberg tables that you can query in Snowflake using an Open Catalog integration. --- # Page 413: Specify Iceberg Schema **URL**: https://docs.redpanda.com/cloud-data-platform/manage/iceberg/specify-iceberg-schema.md --- # Specify Iceberg Schema > For the complete documentation index, see [llms.txt](https://docs.redpanda.com/llms.txt). Component-specific: [cloud-data-platform-full.txt](https://docs.redpanda.com/cloud-data-platform-full.txt) --- title: Specify Iceberg Schema latest-operator-version: v26.2.2 latest-console-tag: v3.11.0 latest-connect-version: 4.107.0 latest-redpanda-tag: v26.2.2 docname: iceberg/specify-iceberg-schema page-component-name: cloud-data-platform page-version: master page-component-version: master page-component-title: Cloud page-relative-src-path: iceberg/specify-iceberg-schema.adoc page-edit-url: https://github.com/redpanda-data/cloud-docs/edit/main/modules/manage/pages/iceberg/specify-iceberg-schema.adoc description: Learn about supported Iceberg modes and how you can integrate schemas with Iceberg topics. page-git-created-date: "2025-07-31" page-git-modified-date: "2026-05-26" --- In [Iceberg-enabled clusters](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/about-iceberg-topics/#enable-iceberg-integration), the `redpanda.iceberg.mode` topic property determines how Redpanda maps topic data to the Iceberg table structure. You can have the generated Iceberg table match the structure of a schema in Schema Registry, or you can use the `key_value` mode where Redpanda stores the record values as-is in the table. After reading this page, you will be able to: - Configure the redpanda.iceberg.mode property when you create or update a topic - Choose the Iceberg mode that produces the table structure your data consumers need - Apply independent translation for record keys, values, and headers ## [](#supported-iceberg-modes)Supported Iceberg modes Redpanda supports the following modes for Iceberg topics: ### [](#key_value)key_value Creates an Iceberg table using a simple schema, consisting of two columns, one for the record metadata including the key, and another binary column for the record’s value. ### [](#value_schema_id_prefix)value_schema_id_prefix Creates an Iceberg table whose structure matches the Redpanda schema for the topic, with columns corresponding to each field. You must register a schema in [Schema Registry](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/) and producers must write to the topic using the Schema Registry wire format. In the [Schema Registry wire format](https://docs.redpanda.com/cloud-data-platform/manage/schema-reg/schema-reg-overview/#wire-format), a "magic byte" and schema ID are embedded in the message payload header. Producers to the topic must use the wire format in the serialization process so Redpanda can determine the schema used for each record, use the schema to define the Iceberg table, and store the topic values in the corresponding table columns. ### [](#value_schema_latest)value_schema_latest Creates an Iceberg table whose structure matches the latest schema registered for the subject in Schema Registry. You must register a schema in Schema Registry. Producers cannot use the wire format in `value_schema_latest` mode. Redpanda expects the serialized message as-is without the magic byte or schema ID prefix in the record value. > 📝 **NOTE** > > The `value_schema_latest` mode is not compatible with the [`rpk topic produce`](#reference:rpk/rpk-topic/rpk-topic-produce) command which embeds the wire format header. You must use your own producer code to produce to topics in `value_schema_latest` mode. The latest schema is cached periodically. The cache period is defined by the cluster property `iceberg_latest_schema_cache_ttl_ms` (default: 5 minutes). ### [](#disabled)disabled Default for `redpanda.iceberg.mode`. Disables writing to an Iceberg table for the topic. > 📝 **NOTE** > > The following modes are compatible with producing to an Iceberg topic using Redpanda Console: > > - `key_value` > > - Starting in version 25.2, `value_schema_latest` with a JSON schema > > > Otherwise, records may fail to write to the Iceberg table and instead write to the [dead-letter queue](https://docs.redpanda.com/cloud-data-platform/manage/iceberg/iceberg-troubleshooting/#dead-letter-queue). ## [](#configure-iceberg-mode-for-a-topic)Configure Iceberg mode for a topic You can set the Iceberg mode for a topic when you create the topic, or you can update the mode for an existing topic. Option 1. Create a new topic and set `redpanda.iceberg.mode`: ```bash rpk topic create --topic-config=redpanda.iceberg.mode= ``` Option 2. Set `redpanda.iceberg.mode` for an existing topic: ```bash rpk topic alter-config --set redpanda.iceberg.mode= ``` ### [](#override-value-schema-latest-default)Override `value_schema_latest` default In `value_schema_latest` mode, you only need to set the property value to the string `value_schema_latest`. This enables the default behavior of `value_schema_latest` mode, which determines the subject for the topic using the TopicNameStrategy. For example, if your topic is named `sensor` the schema is looked up in the `sensor-value` subject. For Protobuf data, the default behavior also deserializes records using the first message defined in the corresponding Protobuf schema stored in Schema Registry. If you use a different strategy other than the topic name to derive the subject name, you can override the default behavior of `value_schema_latest` mode and explicitly set the subject name. To override the default behavior, use the following optional syntax: ```bash value_schema_latest:subject=,protobuf_name= ``` - For both Avro and Protobuf, specify a different subject name by using the key-value pair `subject=`, for example `value_schema_latest:subject=sensor-data`. - For Protobuf only: - Specify a different message definition by using a key-value pair `protobuf_name=`. You must use the fully qualified name, which includes the package name, for example, `value_schema_latest:protobuf_name=com.example.manufacturing.SensorData`. - To specify both a different subject and message definition, separate the key-value pairs with a comma, for example: `value_schema_latest:subject=my_protobuf_schema,protobuf_name=com.example.manufacturing.SensorData`. > 📝 **NOTE** > > If you don’t specify the fully qualified Protobuf message name, Redpanda pauses the data translation to the Iceberg table until you fix the topic misconfiguration. ## [](#configure-key-value-and-header-translation)Configure key, value, and header translation In addition to the [supported modes](#supported-iceberg-modes), `redpanda.iceberg.mode` also accepts a section-based syntax that lets you independently configure how Redpanda translates the record key, value, and headers into the Iceberg table. The `key_value`, `value_schema_id_prefix`, and `value_schema_latest` modes are shorthand for common combinations of these sections (see [Iceberg mode shorthands](#iceberg-mode-shorthands)). The `key` and `headers` sections change fields inside the `redpanda` system struct column (`redpanda.key` and the `value` field of each entry in `redpanda.headers`), while the `value` section changes the columns outside that struct. See [How Iceberg modes translate to table format](#how-iceberg-modes-translate-to-table-format) for the base row structure that every generated table includes. Use the following syntax to configure one or more sections: ```bash
: