Skip to main content
Feedback

Data Hub: CDC flow set up

Updated September 22, 2026

This guide describes how to configure Change Data Capture (CDC) between Boomi Data Hub (producer) and Boomi Data Integration (consumer) using Boomi Event Streams.

note

The Data Hub source connector is available as a limited availability (Beta) release.

With CDC, Data Hub publishes golden-record changes to Boomi Event Streams, and Data Integration consumes them as they occur. This replaces the legacy batch staging-queue path and reduces latency from minutes-to-hours to sub-second.

Setup has two stages: you configure Data Hub and Event Streams once per repository so Data Hub can publish changes, then you create the connection in Data Integration so it can consume them.

For connection field reference, refer to Data Hub connection. For standard (non-CDC) extraction setup, refer to Data Hub walkthrough.

Setting up Data Hub and Event Streams (one-time)

Creating the Event Streams token

  1. Use the service switcher at the top right of the platform and go to Services > Event Streams.
  2. On the Environments dashboard, click View Details next to the target runtime environment (for example, Atom Prod Cloud).
  3. Open the Settings tab.
  4. Scroll to the Tokens section and click + Create a Token.
  5. Configure the token:
    • Name: a unique name, for example DataHub_Outbound_Token.
    • Expiration Date: a policy-compliant expiry date.
    • Authorizations: select both Produce and Consume. Data Hub needs Produce authorization to publish messages, and Data Integration's consumer needs Consume authorization to read them using the same token.
  6. Click Save, then immediately copy the generated token string. You enter this same value in the JWT token field when you create the connection in Data Integration.
warning

Treat this token like a password. It is not displayed in full again. If it is lost, create a new one and bind it to the repository again.

Binding the token to the Data Hub repository

  1. Switch services to Data Hub.
  2. Click Repositories in the left menu and open the target repository.
  3. Open the Configure tab.
  4. Find the Event Streams Environment Token field and click Edit (or Add if it is blank).
  5. Paste the token you created in the previous step.
  6. Click Save. The platform validates the token signature and binds the connection.

Make sure your target Universe model is already published and deployed in this repository. You do not need to change the model's delivery mode or perform any other setup in Data Hub. When you activate a data flow with CDC enabled, Data Integration automatically creates the listener source, attaches it to your deployed model, and resolves the streaming topic. Refer to How CDC activation works.

Topic creation: The underlying streaming topic name is an internal implementation detail. You never see or enter a topic name anywhere in this setup. Data Integration resolves and creates it automatically, from the Repository ID and the other CDC credentials you provide, when you activate a data flow with CDC enabled.

Creating the connection in Data Integration

  1. In Data Integration, create a new Data Hub connection.
  2. Fill in the standard Data Hub fields (below). These are required whether or not CDC is used.
  3. Turn on Enable CDC Connectivity (Events Stream). Four additional fields appear.
  4. Fill in the CDC fields.
  5. Click Test Connection and resolve any errors before saving.
  6. Click Save, then build the data flow, select your models, and activate. Refer to Data Hub walkthrough for the data flow steps.

When the toggle is off, the connection applies to standard extraction only.

Standard fields (required)

FieldDescription
Base URL for API RequestsThe Hub Cloud region for your repository, selected from a dropdown. This same selection also determines the Event Streams host used for CDC, so choose the value that matches both your Data Hub repository and your Event Streams environment.
UsernameYour Boomi Account ID, in the format <account>.<user>, from the repository Configure tab.
My Hub Authentication TokenThe per-repository secret token, from the repository Configure tab.

CDC fields required when Enable CDC connectivity is on

  • JWT token: The token you created in Creating the Event Streams token, with both Produce and Consume authorizations.
  • DataHub API Key: A secret key required in addition to the My Hub Authentication Token above.
  • Repository ID: The identifier of the Data Hub repository this connection targets. This is a required value you enter directly; it is not derived from the My Hub Authentication Token. Data Integration uses it to attach the CDC listener source to your repository during activation.
  • Account Email: The email address associated with your Data Hub account.

Data Integration does not ask for a separate Event Streams Environment ID, region or API host selector, or topic name for CDC. Only the four fields above are required, and the Base URL for API Requests dropdown determines the Event Streams host as well as the Hub host.

Test Connection

Test Connection is read-only: it never creates, attaches, or modifies a Data Hub resource. With CDC on, it performs three checks:

  • Calls the Data Hub Repository API to confirm your Base URL, Username, and My Hub Authentication Token are valid and the repository is reachable.
  • Decodes the JWT token locally (no network call) and confirms it contains the claims Data Integration needs.
  • Confirms DataHub API Key, Repository ID, and Account Email are present. It does not validate their values.

Test Connection never contacts Event Streams, the underlying message broker, or the Platform API, and it does not resolve or check the streaming topic. A successful Test Connection with CDC enabled confirms your Data Hub credentials and JWT token are usable, but it does not guarantee CDC will activate: the DataHub API Key, Repository ID, and Account Email values are first exercised when you activate the data flow.

Common setup errors

What you seeCauseFix
CDC requires Event Streams URL and Token.CDC toggled on with fields left blankComplete the CDC fields.
Topic not found or access denied.The Repository ID, JWT token, or DataHub API Key is incorrect. This does not surface during Test Connection; it appears when the data flow activates and Data Integration attempts to resolve and attach to the topic.Re-check the Repository ID and CDC credentials on the connection.
Insufficient permissions: token must include consume.The JWT token lacks the required Consume authorization. This does not surface during Test Connection; it appears when Data Integration's consumer attempts to read from the topic.Recreate the token with both Produce and Consume authorizations and bind it to the repository again.
Deployment blocked in Data HubThe repository has no Event Streams Environment TokenBind the token to the repository before deploying the model.

How CDC activation works

When you activate a data flow with CDC enabled, Data Integration automatically provisions everything needed to stream from Data Hub.

  1. Creates a listener source scoped to this data flow.
  2. Attaches the source to your deployed Universe model.
  3. Imports the model's domain source configuration.
  4. Resolves the underlying channel.
  5. Reconstructs the streaming topic name.

Data Integration never converts, upgrades, or otherwise modifies your model.

Each of these provisioning steps is idempotent: reactivating the same data flow does not create duplicate sources or topics.

note

If you promote a data flow to a different environment, it gets a new source and connector. The original source is not cleaned up automatically and remains on your Data Hub account after promotion.

CDC event contract

Data Hub delivers changes as XML messages wrapped in a realTimeDeliveryMessage envelope:

FieldMeaning
recordIdStable identifier of the golden record.
recordVersionIdMonotonic version number of the golden record.
opTypeCREATE, UPDATE, or DELETE.
recordUpdatedAtTimestamp of the change at the source.
createdAtTimestamp the event was created. Refer to Expected behavior below.
channelUpdateTypeFULL (entire record). This is the only value Data Integration supports.
recordThe golden record payload, wrapped by the model's root element.
warning

Data Integration only supports a channelUpdateType of FULL. If Data Hub is configured to send DIFF (changed fields only), Data Integration cannot parse the message. This is a deliberate hard failure: streaming stops rather than applying a partial record. Configure the channel to send full records.

CREATE, UPDATE, and DELETE map to insert, update, and delete in your target. An unsupported operation does not fail the data flow. The event is logged and triggers an internal alert; it is not counted and is not shown in your data flow's run.

Event Streams enforces a maximum message size of 1 MB. Delivery is at-least-once, so the same event can arrive more than once, for example after a restart. With Merge loading and a stable primary key, your target holds exactly one correct row per record; in Append mode, or with no primary key, duplicates are possible.

CDC streaming start position

There are two start modes. There is no start from the oldest message mode.

AutomaticRe-initial sync
Start positionResumes from the last committed offsetStarts from now
When it appliesEvery run of an already-running data flow. It restarts, redeploys, or resuming after a pauseFirst activation, an explicit re-sync, or backfilling a newly added column
Historical dataContinues from where it stopped; nothing is skippedComes from the initial load (migration), not from the stream
Existing target dataPreservedReloaded by the accompanying migration

If the stored offset is no longer available, the data flow fails with an error telling you a re-initial sync is required, rather than silently restarting from now.

In the Schema tab, a table's Status column shows Waiting For Migration while the initial load is in progress, then changes to Live once the table has switched over to streaming.

Schema tab showing tables with Status: Waiting For Migration

Schema tab showing a table with Status: Live

CDC schema evolution

Data Integration absorbs Data Hub schema changes while your data flow keeps running:

Change in Data HubTarget actionRows already landedRe-sync needed
Field addedColumn added to the targetKept; new column is nullNo
Field removedColumn retained, no longer populatedPreserved; new rows carry nullNo
Field renamedTreated as remove + addOld column preserved; new column starts nullNo
Type widened compatibly (for example, int to bigint)Column altered in placePreservedNo
Type narrowed or incompatible (for example, string to int)That model is blockedPreserved, untouchedYes
Model added to or removed from the flowTable added or removedOther tables unaffectedNo

The merge key for a Data Hub CDC source is always recordId. Changing a model's own primary key in Data Hub has no effect on how Data Integration merges records.

note

A newly added column reads null for every row landed before the column existed. Data Hub does not resend history, so only a re-initial sync backfills those values.

Data Hub infers a field's type from sampled record values rather than a declared schema. If every sampled value for a newly added column is null, Data Hub reports no type, and the column is coerced to STRING in your target regardless of the type declared in the model.

Where a change cannot be handled safely, only the affected model is blocked. Other models in the same flow keep streaming.

Expected behavior: golden-record metadata fields

The Data Integration target writes the complete record on every change, built entirely from the event-stream message. There is no merge with prior state. This makes the following field behaviors important to understand.

Record creation date

createdAt on a CDC message is the time the event was created, not the time the golden record was created. recordUpdatedAt is the time the record itself was last updated. Because Data Integration writes the full record on every change, the value landed in your target is taken from the event and reflects the update time. A record created months ago but recently updated shows the recent date in your target, even though the Data Hub Repository API still returns the original creation date for that record.

Use recordUpdatedAt, not createdAt, to reason about when a record changed. The original golden-record creation date is available from the Data Hub records query API, not from the CDC stream.

Record title

recordTitle is not delivered through the Data Hub source connector, in either standard extraction or CDC. This is by design: a record title can contain a field value that was deliberately excluded from a source's delivery configuration, so Data Hub disallows recordTitle in delivery by default. If you need a record title, derive it from the model fields the connector already delivers.

Mapping tab for a CDC-enabled model showing recordId, createdDate, updatedDate, and recordTitle offered and checked as source columns

Record identifier: recordId versus id

In CDC, use recordId, which is the top-level identifier on the message envelope as the golden record identifier when mapping to a target identifier column. The id element inside the record payload itself is not delivered in either extraction mode, because it is not visible to Data Integration's metadata discovery. recordId is present and consistent in both standard extraction and CDC.

note

For a source that only reads from Data Hub, recordId is the golden record identifier and mapping it to your target's identifier column covers the need completely. For a source that also contributes to the golden record, recordId is still the golden record identifier, not the identifier that source itself contributed. That value is available separately in Data Hub and is not part of this connector's delivery.

Summary

FieldStandard extractionCDC
Golden record creation dateAvailable from Data HubNot available. createdAt is the event time
recordTitleExcludedExcluded
id (record payload element)Not deliveredNot delivered
recordIdAvailableUse this as the identifier

CDC errors and edge cases

ScenarioBehavior
CDC enabled without Event Streams URL and tokenActivation blocked; invalid fields are focused
Invalid or missing topicActivation blocked
Token lacks consume permissionActivation blocked
Non-XML payload receivedData flow fails
channelUpdateType of DIFF receivedData flow fails; Data Integration only supports FULL
Unsupported operation receivedEvent is logged and triggers an internal alert. Not counted, not shown in the run
Schema change that cannot be handled safelyOnly the affected model is blocked; other models keep streaming
Offset no longer availableData flow fails; a re-initial sync is required

Known limitations

  • The golden record's original creation date is not carried over the CDC stream. createdAt reflects the event time, not the record's creation time.
  • recordTitle is not available through the connector in either extraction mode.
  • Only a channelUpdateType of FULL is supported. A channel configured to send DIFF causes the data flow to fail.
  • A newly added column reads null for history; only a re-initial sync backfills it. Depending on how source metadata is refreshed, a newly added field may not reach the target until you reload metadata from the Source tab.
  • Data Hub infers column types from sampled record values. If every sampled value for a new column is null, Data Hub reports no type, and Data Integration coerces the column to STRING, regardless of the type declared in the model.
  • A removed field's column is retained, not dropped, and simply stops being populated.
  • After promoting a data flow to another environment, the original Data Hub source remains on your account and is not cleaned up automatically.
  • Maximum message size is 1 MB per event.
  • Delivery is at-least-once; duplicates are possible and are resolved by Merge loading with a stable primary key.
On this Page