HubSpot to Snowflake
Load HubSpot contacts, companies, deals and engagements into Snowflake to join CRM activity with product and billing data.
The pipeline as it appears on the Pipeloom canvas. You can add transforms or triggers on the same canvas later.
What to know about HubSpot into Snowflake
Set a lookback window for property history
The contacts, companies and deals property_history streams can miss records because HubSpot's calculated properties always carry the latest sync time. Set Property History Lookback Window, for example 43200 (30 days), and let Append + Deduped absorb the overlap.
Your HubSpot tier sets the speed limit
The connector shares HubSpot's rate limits: 100 requests per 10 seconds on Free and Starter, 190 on Professional and Enterprise, with daily caps of 250,000 to 1,000,000 per account. Many custom properties lengthen syncs, so schedule large initial loads off-peak.
Keep engagement syncs frequent
The engagements stream is fast when the last sync was under 30 days ago and fewer than 10,000 records are new. After a longer gap it reads everything and can take far longer, so a regular schedule is the cheaper option.
Custom objects need one extra scope
Custom CRM objects appear as streams once the app has the crm.objects.custom.read scope and you refresh the connection's source schema. Then they load into Snowflake like any other table.
Set it up
- Create the HubSpot sourceAuthenticate with A Private App access token or a Service Key. You need: A HubSpot account. A Private App (or Service Key) with the scopes for the streams you want, for example CRM read scopes for contacts, companies and deals.
- Create the Snowflake destinationAuthenticate with Key pair (an RSA key of 2048 bits or more). You need: A Snowflake account and the ACCOUNTADMIN role for the one-time setup. A warehouse, database, user and role for Pipeloom, with a key pair attached to the user. Network policies that allow Pipeloom to connect.
- Connect them and pick streamsDraw the edge on the canvas and choose streams and sync modes. HubSpot supports full refresh and incremental; Snowflake supports every mode: full refresh overwrite, overwrite deduped, append, incremental append and incremental append deduped.
- Run, then scheduleThe first sync loads each stream in full. Then set the schedule.
What lands in Snowflake
- A final table per stream with typed columns, plus _AIRBYTE_RAW_ID, _AIRBYTE_GENERATION_ID, _AIRBYTE_EXTRACTED_AT, _AIRBYTE_LOADED_AT and _AIRBYTE_META.
- A raw table per stream in the airbyte_internal schema (or the internal dataset name you set), holding each record as JSON.
Type mapping
- object → OBJECT
- array → ARRAY
- union and unknown → VARIANT
- timestamp with time zone → TIMESTAMP_TZ
- integer → NUMBER
- number → FLOAT (or NUMBER(38,9))
When HubSpot's schema changes
New columns are added and types changed as the source changes. A column removed at the source is kept, with its data, and new rows get NULL. Column names that are SQL reserved words get a leading underscore.
Streams you can sync
Includes Contacts, Companies, Deals, Deal Pipelines, Contact Lists, Engagements, Tickets, Forms and more. The full list is in the reference.
What it costs
A sync is metered at 1 credit per vCPU-minute, and a default sync uses 2.5 vCPUs. As an example, not a benchmark, a sync that takes 5 minutes uses 13 credits. Run 24 times a day, that is about 9,360 credits a month, more than Free's 1,500 but within Starter's 12,000 ($19 a month). Your own sync time decides the real figure, so try the calculator with your numbers.
Questions and fixes
Which HubSpot credentials should I use?
A Private App access token or a Service Key. Add scopes only for the streams you need.
Why is my HubSpot sync slow?
Custom properties make syncs longer, and the connector shares HubSpot's rate limits. Burst limits are 100 requests per 10 seconds on Free and Starter and 190 on Professional and Enterprise, with daily limits per account from 250,000 to 1,000,000.
Why is the engagements stream slow after a gap?
If the last sync was within 30 days and fewer than 10,000 records are new, the connector uses HubSpot's recent-engagements API. Otherwise it reads all engagements, which is much slower. Syncing more often keeps it on the fast path.
Why are records missing from the property history streams?
HubSpot calculated properties (formulas, rollups and hs_analytics_* fields) carry the time of the latest sync, which can move the cursor past unsynced records. Set Property History Lookback Window in the source, for example 43200 (30 days). Those streams use Append + Deduped, so the overlap does not create duplicates.
Why are custom objects or a stream like workflows missing?
Custom objects need the crm.objects.custom.read scope, then Refresh source schema on the connection. Streams such as workflows are skipped, with a warning in the logs, when the app lacks their scope.
Why does the sync fail with "Current role does not have permissions on the target schema"?
The Pipeloom role is missing a grant on the schema it writes to. Re-run the role setup from the Snowflake guide, or grant the role usage and create-table rights on that schema.
Sync HubSpot to Snowflake today
Free forever, no card. Starter is $19 a month for three seats.