Skip to main content

Overview

Collections run against external systems — databases, APIs, and file stores — that occasionally return temporary errors. To keep collections resilient, Entegrata automatically retries errors it recognizes as transient and fails fast on errors it recognizes as permanent. This happens behind the scenes, so most temporary glitches never result in a failed job. This page explains, at a high level, what gets retried and what happens when it doesn’t.

Transient vs. permanent errors

When a collection step hits an error, Entegrata classifies it before deciding what to do.

Transient (retried)

Temporary conditions that usually clear on their own:
  • Rate limiting (too many requests)
  • Service temporarily unavailable
  • Expired authentication (a re-login is attempted automatically)
  • Request timeouts and deadlines
  • Network, DNS, or TLS interruptions
  • Database deadlocks and short-lived consistency lags

Permanent (fail fast)

Conditions that won’t improve by trying again:
  • Bad request or invalid configuration
  • Forbidden / insufficient permissions
  • Resource not found
  • Unrecoverable errors from the source system
Classification can be tuned per connector, so a small number of data sources may treat specific errors differently from the defaults above.

How retries work

When an error is transient, Entegrata retries the operation automatically:
  • Multiple attempts. Collection steps retry up to several times before giving up.
  • Exponential backoff with jitter. The wait between attempts grows with each retry, and a small random offset is added so many concurrent retries don’t hammer the source system in sync.
  • Respecting the source system. When a source returns a rate-limit response with a Retry-After hint, Entegrata honors that wait (up to a reasonable cap) instead of retrying immediately.
Retries are applied at each stage of collection independently — fetching data from the source, and loading it into the Lakehouse each have their own retry handling.

When retries are exhausted

If an operation keeps failing after all retry attempts, the job — or the specific resource within it — is marked as Failed, and the error reason is recorded on the job so you can review it. Two outcomes are treated as not failures:
  • Clean early stops. When a source signals there’s no more data to collect (for example, the last page of an API), the job finishes normally.
  • Resizing. If a job needs more capacity than its assigned worker, it is automatically re-queued onto a larger worker rather than failing.

Partial failures

Collection is tracked per resource. If a connection collects several resources and one of them fails, the others continue and report their own outcomes — a single failing resource does not abort the rest. Entegrata also validates row counts when loading data. If the number of records loaded into the Lakehouse doesn’t match what was collected (after an automatic refresh retry), that resource is failed rather than reported as a partial success. This prevents a job from silently completing with incomplete data.
A job that failed after retries will show its final error reason in the job detail view. Include the Job ID when contacting support so the execution can be traced.

Job Statuses

Understand job statuses and what they mean

Troubleshooting

Diagnose and fix collection job issues

View & Filter Jobs

View and filter collection jobs

Overview

Understand collection jobs and performance monitoring