NetSuite CSV import reconciliation: why your retry window should be time-based
- Published on
- -6 mins read
- Authors
- Name
- Andy Nur (Andy)
- @andynur
Retry logic is usually written as "try up to N times". That works when the time between tries is predictable. When it isn't, "N times" can mean several minutes on a busy day and about a minute on a quiet one.
In this article, I'll walk through a bug in a NetSuite integration where exactly that happened, and the small change that fixed it. Client details are anonymized and the code is simplified.
We'll cover:
- How the integration imports journals with CSV import tasks
- Why the import needs a reconciliation step at all
- How a self-chaining Map/Reduce used up its retries in about a minute
- Switching to a time-based retry window with a hard cap
Prerequisites
You should know N/task (especially CsvImportTask and task status), Map/Reduce scripts, and saved searches. Some integration background helps. My article on NetSuite concurrency limits and retries covers retry basics.
The setup
An external gateway drops journal files on SFTP. On the NetSuite side, a pipeline turns each file into a CSV import task that creates Journal Entries. A monitor Map/Reduce watches the running import tasks. When it finishes a pass, it queues itself again (it self-chains) as long as there is work left.
When an import task completes, the monitor records how many journals were imported and how many failed, and marks the job done. The whole pipeline is covered in Importing ETL and payroll journals into NetSuite with CSV import tasks.
Why reconciliation is needed
Here's the catch: for CSV imports, task.checkStatus() tells you the task is COMPLETE, but not how many rows were imported. So the service counts them itself, by searching for the Journal Entries that the import created.
Right after a large import, that search can come back short or empty. The records exist, but the search index hasn't caught up yet. So the service treats "counts unknown" as "try again on the next monitor pass", and gives up after a limit.
The bug
In a quiet Sandbox, jobs kept finishing as unresolved, even though every journal had been imported correctly.
How 'retry 5 times' became 'retry for about a minute'
Step 1 / 51. Task COMPLETE
The CSV import task has finished. The monitor asks the service to count the imported journals.
On a busy account, other work slows the self-chain down, so the same attempt count covered several minutes. On a quiet account, nothing slowed it down, and the whole budget was spent in about a minute. The index needed longer than that. The retry limit was really measuring how busy the account was.
The fix
Store when the first unresolved attempt happened, and keep retrying until a fixed amount of wall-clock time has passed. The attempt count stays, but only as a hard safety cap.
const RECONCILE_WINDOW_MS = 5 * 60 * 1000 // retry for 5 real minutesconst RECONCILE_MAX_ATTEMPTS = 30 // safety cap only
function handleUnresolved(job) { const state = job.resultRef || {} const attempt = (state.reconcileAttempts || 0) + 1 const firstAttemptAt = state.reconcileFirstAttemptAt || Date.now() const elapsedMs = Date.now() - firstAttemptAt
if (elapsedMs < RECONCILE_WINDOW_MS && attempt < RECONCILE_MAX_ATTEMPTS) { saveResultRef(job.id, { ...state, reconcileAttempts: attempt, reconcileFirstAttemptAt: firstAttemptAt, }) return 'RETRY_NEXT_PASS' }
log.audit('Reconcile retries exhausted', { jobId: job.id, attempt, elapsedMs }) return 'FINALIZE_UNRESOLVED'}Two details matter here:
- The timestamp is stored, not kept in memory. Each monitor pass is a new script run. The first-attempt time lives in the job's stored state, so it survives across passes.
- The hard cap stays. If something goes wrong and the window never closes, 30 attempts still stops the loop.
When to use time, when to use attempts
Count attempts when each attempt is expensive and the delay between them is fixed. Use a time window when you're waiting for something else to catch up, like a search index, a queue, or an external system.
Testing it
The tricky part is testing behavior that depends on time. The Jest tests use fake timers (jest.useFakeTimers().setSystemTime(...)) and a stored first-attempt time, so each case can sit at a known point in the window. They check that:
- An unresolved first check defers finalization to the next monitor pass
- Retries continue inside the window, even after many attempts
- The job finalizes as unresolved once the window has passed
Conclusion
The bug looked like a search problem, but the search was fine. The retry policy assumed a pace that the self-chain didn't keep.
If your retries wait for something to catch up:
- Measure the window in time, and store when it started
- Keep an attempt count only as a safety cap
- Test with a fake clock, so time is part of the test input