The batch model

Everything is a batch, including a batch of one.

On this page

Every call that touches LinkedIn is a batch. Submitting returns 202 and a batch id; reading is a separate, free call. That holds whether you sent one target or fifty thousand.

It still holds under /sync, which creates the same batch and drains it inside your request. What you send there is one target and what comes back is one result, because that path answers a single entity - but the batch under it is real: the same id, the same entry read through the same cursor, the same line in the usage rollup. That is the whole reason the synchronous path is an addition rather than a second product - see Synchronous requests.

The unit is a batch, never a job - job already means a LinkedIn job posting in this API, and a docs page cannot survive both meanings.

Why a batch of one is still a batch #

Because the alternative is two result shapes, two cursors and two failure vocabularies. One shape, learned once, holds from 1 to 50,000. It is also why the synchronous path is not a batch of one: an endpoint that answers a single entity should say so in the body rather than hand you an array to unwrap, and the entry inside its result is the entry the cursor returns, field for field.

Lifecycle #

queued -> processing -> completed | failed | cancelled

failed is reserved for a batch that could not run at all. A batch in which every single item failed is still completed, with succeeded: 0. That distinction is what lets you treat completed as "results are final" without a special case for the all-failed shape.

Results stream #

GET /batches/{id}/results returns entries as they land, while status is still processing. A 50,000-profile batch is consumable from the first second rather than after the last row.

Shell
cursor=""
while : ; do
  page=$(curl -s "https://api.easydata.win/v1/batches/$BATCH/results?cursor=$cursor" -H "X-API-Key: $API_KEY")
  echo "$page" | jq -c '.data.entries[]'
  [ "$(echo "$page" | jq -r '.meta.hasMore')" = "true" ] || break
  cursor=$(echo "$page" | jq -r '.meta.nextCursor')
  sleep 2
done

The cursor is monotonic and stable. Walking it to exhaustion yields every entry exactly once - no duplicates, no gaps - even across a batch that is still filling in behind you.

Entries carry their input #

Every entry echoes item_index and your original input, so you join back to your own rows without assuming an order. Order is never guaranteed: items are claimed by whichever worker in the pool is free, and a slow target does not hold up the ones behind it.

Failure is per item #

One unreachable profile does not fail a batch. It becomes an entry with status: "failed", an error.type, and credits_used: 0.

Paged operations #

A Sales Navigator search emits one entry per upstream page of up to 100 rows, with page set, rather than one enormous entry per target. That is also the metering unit: each page is charged for the rows it actually returned, as it lands.

This is why the counters mean two different things. total, succeeded, failed and pending count TARGETS and always sum to total. results_available counts delivered ENTRIES, which for a paged operation is a much larger number.

Idempotency #

Send an Idempotency-Key header on any create. A retried submission returns the original batch_id rather than duplicating the work - which matters most in exactly the case you cannot see, where our response was lost on the way back to you.

Shell
curl -X POST "https://api.easydata.win/v1/profiles/enrich" \
  -H "X-API-Key: $API_KEY" \
  -H "Idempotency-Key: 6b1f0e2a-3c4d-4e5f-8a9b-0c1d2e3f4a5b" \
  -H "Content-Type: application/json" \
  -d '{"targets": ["williamhgates"]}'

The key is checked before the body is validated. A retry of something we already accepted returns the original batch even if your client has since started sending something we would now reject.

Cancelling #

POST /batches/{id}/cancel stops items that have not started. Everything already delivered stays delivered and stays readable, and anything in flight finishes rather than being thrown away half-scraped - you were going to be charged for that fetch either way.

Exports #

Add format=jsonl for one entry per line, or format=csv for a flat file. The CSV carries identity, outcome and cost as columns and leaves the record whole in a JSON cell: flattening a Profile into columns would mean choosing a shape every caller then works around.

Neither format has an envelope, so the cursor comes back in an X-Next-Cursor header rather than in meta, and the server echoes back the cursor you sent once a page holds no entries - which is how you know you have reached the end. Every CSV page carries its own header row: keep the first and drop the rest, or a file most tools refuse halfway through is what you get. The official clients page and de-header it for you, as ed.export(batch_id, format="csv").