Skip to content

Cursors and Feeds

A source is rarely a single stream. A Substack source follows several publications; an RSS source follows many feeds; a Twitter source follows bookmarks, the home feed, and individual profiles. The feeds model gives each of these its own identity, cursor, and lifecycle — without forcing the user to create a separate source per feed.

This page covers the data model (source → feed → document), the cursor types a cursor can carry, and the compare-and-swap protocol that keeps cursors correct when syncs overlap. It reflects the architecture in the V1 feeds/schema design.

Data is organized in three levels:

source → feeds → documents
  • Source — the top-level “thing you connected.” Holds the type, user preferences (no credentials), and overall status. One row per connected thing.
  • Feed — a specific publication, section, channel, or profile within the source. Has its own cursor, document count, status, and error state.
  • Document — belongs to a feed (and transitively to the source). Deduplicated by (feed, externalId).

A single-feed source (one blog) has exactly one feed with externalKey: "default". A multi-feed source gets one feed per publication.

Substack with three publications
source s1 type: substack name: "Substack" config: {}
feed f1 externalKey: "https://stratechery.com" name: "Stratechery" cursor: { lastDate: "2026-03-20T…" }
feed f2 externalKey: "https://platformer.news" name: "Platformer" cursor: { lastDate: "2026-03-22T…" }
feed f3 externalKey: "https://thegeneralist.com" name: "The Generalist" cursor: { lastDate: "2026-02-15T…" }

Per-feed cursors are why you can pause one publication, see per-publication counts, and add a fourth feed later without its backlog being silently skipped by a source-wide cursor.

A sync client ensures a source, then ensures each feed within it:

  1. createSource — once per connected thing.
  2. addFeed — once per feed. Pass an externalKey (the feed URL, username, or section) and a display name. A single-feed source uses externalKey: "default".
  3. ingestDocuments — every sync, passing both the sourceId and the feedId.

Trove’s cloud runner does this for you for cloud-run sources, and the Mac app does it for sources that run there; if you push documents yourself via the Sync API, you do it explicitly.

A cursor is an opaque string stored on the feed (Feed.cursor). Trove parses it into a cursor that describes how to resume — and the SDK surfaces exactly that typed Cursor on ctx.cursor (a tagged union, never a bare string). Three cursor types cover the common cases:

TypeCursor valueWhen to useResume strategy
dateAn ISO 8601 timestamp, e.g. "2026-06-14T09:00:00Z"The upstream exposes content in time order and supports “since <date>” filtering (RSS, most APIs)Fetch items newer than the cursor; the newest date you push becomes the next cursor.
idSetThe set of ids already stored, in values, bounded by an optional max retention capThe upstream has no reliable date filter, so resuming means remembering what you haveFetch until you reach ids already in the set, then carry every id you saw forward. max bounds the set — it is a count, never an id.
nonenullThe feed is small enough to re-fetch in full every run, or you rely purely on dedupRe-fetch everything each run; (feed, externalId) dedup skips what’s already indexed.

Choosing a cursor is a design decision for the source adapter: pick the cheapest one the upstream API supports. A none cursor is always correct (dedup makes it safe) but re-fetches more; a date cursor is the most efficient when the upstream supports it.

A date-cursor sync
async sync(ctx) {
// ctx.cursor is a Cursor tagged union — read the date variant's value.
const since = ctx.cursor.type === "date" ? ctx.cursor.value : "1970-01-01T00:00:00Z";
const res = await ctx.fetch(`${api}?updated_after=${since}`);
const items = await res.json();
const documents = items.map((i) => ({ id: i.id, text: i.body, date: i.published }));
// Advance the cursor: the newest item's date becomes the next cursor.
const newest = documents.reduce((max, d) => (d.date > max ? d.date : max), since);
return { documents, cursor: { type: "date", value: newest } };
}

The cursor is owned by the cloud (stored on feeds.cursor) but advanced by whoever runs the sync — Trove’s cloud runner, the Mac app, or your own Sync API code — sent on each ingest. Because runs can overlap (a cloud run and a manual sync, two machines, or an eager launch-sync racing a timer), ingestDocuments uses compare-and-swap to stay correct.

The contract, per sync:

  1. Read the current cursor from Feed.cursor before syncing (ctx.cursor gives you this).
  2. Fetch documents newer than that cursor position.
  3. Push with ingestDocuments(sourceId, feedId, documents, cursor, cursorBefore):
    • cursor — the new position you’re advancing to.
    • cursorBefore — the value you read in step 1 (the expected current value).
  4. The server verifies cursorBefore matches the stored cursor. If it doesn’t, another sync ran concurrently: the server still ingests the documents (dedup handles any overlap) but keeps the more-advanced cursor and logs a warning.
  5. The server advances feeds.cursor only if the new cursor is monotonically ahead of the stored one. Cursors never regress.
  6. The server returns IngestResult with the final cursor.
mutation {
ingestDocuments(
sourceId: "src_abc123"
feedId: "feed_def456"
cursor: "2026-06-14T09:00:00Z" # advancing to here
cursorBefore: "2026-06-13T21:00:00Z" # what we read before syncing
documents: [ /* … up to 50 … */ ]
) {
documentsIndexed
documentsSkipped
cursor # final, authoritative cursor
}
}
  • Idempotent retries. If the client crashes after the server ingests but before receiving the response, retrying the batch is safe — (feed, externalId) dedup skips the already-indexed documents, and the client reads the already-advanced cursor on retry rather than re-fetching from the old position.
  • Concurrency safety. Two devices syncing the same feed race to push. Dedup prevents duplicate documents; monotonic cursor advancement prevents the cursor from regressing to an earlier device’s position. Redundant fetching is the only cost.
  • No lost backlog. Because the cursor only moves forward and only when ahead, a slow sync can’t stomp a faster one’s progress.

Send up to 50 documents per call; page through larger backlogs across multiple ingestDocuments calls, advancing the cursor each call.

A source’s sync(ctx) runs in one of three places, and the cursor protocol above keeps it correct in every one:

  • Trove’s cloud (the default). For public feeds and keyless APIs, Trove schedules the sync, runs sync(ctx), and ingests the result — no app required, your Library stays fresh on its own.
  • The Trove Mac app. Sources that need a real browser, local files, or your own network run on your Mac (badged “Runs on your Mac”); you can also choose to run any cloud-eligible source there. The app owns per-feed scheduling (sleep/wake, eager sync on launch, battery-aware throttling) and holds credentials in the macOS Keychain, never sent to the cloud.
  • Your own code. Drive the Sync API directly — run sync(ctx) wherever you like and push with ingestDocuments.

Whichever runs it, the cloud is the authoritative recorder: it stores cursors, document counts, and sync runs, and runs the ingest pipeline (dedup, chunking, embedding, storing). The executor reads the cursor, fetches what’s new, pushes, and records a sync run for the queryable audit trail — and the compare-and-swap protocol above is exactly what lets a cloud run and a manual sync (or two machines) overlap safely. On a failure, retrying the whole batch is safe: dedup skips what’s already indexed.

Executor (Trove cloud, the Mac app, or your code) Cloud (recorder + ingest)
1. sync is due for feed f1
2. read feed cursor (cursorBefore)
3. run sync(ctx) → documents
4. ingestDocuments ────────────────────────────► verify cursorBefore (CAS)
{ sourceId, feedId, dedup by (feed_id, external_id)
documents, cursor, cursorBefore } store text in R2, embed, index
◄──────────────────────────── advance cursor if ahead
5. record the sync run (audit) return IngestResult