Deletion & blacklist¶
POST /v1/forget purges content and blacklists it in a single atomic
operation. Blacklisting matters because it tells upstream capture to stop wasting
work — a subsequent POST /v1/pages for a blacklisted pattern returns 403
rather than a silent 202.
Forget¶
curl -s -X POST http://127.0.0.1:8000/v1/forget \
-H "Authorization: Bearer $REFINDERY_AUTH_TOKEN" \
-H "Content-Type: application/json" \
-d '{"domain": "example.com"}' # or {"url": "https://example.com/page"}
The purge:
- deletes the page row (cascading chunks, mentions, and cluster memberships);
- deletes the page's vectors from every model's store;
- recomputes entity counts and garbage-collects orphaned entities;
- marks affected clusters stale.
Vector deletes are queued as PURGE_VECTORS jobs and reconciled by a periodic
verify_tombstones task (the tombstone moves pending → deleted → verified),
so a store that is briefly unavailable still converges. See
Operations.
Purge is irreversible
Forgetting deletes content permanently. Removing a blacklist rule later does not restore purged pages.
Blacklist management¶
A forget adds a blacklist rule — an exact canonical URL or a domain suffix. Manage rules directly:
GET /v1/blacklist list rules
DELETE /v1/blacklist/{id} remove a rule (un-blacklist; does not restore content)
write scope is required for forget and un-blacklisting. See
Authentication.
Related¶
- Upstream ingest API — full contract and the
403blacklisted response. - Ingesting pages — where the blacklist check happens.
- Operations — tombstone reconciliation and lease model.