{
  "$type": "site.standard.document",
  "bskyPostRef": {
    "cid": "bafyreif46wwmu2uu4xe7pmltzsy3ylbwkuvyh6m5tiimjyfnmqysrdea2q",
    "uri": "at://did:plc:pgryn3ephfd2xgft23qokfzt/app.bsky.feed.post/3mowybt5yxto2"
  },
  "path": "/t/seeking-indic-document-dataset-india-invoices-receipts-utility-bills-payment-advices-packing-lists-commercial-invoices-credit-notes/177055#post_5",
  "publishedAt": "2026-06-23T07:09:40.000Z",
  "site": "https://discuss.huggingface.co",
  "textContent": "masonreed11:\n\n> Public datasets can help, but for enterprise-grade Indic document AI, custom annotation or synthetic data generation is often needed. Quality, consistency, and licensing matter more than just having a large dataset.\n\nCompletely agree — quality, consistency, and licensing are non-negotiable for us, which is exactly why we’re not relying on public datasets alone. We’re currently evaluating both custom annotation vendors and synthetic generation pipelines. If anyone has specific vendor recommendations with Indic script experience, we’d love to hear them!",
  "title": "[SEEKING] Indic Document Dataset (India) — Invoices, Receipts, Utility Bills, Payment Advices, Packing Lists, Commercial Invoices, Credit Notes"
}