{
"$type": "site.standard.document",
"bskyPostRef": {
"cid": "bafyreihyg4j4ul57dppgpzof4tgppj66yjjwih7ilxg554kohf4rapfusi",
"uri": "at://did:plc:pgryn3ephfd2xgft23qokfzt/app.bsky.feed.post/3mpumhahl4w42"
},
"path": "/t/is-selling-datasets-way-harder-than-building-them-or-is-it-just-me/177440#post_2",
"publishedAt": "2026-07-05T01:35:09.000Z",
"site": "https://discuss.huggingface.co",
"tags": [
"Data Appraisal Without Data Sharing",
"Try Before You Buy",
"(click for more details)"
],
"textContent": "It really does seem that selling datasets comes with all kinds of difficulties…\n\n* * *\n\n## Short answer\n\nNo, I do not think you are imagining the problem.\n\nA buyer needs enough evidence to decide whether a dataset is worth purchasing, but sufficiently detailed inspection may disclose much of what is being sold. Research such as Data Appraisal Without Data Sharing and Try Before You Buy starts from essentially this dilemma.\n\nHowever, I would avoid treating every stalled transaction as a single “dataset quality” problem. At least four different uncertainties may be involved:\n\n 1. **Is the dataset technically and substantively sound?**\n 2. **Can it legally and operationally be used?**\n 3. **Will it help this buyer’s particular model and product?**\n 4. **Who is responsible after payment if something is wrong or difficult to use?**\n\n\n\nAn independent report could help substantially with the first two, provide limited evidence for the third, and support the fourth only when it is connected to clear contractual responsibilities.\n\nMy current view would be:\n\n> A useful product is probably not one universal “quality certificate.” It is a version-specific evidence package, a limited path for buyer-specific evaluation, and explicit acceptance, correction, support, and refund terms.\n\nThere are also important things that remain unknown. I have not found public evidence showing that ML-dataset buyers on the Hugging Face Hub will pay for a generic certification service, or that a certificate materially increases completed sales. I have also not found public statistics on HF dataset sales, failed purchases, refunds, or disputes. Those may be central unknowns rather than minor missing details.\n\n## A practical way to separate the problem\n\nIf the transaction is blocked by… | The most relevant response is likely to be…\n---|---\nBroken files, schema problems, duplicates, or label errors | A reproducible technical and sampled-content audit\nUnclear source or commercial rights | Provenance records, license evidence, and legal review\nUncertainty about model improvement | A fixed reference experiment or buyer-specific pilot\nA buyer without a stable training/evaluation pipeline | Reference code, an evaluation harness, and integration support\nProcurement or internal accountability | A named seller, contract owner, acceptance criteria, and support terms\nFear of abandonment after payment | Correction, update, response-time, and refund commitments\nUnclear price relative to value | A paid pilot, staged payment, or narrowly scoped license\nA marketplace reselling third-party data | Explicit creator, seller, support, correction, and refund roles\n\nSo, before building a certification service, I would first identify which branch is actually stopping transactions. A quality audit will not solve an internal procurement problem; a better contract will not prove model lift; and a successful benchmark result will not settle provenance.\n\n## A small starting point\n\nA relatively low-risk first experiment could be:\n\n 1. Select one kind of dataset and one target buyer profile.\n 2. Fix one exact Hub revision or file manifest.\n 3. Create a public **Evidence Pack** for that revision.\n 4. Offer three or four evaluation paths:\n * evidence only;\n * a representative gated sample;\n * a fixed-budget seller-side pilot;\n * delivery with objective acceptance tests.\n 5. Record which path moves a buyer to the next concrete step.\n\n\n\nThe result to measure is not only model accuracy. It is also whether the buyer proceeds to a pilot, legal review, pricing discussion, purchase approval, or stops for a clearly identified reason.\n\n1. Why “dataset quality” should be divided into several layers (click for more details) 2. Why “good for fine-tuning” is the hardest claim (click for more details) 3. What can be shown before purchase without transferring the whole asset? (click for more details) 4. What could an independent report credibly say? (click for more details) 5. Why the transaction design may matter as much as the audit (click for more details) 6. Why third-party marketplaces make this much harder (click for more details) 7. A small pilot before building a general certification service (click for more details)\n\n## Important unknowns\n\nThe following seem important enough to remain visible rather than being hidden inside a general conclusion:\n\n * Public evidence of B2B willingness to pay for generic ML-dataset certification is limited.\n * Public HF statistics on dataset sales, refunds, disputes, and failed purchases do not appear to be available.\n * A certificate’s effect on completed transactions is not established.\n * It is unclear whether buyers prefer generic audits, representative samples, buyer-specific pilots, or stronger contractual recourse.\n * A benchmark improvement does not establish buyer-specific value.\n * A complete Dataset Card does not establish that its claims are true.\n * A gated dataset controls access but does not provide payment, escrow, acceptance, refund, or recovery of downloaded copies.\n * Protected evaluation can reduce disclosure but creates new security, cost, and trust questions.\n * Legal usability cannot generally be reduced to the presence of a license tag.\n * A marketplace cannot reliably promise corrections unless someone in the responsibility chain has the source material, authority, incentive, and resources to perform them.\n\n\n\n## Direct answers\n\n### Is this a real problem?\n\nYes. The conflict between buyer inspection and seller protection is recognized in data-market and data-valuation research.\n\n### What may be delaying deals?\n\nQuality uncertainty is one plausible cause, but it should be separated from provenance, licensing, pricing, procurement, format, integration cost, buyer readiness, support, and organizational accountability.\n\n### How can trust be built before purchase?\n\nThrough layers:\n\n * public documentation and statistics;\n * version-bound audit evidence;\n * representative limited disclosure;\n * reproducible reference experiments;\n * buyer-specific evaluation where justified;\n * objective post-delivery acceptance tests;\n * explicit support and correction terms;\n * visible responsibility for third-party data.\n\n\n\n### Would an independent report help?\n\nProbably, especially for:\n\n * technical integrity;\n * sampled content quality;\n * duplicate measurements;\n * documentation and provenance evidence;\n * reproducibility under stated conditions;\n * supplier ownership and correction processes.\n\n\n\nIt should not imply universal model improvement or universal legal safety.\n\nA defensible description might be:\n\n> A version-specific, attributable verification report stating what was tested, how it was tested, what was not tested, and what happens when the dataset changes or a defect is found.\n\n## Overall\n\nSelling may genuinely be harder than building because the buyer is purchasing much more than rows or files:\n\n * provenance and usage rights;\n * quality claims;\n * uncertain future model utility;\n * integration work;\n * documentation;\n * support;\n * correction capability;\n * continuity;\n * and a responsible counterparty.\n\n\n\nSo the main design question may be less:\n\n> Can this dataset receive one universal quality certificate?\n\nand more:\n\n> For this exact dataset revision, what evidence can be produced before purchase, what remains buyer-specific, how can the remaining uncertainty be tested, and who is responsible if the data, integration, or transaction fails?\n\nReferences and reusable building blocks (click for more details)",
"title": "Is selling datasets way harder than building them? Or is it just me?"
}