{
"$type": "site.standard.document",
"bskyPostRef": {
"cid": "bafyreibcxmiyjlxqds4bsanynqgtdghb5mqr23uouxq6uaiuk6tg5blq4i",
"uri": "at://did:plc:pgryn3ephfd2xgft23qokfzt/app.bsky.feed.post/3mpzu7mhibx42"
},
"path": "/t/the-case-for-an-nvc-annotated-dataset/177505#post_2",
"publishedAt": "2026-07-07T05:14:25.000Z",
"site": "https://discuss.huggingface.co",
"tags": [
"PersonaConflicts / Words Like Knives",
"EMNLP paper",
"(click for more details)"
],
"textContent": "For now, I found a few related resources that might be useful:\n\n* * *\n\nI did not find an exact duplicate of the proposed resource, but I found several projects that cover different parts of it: NVC-informed conflict annotation, empathy and counseling schemas, paired rewriting, recipient-side evaluation, conversational sycophancy, and the practical mechanics of running a small community annotation project on the Hub.\n\nThe closest starting point I found is PersonaConflicts / Words Like Knives, which provides 5,772 simulated dialogues between familiar social partners, relationship backstories, and a human-annotated subset using NVC-informed communication-breakdown labels. The accompanying EMNLP paper may be useful for comparing schemas and identifying what a new dataset would add beyond synthetic conflict dialogues.\n\nA practical default route might be:\n\n 1. use existing conflict and empathy data to test a draft annotation schema;\n 2. distinguish directly observable text from inferred feelings or needs;\n 3. retain evidence, uncertainty, multiple interpretations, and unaggregated annotations;\n 4. evaluate extraction, rewriting, recipient perception, de-escalation, and sycophancy separately;\n 5. release the schema, annotation guide, difficult examples, and a small pilot before committing to large-scale annotation.\n\n\n\nThe smallest useful first release may therefore be a dataset repository containing:\n\n * a draft annotation guide;\n * 25–100 independently annotated examples;\n * raw annotations as well as any adjudicated version;\n * examples of uncertainty and disagreement;\n * a small controlled evaluation set;\n * a Dataset Card documenting provenance, intended use, and limitations.\n\n\n\nOne distinction seems particularly useful from the beginning: a feeling or need explicitly stated by the speaker is a different kind of label from a feeling or need inferred by an annotator. For inferred fields, it may help to preserve supporting text, confidence, alternative interpretations, and an `insufficient context` option rather than requiring one definitive answer.\n\nClosest adjacent resources (click for more details) NVC-specific schema ideas (click for more details) Annotation workflow and disagreement (click for more details) Rewriting and controlled comparisons (click for more details) Evaluation map and controls (click for more details) Possible pilot routes (click for more details) Hub implementation, governance, and reproducibility (click for more details)\n\nIn short, many of the pieces needed to test the idea already exist without first building the full corpus. PersonaConflicts provides a close conflict-data starting point; EPITOME, AnnoMI, and ESConv provide annotation-process precedents; CNVC materials provide distinctions beyond OFNR; SycophancyEval and BenSyc provide different controls for agreement and support; and the Hub provides repository, documentation, discussion, versioning, and annotation tooling.\n\nA particularly reusable first contribution might be a small public pilot preserving:\n\n * explicit versus inferred content;\n * evidence and uncertainty;\n * multiple Need candidates;\n * raw annotator disagreement;\n * clarification as an alternative to guessing;\n * meaning preservation in rewrites;\n * separate evaluation of support, recipient perception, de-escalation, and factual stability.\n\n\n\nThat would make it easier to estimate annotation cost, revise the schema, identify where dedicated training data is actually useful, and give future collaborators something concrete to inspect or extend.",
"title": "The Case for an NVC-Annotated Dataset"
}