{
  "$type": "site.standard.document",
  "bskyPostRef": {
    "cid": "bafyreigqauhyyqexvfnok62t5ui4uzez4ede3okjg4fo3osut3sgjwe7mq",
    "uri": "at://did:plc:lk3jfj3zq4k4wxnk474axylu/app.bsky.feed.post/3mp4liptdn3f2"
  },
  "path": "/t/gpt-4o-returning-malformed-unicode-like-u0000e6-instead-of-ae-encoding-bug/1323897#post_4",
  "publishedAt": "2026-06-25T13:54:33.000Z",
  "site": "https://community.openai.com",
  "textContent": "Welcome new user with a web site to promote, multiple old thread bumper.\n\nThe AI can certainly write \\u000000000000999 if it wants. Just as it could write the correct multi-byte Unicode code point if it wanted to and does in the majority of the cases for a topic creator that says “occasionally”.\n\nThe API cannot “ban escaped unicode sequences”.\n\nThe answer you seek: logprobs.\n\nGet the actual predictions of tokens, and alternates that the AI is making.\n\nYou’ll see this is a training issue, likely large amounts of 8-bit code pages were ingested in corpus with flaky conversion on these models, two years old, and even more major fault with this symptom was seen in the initial release of gpt-4-turbo.",
  "title": "GPT-4o returning malformed Unicode like \\u0000e6 instead of æ — encoding bug?"
}