# How to enter certain UTF-8 characters

**URL:** <https://forum.graphviz.org/t/how-to-enter-certain-utf-8-characters/1352>\
**Category:** Help\
**Created:** [November 6, 2022, 9:18pm UTC](https://forum.graphviz.org/t/how-to-enter-certain-utf-8-characters/1352 "2022-11-06T21:18:42Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![steveroush](https://avatars.discourse-cdn.com/v4/letter/s/a9adbd/32.png) [@steveroush](https://forum.graphviz.org/u/steveroush)\
**Post date:** [November 6, 2022, 9:18pm UTC](https://forum.graphviz.org/t/how-to-enter-certain-utf-8-characters/1352/1 "2022-11-06T21:18:42Z")

</div>

How can I get this character to dispay: U+1D404 (bold Capital E) (see [UTF-8 Character Set 1D400-1D4FF](https://www.obliquity.com/computer/html/unicode1D400.html))?

---

<div class="post-metadata">

**Author:** ![smattr](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/smattr/32/85_2.png) [@smattr](https://forum.graphviz.org/u/smattr)\
**Post date:** [November 7, 2022, 3:21am UTC](https://forum.graphviz.org/t/how-to-enter-certain-utf-8-characters/1352/2 "2022-11-07T03:21:32Z")

</div>

Not sure I understand the question. Strings are UTF-8 internally, so just put it in your input file? 𝐄

Are you getting some error when you try this?

---

<div class="post-metadata">

**Author:** ![erg](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/erg/32/34_2.png) [@erg](https://forum.graphviz.org/u/erg)\
**Post date:** [November 7, 2022, 11:14pm UTC](https://forum.graphviz.org/t/how-to-enter-certain-utf-8-characters/1352/3 "2022-11-07T23:14:23Z")

</div>

You can identify a unicode character numerically in ascii, either using decimal or hexadecimal notation, but it needs to be in an HTML string. For your case, you can do something like

```auto
graph {
  vn [label=<&#x1D400;>]
  un [label=<&#119808;>]
}

```

The documentation notes the decimal case but not the hex case. Also, I just noted there is a problem in the expat library, in that it doesn’t accept a capital X in the hex case, but only allows a lowercase x. The documentation should have a few more examples.

Looks like only ‘x’ is allowed in HTML entities, so the Graphviz code allowing ‘X’ will never get called.

---

<div class="post-metadata">

**Author:** ![smattr](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/smattr/32/85_2.png) [@smattr](https://forum.graphviz.org/u/smattr)\
**Post date:** [November 16, 2022, 4:56am UTC](https://forum.graphviz.org/t/how-to-enter-certain-utf-8-characters/1352/4 "2022-11-16T04:56:27Z")

</div>

> [@erg](#):
>
> Looks like only ‘x’ is allowed in HTML entities, so the Graphviz code allowing ‘X’ will never get called.

Can you elaborate? I assume you’re referring to lib/common/xml.c:29:

```c
 17 /* return true if *s points to &[A-Za-z]+; (e.g. &Ccedil; )
 18 * or &#[0-9]*; (e.g. &#38; )
 19 * or &#x[0-9a-fA-F]*; (e.g. &#x6C34; )
 20 */
 21 static bool xml_isentity(const char *s)
 22 {
 23 s++; /* already known to be '&' */
 24 if (*s == ';') { // '&;' is not a valid entity
 25 return false;
 26 }
 27 if (*s == '#') {
 28 s++;
 29 if (*s == 'x' || *s == 'X') {
…

```

There are code paths into this XML processing from several places, none of which I see restricting the input to containing only lower case “x”.

I’m also a little unclear about this:

> [@erg](#):
>
> Also, I just noted there is a problem in the expat library, in that it doesn’t accept a capital X in the hex case, but only allows a lowercase x.

I’m not too familiar with the expat codebase, but I think you’re referring to expat/lib/xmltok\_impl.c:516 (as of commit 441f98d02deafd9b090aea568282b28f66a50e36):

```c
 512 static int PTRCALL
 513 PREFIX(scanCharRef)(const ENCODING *enc, const char *ptr, const char *end,
 514 const char **nextTokPtr) {
 515 if (HAS_CHAR(enc, ptr, end)) {
 516 if (CHAR_MATCHES(enc, ptr, ASCII_x))
 517 return PREFIX(scanHexCharRef)(enc, ptr + MINBPC(enc), end, nextTokPtr);
…

```

This comparison does indeed appear to bottom out on exclusively lower case “x”, regardless of encoding. But my reading of [the official guidance](https://www.w3.org/International/questions/qa-escapes) is that lower case “x” is the only legal form. I don’t think this is an expat problem. This seems to be Graphviz (and Chrome and Firefox, now that I check) being overly liberal.
