# Large graph builders - should Graphviz provide more help?

**URL:** <https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097>\
**Category:** Dev\
**Created:** [March 27, 2024, 1:58am UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097 "2024-03-27T01:58:53Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![steveroush](https://avatars.discourse-cdn.com/v4/letter/s/a9adbd/32.png) [@steveroush](https://forum.graphviz.org/u/steveroush)\
**Post date:** [March 27, 2024, 1:58am UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/1 "2024-03-27T01:58:53Z")

</div>

And if so - what?

Most of the questions to the Forum, to stackoverflow, and to the Issues system are about small-ish graphs. Does that mean that large graph builders have everything under control?  
My intuition says “no”, but my intuition is often confused.

So, if there are any users who build large graphs (say, 1500 edges or more +/-), please share your opinions - how can Graphviz be more supportive?

p.s. I know that faster is always better. What else? Especially in the area of documentation, tips, tools.

---

<div class="post-metadata">

**Author:** ![schmoo2k](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/schmoo2k/32/29_2.png) [@schmoo2k](https://forum.graphviz.org/u/schmoo2k)\
**Post date:** [March 27, 2024, 6:46pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/2 "2024-03-27T18:46:10Z")

</div>

One quick win would be the ability to have some sort of “progress” callback (at the API level)

FWIW My main use case involves graphs that are too big and I present a tree view + breadcrumbs view of the data so they can pick the “top” cluster to render from…

 ![image](https://global.discourse-cdn.com/graphviz/original/2X/f/fa69becbb62a5b8b278776c61d55a042afc0b982.png)

---

<div class="post-metadata">

**Author:** ![scnorth](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/scnorth/32/89_2.png) [@scnorth](https://forum.graphviz.org/u/scnorth)\
**Post date:** [March 28, 2024, 3:38pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/3 "2024-03-28T15:38:08Z")

</div>

We might provide more tutorial guidance about what to expect. In terms of layout, if a graph is directed but is not a tree, layered graph layout (dot) isn’t usually practical for more than a few dozen nodes. You can achieve somewhat of the same effect using neato -Gmode=hier (as explained in this report) up to maybe a few hundred nodes. sfdp for undirected graphs was engineered for thousands of nodes.

For large graph viewing, I’m not sure of the state of the art today, but it used to be the case that web clients run out of gas around 100K DOM objects. A node or edge typically has a couple of DOM objects. (I forget whether piecewise cubic Bezier splines are represented by one DOM object or possibly one per segment.) We like d3-graphviz (which is presently pinned to the top of this forum) it does rely on DOM rendering, and it’s probably good for several tens of thousands of nodes and edges. (Someone should let us know.)

cytoscape may be more scalable.

When graphs get really large, for undirected graphs (spring models) there is technology like tSNE and uMAP implementations in python that may be more satisfactory if you don’t need to see the edges. Also the “graphistry” commercial software uses GPUs for layout and rendering and is more scalable up to the memory limit of the GPU. I don’t think it has the concrete rendering features of graphviz (like all the node shapes and text layout options) but when I looked at it a few years ago it seemed good.

---

<div class="post-metadata">

**Author:** ![smattr](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/smattr/32/85_2.png) [@smattr](https://forum.graphviz.org/u/smattr)\
**Post date:** [March 30, 2024, 1:46am UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/4 "2024-03-30T01:46:33Z")

</div>

> [@schmoo2k](#):
>
> have some sort of “progress” callback

The vast majority of expensive graphs I’ve profiled are bottlenecked in `dfs_range`. I wonder if simply plumbing through a progress callback for that single function would be enough.

---

<div class="post-metadata">

**Author:** ![scnorth](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/scnorth/32/89_2.png) [@scnorth](https://forum.graphviz.org/u/scnorth)\
**Post date:** [March 30, 2024, 6:01pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/5 "2024-03-30T18:01:08Z")

</div>

When I profiled, long ago, I thought the cost of a large layout was fairly well balanced between phases 2-4 (mincross, X coord solving, spline routing); cost of phase 1 (ranking = Y coord solving) i negligible. dfs\_range is only involved in phase 3. A progress bar would probably need to account for this.

---

<div class="post-metadata">

**Author:** ![FeRDNYC](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/ferdnyc/32/1428_2.png) [@FeRDNYC](https://forum.graphviz.org/u/FeRDNYC)\
**Post date:** [May 3, 2024, 10:19pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/6 "2024-05-03T22:19:51Z")

</div>

> [@steveroush](#):
>
> Most of the questions to the Forum, to stackoverflow, and to the Issues system are about small-ish graphs.

Beware inferences like that, though — there’s a strong, _strong_ bias on Stack Exchange sites for questions to be accompanied by a [Minimal, Reproducible Example](https://stackoverflow.com/help/minimal-reproducible-example). It’s kind of learned behavior for a _lot_ of support/assistance contexts, really.

So, people dealing with large-graph issues may simply be _reducing them down_ to small-graph examples, for the purpose of discussing their issues with others.

---

<div class="post-metadata">

**Author:** ![steveroush](https://avatars.discourse-cdn.com/v4/letter/s/a9adbd/32.png) [@steveroush](https://forum.graphviz.org/u/steveroush)\
**Post date:** [May 3, 2024, 10:33pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/7 "2024-05-03T22:33:35Z")

</div>

Fair enough. Any thoughts about how to help those who wrestle with large graphs?

---

<div class="post-metadata">

**Author:** ![FeRDNYC](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/ferdnyc/32/1428_2.png) [@FeRDNYC](https://forum.graphviz.org/u/FeRDNYC)\
**Post date:** [May 4, 2024, 4:11am UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/8 "2024-05-04T04:11:05Z")

</div>

Nothing I think would be particularly helpful.

Though I will remind everyone of [the graphviz repo issue](https://gitlab.com/graphviz/graphviz/-/issues/2486) I filed a while back, wondering what the `world.gv` and `4elt.gv` files in the graphviz distribution were all about. Both of those _definitely_ fall into the “large graph” category, as 4ELT is 15,000+ nodes and over 45,000 edges, while the World map is 148,201 nodes and 1 fewer edges.

(_You_ were the one who eventually figured out that `neato` could render either graph in mere seconds, after others of us had hit program crashes and 5±minute runtimes trying to process them through `dot` or `sfdp`.)

---

<div class="post-metadata">

**Author:** ![FeRDNYC](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/ferdnyc/32/1428_2.png) [@FeRDNYC](https://forum.graphviz.org/u/FeRDNYC)\
**Post date:** [May 4, 2024, 4:31am UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/9 "2024-05-04T04:31:00Z")

</div>

Actually, I guess that’s one area that could use expansion.

There are a _lot_ of tools documented as part of the graphviz swiss army knife. I count 29 tools [in the Command Line section](https://www.graphviz.org/doc/info/command.html), _plus_ another 10 [in the Layout Engines section](https://www.graphviz.org/docs/layouts/). And clearly, they all have different strengths and weaknesses. So, how does a hypothetical large-graph user go about _figuring out_ what criteria they should use when selecting a tool for their needs, never mind actually making that selection?

The documentation doesn’t necessarily offer that much guidance in that regard. In fact, sometimes it’s almost the exact opposite.

The [docs for `neato`](https://www.graphviz.org/docs/layouts/neato/), for example, introduce it with:

> `neato` is a reasonable default tool to use for undirected graphs that aren’t too large (about 100 nodes), when you don’t know anything else about the graph.
> 
> `neato` attempts to minimize a global energy function, which is equivalent to statistical multi-dimensional scaling.

And yet, all evidence in that description to the contrary, it was the only tool capable of rendering those large graphs on less-than-glacial time scales.

So I guess the questions are,

1. Why is `neato` the only capable choice for rendering those example fles?
2. How did you _know_ it was the right choice?
3. Can more of that logic/guidance be encoded into the documentation, somehow?

---

<div class="post-metadata">

**Author:** ![smattr](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/smattr/32/85_2.png) [@smattr](https://forum.graphviz.org/u/smattr)\
**Post date:** [May 4, 2024, 8:46pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/10 "2024-05-04T20:46:25Z")

</div>

> [@FeRDNYC](#):
>
> Both of those _definitely_ fall into the “large graph” category

If you want real world example large graphs, look for Gitlab issues labelled `performance`. These are mostly graphs that are too large (either in runtime or memory) for Graphviz to currently handle.

---

<div class="post-metadata">

**Author:** ![scnorth](https://sea2.discourse-cdn.com/graphviz/user_avatar/forum.graphviz.org/scnorth/32/89_2.png) [@scnorth](https://forum.graphviz.org/u/scnorth)\
**Post date:** [May 5, 2024, 12:03pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/11 "2024-05-05T12:03:58Z")

</div>

The documentation could definitely be improved. The tools in Graphviz accumulated organically, so doesn’t provide a good overview of how everything fits together, or what to use when.

Several times we tried working on a book, but didn’t get that far.

Someone (who did write a leading visualization textbook) suggested to start differently, by working on a better FAQ or Wiki, with plenty of examples, and see how that goes.

I was going to write a brief explanation of the layout tools and viewers here. Even that started to seem kind of complicated and pointlessly anecdotal.

---

<div class="post-metadata">

**Author:** ![steveroush](https://avatars.discourse-cdn.com/v4/letter/s/a9adbd/32.png) [@steveroush](https://forum.graphviz.org/u/steveroush)\
**Post date:** [May 5, 2024, 3:46pm UTC](https://forum.graphviz.org/t/large-graph-builders-should-graphviz-provide-more-help/2097/12 "2024-05-05T15:46:40Z")

</div>

- Why is **neato** the only capable choice for rendering those example files?  
**neato** (and only **neato** ) has _special_ features that allows it to accept pre-defined node layouts ( **pos** attribute) and predefined node + edge layouts (see [FAQ | Graphviz](https://www.graphviz.org/faq/#FaqDotWithNodeCoords) and [FAQ | Graphviz](https://www.graphviz.org/faq/#FaqDotWithCoords)). It then takes these **pos** values as given and creates finished graphs - much faster.
- How did you know it was the right choice?  
Eventually, I really looked at the source. Upon noticing the **pos=** node attributes, **neato -n** was the obvious choice.
- Can more of that logic/guidance be encoded into the documentation, somehow?  
Oddly, this (presence of **pos=** attributes) is a situation that _is_ reasonable documented - but who gazes at a 1Mb input file?  
I like all kinds of documentation styles (see [scnorth](https://forum.graphviz.org/u/scnorth)’s answer below) I also like _active_ documentation (my term, not trademarked) - systems that volunteer answers to questions the user does not know to ask.  
Specifically, one thing we might do is enhance the layout engine front-end to
  - flag language misuse (e.g. **rankdir** applied to a cluster)
  - engine use suggestions (e.g. use **neato -n** if **pos** attributes present)
  - flag implementation weaknesses/bugs (e.g. **ortho** splines fail if ports used)
  - flag invalid or dubious attribute values (e.g. **height=25** )

- See [Linter for the DOT language](https://forum.graphviz.org/t/linter-for-the-dot-language/1594/) & [Gvstats - info about your Graphviz files](https://forum.graphviz.org/t/gvstats-info-about-your-graphviz-files/2029/) for more info.
