Rendered at 10:10:49 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
PeterStuer 2 hours ago [-]
"The choice of Haskell has also led to a high quality and low volume of contributors"
I feel this influence of choosing a tech stack and its impact on self selected and auto-reenforced culture is most often underestimated.
From my own experience, at a time I was (involuntarily) working in Java, and when .Net was released, from a pure technical point of view it was like a breath of fresh air. Java was suffering from overengineering, archtecture astronauts galore and no sensible UX framework. .Net, the new kid, came in lean and clean with a UX library that 'just worked'.
Problem later was that for all its flaws and being overly 'academic', in teams (the real thing, not the awfull app), you could have indepth discussions about non trivial aspects of SWE topics in the Java world, whereas for all its technical prowess, in .Net land you were mostly dwelling amongst the 2 week CRUD app bootcamp folks. This ofc is a gross oversimplication.
You had brilliant engineers and challanged codemonkeys on both sides. But the skew was more than a little biased.
adamddev1 12 hours ago [-]
> by writing N parsers (“readers”) and M renderers (“writers”), one could support N × M conversions.
Beautiful writeup for a wonderful project. In an age of vibe-coding hype it's also so nice to see how things can be extended and snowball in usefulness when things are built correctly, by hand, from basic principles.
> Perhaps, then, in the future, people will no longer have a need for tools like pandoc.
I think we will need wonderful things like pandoc more and more. As mentioned there is a huge ecological and practical difference. Even if LLMs could get infintisamally close to deterministic-level reliability, it's still so many more orders of magnitude better in efficiency, especially with big batch jobs etc.
vatsachak 8 hours ago [-]
Well the difficult part is finding a good intermediate representation, which they did
aanet 10 hours ago [-]
That a professor of philosophy [1] made tools [2] that are used by millions around the world... that's just mind boggling. In a good (great!) way.
A few years ago, I gobbled together some bash scripts around pandoc to build a site generator for my personal website. Works great. I use html templates, markdown for the content, etc. Mostly the bash scripts just serve to list files and process them one by one. I actually process them concurrently by forking processes so it's reasonably fast. A bit wonky but it works fine for my use case.
As for Haskell, I guess tree transformations and parsing are the perfect use case for functional programming. I studied in Utrecht in the nineties when Erik Meijer was still teaching there (later went to work at Microsoft Research where he contributed to things like F# and Linq). In short, my compiler course was taught using functional programming. We were toying around with writing our own parser generators to implement a subset of Modula 3 or our own toy languages. Lots of monads and other esoteric abstractions.
I haven't really done much professionally with any of that since except having a really easy time when languages like Kotlin, Javascript, etc. started borrowing liberally from functional programming. These days, if you have a list, calling map or forEach on it with another function is perfectly normal in many languages. Very nice alternative to a for or while loop.
huijzer 2 hours ago [-]
I can imagine that it works great, but at the same time I have also used plain Rust to generate HTML. It’s not that hard to generate valid HTML. Parsing however is a different thing. That’s hard. For Markdown source, there are libraries to parse the Markdown. That’s a bit more work to get right, but at the same time Pandoc also requires some tweaks and configuration to get right, plus getting the dependency installed in your environment(s). I personally would go for the Rust libraries way for a new project for flexibility sake but wouldn’t say the other way is wrong per se.
malkosta 10 hours ago [-]
Love pandoc, I use it for all sorts of stuff…here is my minimal static site generator:
find . -name '*.md' -type f -exec sh -c '
for file do
out="docs/${file#./}"
out="${out%.md}.html"
mkdir -p "$(dirname "$out")"
pandoc --quiet --template template.html "$file" -o "$out"
done
' sh {} +
raybb 8 hours ago [-]
To top it all off Pandoc has a great experience for contributors. Over the past few years I've opened several bug reports related to Typst and docx and all of them got responses that were kind and helpful. I even had a few PRs merged in despite knowing almost nothing of Haskell.
jwr 3 hours ago [-]
Thank you for Pandoc! I used it for many things over the years, and it was always there as a good option. I mostly use it for Markdown->Typst these days and it does this job very well.
which strips out styling, wrapper divs, spans, inline attributes, etc from (for instance) HTML copied from a google or word doc. Just the semantic goodness!
koolba 6 hours ago [-]
Pandoc is awesome. My fav usage is configuring git to use it to normalize binary docs to markdown (like a .docx) so they can be diffed. Works amazing for redlining contacts.
sdoering 56 minutes ago [-]
Wow. Thanks. This is an awesome idea. Will blatantly steal. Thanks for sharing.
aleks_me2 2 hours ago [-]
Thanks also that new reader/writer are also integrated like typst.
I use it to create pdfs for my blog posts with that code :-)
Thank you for pandoc. I made the (at the time perhaps not transparently wise) choice to go all in on it when I started my PhD, and that decision has paid nothing but dividends since. I owe my career as a scientist to it.
maleldil 3 minutes ago [-]
How do you usually collaborate with other authors? I don't know your area, but it's common to use Overleaf to collaborate on papers, and you're stuck with LaTeX there. I wrote the first draft of my thesis in Typst, but then converted to LaTeX to get feedback on Overleaf from other people.
rao-v 12 hours ago [-]
I have at times mused that writing the internal state of pandoc to disk (yes 13th standard etc) would be the most interoperable file format
boramalper 4 hours ago [-]
I don’t think it has backward compatibility / stability guarantees though, does it?
mjmas 9 hours ago [-]
pandoc document.md --to native
applicative 4 hours ago [-]
Since your document is a haskell value, syntax highlighting gives added value to the terminal drug trip
pandoc -f html -t native "https://news.ycombinator.com/item?id=49156750" | bat -l hs --style=plain --paging=never
Curiositry 6 hours ago [-]
Pandoc is a fantastic piece of software. I have never had any issues with it, and it's the tool I reach for anytime I need to covert documents. I'm super grateful to John MacFarlane for creating it and maintaining it for all these years!
zaqr 3 hours ago [-]
TIL of this wonderful tool, how did I never used this before, boggles the mind
graemep 12 hours ago [-]
Pandoc is amazing. I have HTML and want a PDF? One command and it just works painlessly. That was just my last use of pandoc this weekend. I do not use it that often, but when I do its perfect.
jack0111 5 hours ago [-]
It's not easy to maintain so many parsers and renderers, and at the same time to keep the qualities of them.
w10-1 6 hours ago [-]
unending thanks, as much for pandoc as for the example of a clear design+implementation that lasts.
zombot 2 hours ago [-]
Happy birthday and thank you so much!
BeetleB 11 hours ago [-]
pandoc was (and still is) critical to some of my flows. I hope it never dies!
oytech 11 hours ago [-]
Effort worth admiration. Thank you for creating pandoc! It allowed me to write my CV in more readable markdown, but also get nice pdf with latex.
I feel this influence of choosing a tech stack and its impact on self selected and auto-reenforced culture is most often underestimated.
From my own experience, at a time I was (involuntarily) working in Java, and when .Net was released, from a pure technical point of view it was like a breath of fresh air. Java was suffering from overengineering, archtecture astronauts galore and no sensible UX framework. .Net, the new kid, came in lean and clean with a UX library that 'just worked'.
Problem later was that for all its flaws and being overly 'academic', in teams (the real thing, not the awfull app), you could have indepth discussions about non trivial aspects of SWE topics in the Java world, whereas for all its technical prowess, in .Net land you were mostly dwelling amongst the 2 week CRUD app bootcamp folks. This ofc is a gross oversimplication. You had brilliant engineers and challanged codemonkeys on both sides. But the skew was more than a little biased.
Beautiful writeup for a wonderful project. In an age of vibe-coding hype it's also so nice to see how things can be extended and snowball in usefulness when things are built correctly, by hand, from basic principles.
> Perhaps, then, in the future, people will no longer have a need for tools like pandoc.
I think we will need wonderful things like pandoc more and more. As mentioned there is a huge ecological and practical difference. Even if LLMs could get infintisamally close to deterministic-level reliability, it's still so many more orders of magnitude better in efficiency, especially with big batch jobs etc.
Pandoc is my go-to tool. Thank you, Sir!
[1] https://johnmacfarlane.net/index.html
[2] https://johnmacfarlane.net/tools.html
https://gist.github.com/rahimnathwani/210b1f9cb6ce731a304322...
As for Haskell, I guess tree transformations and parsing are the perfect use case for functional programming. I studied in Utrecht in the nineties when Erik Meijer was still teaching there (later went to work at Microsoft Research where he contributed to things like F# and Linq). In short, my compiler course was taught using functional programming. We were toying around with writing our own parser generators to implement a subset of Modula 3 or our own toy languages. Lots of monads and other esoteric abstractions.
I haven't really done much professionally with any of that since except having a really easy time when languages like Kotlin, Javascript, etc. started borrowing liberally from functional programming. These days, if you have a list, calling map or forEach on it with another function is perfectly normal in many languages. Very nice alternative to a for or while loop.
find . -name '*.md' -type f -exec sh -c '
' sh {} +tidyhtml () { pandoc -f html-native_divs-native_spans -t markdown-raw_html-raw_attribute | pandoc -f markdown -t html }
which strips out styling, wrapper divs, spans, inline attributes, etc from (for instance) HTML copied from a google or word doc. Just the semantic goodness!
I use it to create pdfs for my blog posts with that code :-)
If you compare this with the HTML produced by typst or hevea, this is super useful, as we can then roll out our own styles.
I never heard about djot [0], is anyone using it?
[0]: https://djot.net/
[0] https://github.com/gn0/nvim-web-server