Skip to content

substitute leaves placeholders unfilled, or garbles the line, in files saved by LibreOffice #187

Description

@Andy9822

We use docx to fill contract templates: a template has placeholders like {{name}}, and we call paragraph.substitute("{{name}}", "Jane") on each paragraph. Thanks for #181, which made this work when Word splits a placeholder internally.

What happened

It worked on the Word files we tried. Then we opened one of them in LibreOffice, saved it (still as .docx), and ran the same code. Some placeholders stayed as {{name}}, and one line came out mangled. Nothing raised an error. Files converted with macOS's textutil behave the same way.

It happens when a line break (Shift+Enter) or a tab sits on the same line as the placeholder, or when the placeholder is partly inside a link:

In the document After substitute Expected
Name: {{name}}, a line break, Next line unchanged Name: Jane, a line break, Next line
The same, with the placeholder split internally Name: JaneNext lineme}}Next line Name: Jane, a line break, Next line
A tab inside the placeholder's braces A Jane Bam A Jane B
A placeholder that starts inside a link See site {{na (the rest of the line is gone) See site Jane end
A space right after the placeholder shows as Janeend Jane end

The tab case is the worst one: the text is wrong and no braces are left, so checking the output for a leftover {{ doesn't catch it.

To reproduce

The script builds the sample file itself, then runs substitute on each paragraph:

require "docx"
require "zip"

W = "http://schemas.openxmlformats.org/wordprocessingml/2006/main"
PARAGRAPHS = [
  # 1. One run with two w:t around a line break (LibreOffice writes Shift+Enter this way).
  '<w:r><w:t>Name: {{name}}</w:t><w:br/><w:t>Next line</w:t></w:r>',
  # 2. Placeholder split across runs; the last run holds a w:t, w:br, w:t.
  '<w:r><w:t xml:space="preserve">Name: {{na</w:t></w:r><w:r><w:t>me}}</w:t><w:br/><w:t>Next line</w:t></w:r>',
  # 3. A run with several w:t in the middle of the match.
  '<w:r><w:t xml:space="preserve">A {{n</w:t></w:r><w:r><w:t>a</w:t><w:tab/><w:t>m</w:t></w:r><w:r><w:t xml:space="preserve">e}} B</w:t></w:r>',
  # 4. A hyperlink with two runs.
  '<w:r><w:t xml:space="preserve">See </w:t></w:r><w:hyperlink w:anchor="x"><w:r><w:t xml:space="preserve">site </w:t></w:r><w:r><w:t>{{na</w:t></w:r></w:hyperlink><w:r><w:t xml:space="preserve">me}} end</w:t></w:r>',
  # 5. A trailing space moved into a w:t that lacks xml:space="preserve".
  '<w:r><w:t>{{na</w:t></w:r><w:r><w:t xml:space="preserve">me}} </w:t></w:r><w:r><w:t>end</w:t></w:r>',
]
Zip::File.open("repro.docx", create: true) do |zip|
  zip.get_output_stream("[Content_Types].xml") { it.write %(<?xml version="1.0"?><Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types"><Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/><Default Extension="xml" ContentType="application/xml"/><Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/></Types>) }
  zip.get_output_stream("_rels/.rels") { it.write %(<?xml version="1.0"?><Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships"><Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/></Relationships>) }
  zip.get_output_stream("word/_rels/document.xml.rels") { it.write %(<?xml version="1.0"?><Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships"/>) }
  zip.get_output_stream("word/styles.xml") { it.write %(<?xml version="1.0"?><w:styles xmlns:w="#{W}"/>) }
  zip.get_output_stream("word/document.xml") { it.write %(<?xml version="1.0"?><w:document xmlns:w="#{W}"><w:body>#{PARAGRAPHS.map { "<w:p>#{it}</w:p>" }.join}</w:body></w:document>) }
end

doc = Docx::Document.open("repro.docx")
doc.paragraphs.each do |paragraph|
  before = paragraph.text
  paragraph.substitute("{{name}}", "Jane")
  puts "#{before.inspect} => #{paragraph.text.inspect}"
end
doc.save("repro_out.docx")
puts Docx::Document.open("repro_out.docx").doc.xpath("//w:t", "w" => W).to_a.last(3).map(&:to_xml)

Output:

"Name: {{name}}Next line" => "Name: {{name}}Next line"
"Name: {{name}}Next line" => "Name: JaneNext lineme}}Next line"
"A {{name}} B" => "A Jane Bam"
"See site {{name}} end" => "See site {{na"
"{{name}} end" => "Jane end"
<w:t>Jane </w:t>
<w:t xml:space="preserve"/>
<w:t>end</w:t>

Environment: Ruby 4.0.7, docx 0.13.0 (the same code on master at e083d0b), nokogiri 1.19.4, rubyzip 3.7.0, macOS 26.

Technical details

In the file, a paragraph is made of runs (stretches of text with the same formatting), and a run keeps its characters in one or more w:t nodes. Word starts a new run after a line break. LibreOffice and textutil keep one run with two w:t around the break or tab, and a link made of two runs has two w:t as well.

  • TextRun#text= only writes when the run has exactly one text node (lib/docx/containers/text_run.rb:37-45); with two, it does nothing. A hyperlink counts its runs' text nodes too (text_run.rb:27, paragraph.rb:60).
  • Paragraph#substitute (paragraph.rb:108-113) writes the merged text into the first run the match spans and empties the others, without checking that the first write happened. So text is lost, duplicated or left as it was.
  • Separately, Elements::Text#content= never sets xml:space="preserve". When a node ends up starting or ending with a space and has no such attribute, the space is dropped when the file is opened (the Janeend row).

A possible fix: run the same loop over the paragraph's w:t nodes (.//w:t, in document order) instead of its TextRuns. Put the replacement in the node holding the match's first character, cut the rest of the match from the nodes after it, and leave the breaks and tabs outside the match where they are. Set xml:space="preserve" on every node written. We did this in our own code and it gives the expected result on every row above, and on files saved by LibreOffice and textutil.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions