Skip to content
aistructured-contentguides

AI Meets Structured Content Publishing: What Changes, What Doesn't

Shakewell ·

Structured content publishing has a reputation for being slow, expensive and staffed by specialists. So when a technology arrives that can read a document and describe what’s in it, the obvious conclusion is that the specialists were the problem.

That’s the wrong read. AI changes a real and specific set of tasks in a publishing pipeline. It changes almost nothing about why the pipeline exists.

What actually changes

Conversion stops being the bottleneck. Migrating unstructured documents into a standard — Word and FrameMaker into DITA, legacy SGML into something current — used to be the largest line item in any structured-content programme, because a human had to decide what every heading and paragraph was. A model that proposes the markup and a human who reviews it is a genuinely different economics. Not unsupervised, but the ratio changes by an order of magnitude.

Tagging and classification get cheap. Metadata, taxonomy terms, cross-references, index entries: the work nobody enjoys and everybody skips under deadline. This is the strongest fit, because the cost of being wrong is low and the error is visible.

Retrieval finally works. This is the one that matters to readers rather than authors. Structured content has always been theoretically findable — the structure is right there. In practice, search over technical documentation has been poor for decades. Retrieval over properly structured, entitlement-aware source is the first genuine improvement in a long time, because the model can be pointed at exactly the subset a given reader is allowed to see.

Format transformation gets less bespoke. Publishing one source to a manual, a portal, a device and a partner feed is what structured content is for, and the transforms have always been custom engineering. Some of that is now generatable.

What doesn’t change

The standards still matter, and arguably more. S1000D, DITA, MIL-STD-40051, JATS and the ATA specs exist because aviation, defence and regulatory publishing need a document to mean the same thing to everyone who touches it, permanently. A model can help produce conforming output. It cannot decide that conformance is optional. If anything, generated content raises the value of a schema — it’s the thing that catches what the model got wrong.

Review workflows are the control, not the friction. Every AI-assisted step in a regulated pipeline needs a human accountable for the output. That is not a temporary state pending better models. In an industry where the wrong revision reaching an engineer has physical consequences, the review step is the product.

Provenance and versioning become more important, not less. A reader has always needed to know which revision applied on a given date. Add generated or assisted content and you also need to know what produced it, from what source, and who approved it. The systems that handle this well are the ones that treated versioning as a first-class concern before AI was in the conversation.

Entitlement doesn’t get to be approximate. Retrieval systems are excellent at surfacing things. In technical publishing, surfacing a document to an operator not entitled to see it is a serious failure, not a relevance bug. Entitlement has to be enforced at the data layer, so a query can only ever return what that customer is allowed to see — for SES we enforce it there rather than in the interface, precisely because that is the layer furthest from anything a model influences.

The uncomfortable prerequisite

Most organisations wanting AI over their documentation don’t have a model problem. They have a content problem.

If the source is a decade of PDFs with inconsistent structure, no revision metadata and no reliable notion of which version is current, then a retrieval system built over it will be confidently wrong — and confidently wrong is worse than a bad search box, because people believe it.

Structured content is the prerequisite. It always was; AI just made the consequence of skipping it more visible. A system grounded in versioned, entitlement-aware, approved source can cite where an answer came from and be constrained to what has been signed off. One pointed at a folder cannot show its working.

That is also the honest reason a readiness assessment comes before an AI project rather than after. For the AEMC, the value was giving legal and policy teams a browser-based structured workflow that produces conforming output without XML specialists — the structure is what makes anything downstream possible.

What we’d tell you to do first

  1. Audit what you actually have. Not how many documents — how consistently they were produced. That single answer determines whether conversion is weeks or quarters.
  2. Fix versioning and entitlement before retrieval. They are harder to add later and they are what makes the output trustworthy.
  3. Pick the narrow win. Search across a corpus nobody can navigate, or classification of an unclassified backlog. Both are measurable and neither requires betting the pipeline.
  4. Keep the human in the loop where being wrong is expensive. In regulated publishing that is most places, and that is fine.

The organisations getting value here are not the ones that adopted AI fastest. They are the ones whose content was already structured enough to point something at.

If you’re weighing this up, we’re happy to talk it through.


Related: AI consulting & implementation · Technical publications & structured content · SES case study · AEMC case study

Start a conversation

Start a conversation

Tell us what you want to build, fix or scale — we’ll come back with a clear way forward.