Start with the content you already have
A structured-content program does not begin with an empty repository. It begins with a few thousand documents that are already written, already approved and already in use.
Four things have to happen before a document is a component
A document carries its structure in its formatting. A component carries its structure in its markup, its metadata and its links. The distance between those two states is the work, and it is the same four steps every time.
Word, HTML, Markdown, text
Content arrives in the formats it was written in. Files are read where they are, without a hand-built intermediate step for each one.
Metadata and styling
Metadata is captured on the way in rather than reconstructed afterward, and styling is aligned to the target model instead of carried over document by document.
Links and cross-references
References that pointed at a page number, a bookmark or a file path have to point at a component instead. Links are fixed as part of the import, not left for a reviewer to find.
Componentization
The document is split into topics that can be approved, reused and published on their own. This is the step that makes every later reuse possible, and the step a copy of the file in SharePoint never delivers.
An import path is not an authoring surface
The argument this site makes against most of the field is about the daily authoring surface. A platform that treats Microsoft Word as an import format asks the subject matter expert to hand a file to someone who converts it, every time, forever, and that is where the chain of custody breaks. DxAuthor+ is the answer to it: experts write in Word, and the structure is applied at save time.
Onboarding an estate that already exists is a different question. It happens once per body of content, and it ends. It also does not contradict the no-migration claim made everywhere else here, because that claim is about the platform: nothing moves off the customer's SharePoint, and there is no parallel repository, no second identity model and no separate audit log. The documents are a separate matter. They still have to become components, and that one-time conversion is better done by a tool that captures metadata, aligns styling, fixes links and componentizes than by a team pasting sections in by hand.
The estate does not have to move in one pass
Migrations of this shape usually run as a phased project rather than a single event, and the phasing is what keeps the program shippable while it is happening.
- 01
New content goes straight in
From the day the repository exists, anything newly written is authored as components. This stops the estate growing in the old shape while the old shape is being converted.
- 02
One document set goes first
Usually the set that costs the most to translate, or the one reviewed most often, because that is where component reuse pays for the conversion soonest.
- 03
Legacy content follows as it is touched
A document scheduled for revision is a document worth converting. The revision and the conversion are the same piece of work, done once.
- 04
The cold archive waits
Content that is retained but not maintained does not need to be componentized to be compliant. Converting it early spends effort on documents nobody will edit again.
Content comes in, lives in SharePoint, is written in Word, and is checked against your rules.
One place to look for each of the four jobs: bringing existing content in, managing and publishing it, writing it, and checking it against your rules.
DxMigrationTool
The content you already have, brought in as components.
- Imports Word, HTML, Markdown and text files
- Metadata captured, styling aligned
- Links and cross-references fixed
- Documents componentized into topics
Dx5
Content management and publishing inside Microsoft SharePoint.
- Browser-based topic-tree authoring
- Built-in approval workflows and snapshot capability
- One-click publishing: Word, HTML, XML
- Translation packages per language
DxAuthor+
Structured DITA authoring inside Microsoft Word.
- Word add-in, authors stay in the tool they know
- Topic tagging from SharePoint Term Store
- Cross-topic linking with SharePoint search and preview
- Single sign-on through Entra ID
DxChecker
Rule-based content quality across the SharePoint repository.
- Configurable rule sets: your standards as machine-checkable rules
- Runs at authoring time, on schedule, or via workflow
- Controlled-language enforcement (ASD-STE100 and equivalents)
- Roadmap: deeper DxAuthor+ integration; AI-supported content rules
Frequently asked questions
Which formats does DxMigrationTool read?
Microsoft Word files, HTML files, Markdown files and plain-text files. Those are the four inputs. Content already held as DITA XML does not need this tool: a move from one DITA repository into SharePoint is a repository move, and it keeps topic types, references and metadata as they are.
Does an import tool contradict the no-migration argument?
No, because the two words point at different things. No migration means no platform migration: the content stays in the SharePoint tenant the organization already runs, so there is no parallel repository to stand up and no second identity or audit regime. The documents themselves still have to become components, and that conversion is what this tool does.
Does it decide the topic model for us?
No. Which document set goes first, and what the topic model looks like, are decisions a human makes once and then applies. Most programs start with one document set, usually the one that costs the most to translate or is reviewed most often. DxMigrationTool applies the model to the files, captures the metadata, aligns the styling and repairs the links.
How is this different from DxChecker?
DxChecker validates content in place against a ruleset, at authoring time, on a schedule or inside a workflow. It reports and, where a rule allows it, fixes. DxMigrationTool runs once per body of content, on the way in. One is the gate that stays; the other is the door.
What does DxMigrationTool cost?
DitaExchange does not publish pricing. Commercials are scoped per engagement, because deployment, scale and the shape of the content program vary widely across regulated customers. The useful first conversation is about which document set goes first, not about a license count.
Bring one real document set to the first conversation
The useful question is not how large the estate is. It is which document set goes first, what its reuse looks like, and what the links point at today. All three are answerable before anything is imported.