Skip to content
SE
menu

$ proof

Proof, measured — not promised.

Synthetic Engineering is built by the same senior team as Synthetic Users — people who have delivered AI transformation as a preferred partner to demanding organisations. Every workflow carries a baseline and a delta.

Delivered for organisations including

  • JPMorgan
  • Samsung
  • Comcast
  • Capgemini
  • AB InBev
  • TikTok
  • Vitra
  • Square
01 /  Case points

Described by mechanism. We do not publish invented numbers — specific, cleared figures and named references are shared under NDA at the intro call.

Technical documentation /0.1

Authoring that keeps pace with engineering

Tier-1 automotive supplier · service & repair content

Problem
A DITA-based authoring team maintained six document families across model variants and markets. Time-tracking in the diagnostic showed the largest single block of authoring hours went to cross-checking released engineering data — torque specs, part references, variant applicability — not to writing.
What we did
Six-week baseline of hours per document type across the delivery line. Then a two-week throwaway prototype: a drafting agent over an LLM API with retrieval on released PLM extracts, proposing updates with source citations. Two-week limited deployment with five senior authors, every output behind a validation gate owned by the lead author.
How it's measured
Instrumented as hours per document type vs. the pre-agent baseline, plus a citation-accuracy audit on every gated output. Drafting cleared the gate and scaled to the full team; a second candidate — auto-translation of safety warnings — was killed at the two-week mark for unacceptable review burden.
Localization /0.2

Localization past release cadence

Consumer electronics · 100+ locales

Problem
Release windows compressed from weeks to days while human review throughput stayed flat. The bottleneck wasn't machine translation quality — it was reviewers re-reading whole files because they couldn't trust which segments needed attention.
What we did
Mapped the data landscape first: translation memories, termbases and style guides per locale tier. Staged an LLM post-edit pipeline inside the existing TMS — no platform swap — starting with the four highest-volume European pairs plus Japanese, with reviewer-in-the-loop sampling QA and per-segment confidence routing deciding what humans actually read.
How it's measured
Measured as reviewed words per reviewer-hour by language tier, with LQA error rates held against the pre-pipeline baseline. Tier-1 pairs cleared the gate and expanded tier by tier; two low-volume tiers stayed fully human — the delta never justified the review overhead.
Engineering change /0.3

Change impact, traced in minutes

Industrial group · SAP change records

Problem
Every engineering change order triggered a manual hunt: analysts traced affected manuals, parts catalogues and training content by memory and spreadsheet, typically days per change, with misses surfacing months later as field-documentation defects.
What we did
Two-week prototype on off-the-shelf orchestration: an agent walking change records against BOM structures and content references, producing a ranked impact list with confidence per item. Kept deliberately human-gated — the analyst signs off the list; the agent never edits content.
How it's measured
Measured as analyst time-to-impact-list vs. baseline, and missed-impact rate audited against a set of historical changes with known outcomes. Cleared the gate on time; the miss-rate audit is what earned analyst trust and adoption.
Parts data /0.4

One defensible dataset, many workflows

Off-highway equipment · parts & catalogue data

Problem
The same part carried conflicting attributes across the dealer catalogue, ERP and a legacy PIM. Every automation candidate downstream — search, cross-sell, documentation reuse — stalled on the same root cause.
What we did
Data-quality diagnostic first: duplicate and conflict rates quantified by attribute class across the three systems. Then the one custom build the gates justified — a curated golden-record dataset with a matching pipeline and a human adjudication queue for the ambiguous tail, instead of another agent on top of bad data.
How it's measured
Measured as records reconciled per adjudicator-hour and downstream error rate in the workflows the dataset feeds. The dataset is the moat: three previously-killed automation candidates were re-scored against it and two re-entered the build queue.
02 /  Measurement philosophy

The analytics layer decides what productizes and what dies.

Baseline and delta on every workflow. Internal wins are proven first; only what clears the bar is scaled, and only what scales is ever considered for productization. Nothing ships on a hunch.

03 /  Sectors & work types
  • Technical authoring
  • Engineering-change / BOM impact
  • Localization at scale (100+ languages)
  • Parts / catalogue data
  • Training content
  • 3D / visual assets

note: specific engagement detail and references available under NDA at the intro call.

Want the specifics?

We share cleared metrics and named references under NDA on the intro call.

Book an intro call