# Converter API > File conversion and document extraction service for DDS applications. > Base URL: https://converter.ddsdashboard.com > Auth: every endpoint except GET / , GET /health and GET /llms.txt requires > `Authorization: Bearer `. The service never holds storage credentials. Callers pass presigned S3 URLs (or a presigned POST policy) and the service reads and writes through them. All processing of user-supplied files happens here, isolated from application data. ## Endpoints ### GET /health Liveness probe. Returns {"ok": true}. ### GET / HTML landing page: a request builder for every route, the searchable format matrix and the extraction contract. ### POST /convert One file in, one file out. The API writes a job record, enqueues the job and a worker runs it; nothing runs inside the API process. Body (JSON): sourceFormat string e.g. "docx", "png", "mp4" targetFormat string e.g. "pdf", "webp", "mp3" sourceUrl string presigned GET URL the service downloads from destinationUrl string presigned PUT URL the service uploads to outputContentType string? Content-Type sent with the PUT width, height int? output pixels; see "Width and height" below clip object? video only: {"startS": 84, "endS": 141} cuts that range (seconds) into an mp4 instead of converting the whole file; stream copy first, re-encode fallback. targetFormat must be "mp4", else engine_failed. width and height do not apply to a clip. trim object? {"fuzzPercent": 8, "marginPx": 32}. Image pairs: reads the first frame, auto-orients, then trims borders that match the corner color within fuzzPercent (0-100); resize, if given, applies after the trim. html->pdf: runs the same trim on a screenshot of the page and crops the page to what it keeps, plus marginPx on each side, by shrinking the PDF's MediaBox and CropBox; the content is not re-rendered. Every other pair ignores it. mode "sync" (default) | "async" Sync response: 200 {"ok": true, "outputKey", "engine", "commands"} 422 {"ok": false, "errorCode", "message", "commands"} 202 {"ok": true, "jobId", "status"} when the job has not finished within the sync wait cap (45 s); poll GET /jobs/:id. Sync callers must handle 202. Async response: 202 {"ok": true, "jobId", "status": "queued"}; poll GET /jobs/:id. The pair must be in the format matrix on GET /. Pairs outside it answer 422 unsupported_pair. Width and height: ImageMagick image outputs and FFmpeg video outputs scale to them; give one of the two to keep the aspect ratio. Video-to-image thumbnails default to 640x360; a missing side keeps that default, so give both. html->pdf/png use them as the Chromium viewport (default 1200x800). Every other engine ignores them. html->pdf/png: the sourceUrl may be any live http(s) page, fetched directly instead of downloaded. Chromium waits for network idle. png is a screenshot of the viewport. pdf is one page of viewport size, rendered with print media; content past the viewport is cut off. md/markdown->pdf/png/svg: Chromium renders the Markdown and turns Mermaid code blocks into diagrams. Markup pairs run through pandoc with the reader and writer named explicitly (md is pandoc markdown, txt is plain text). eml->html/txt/pdf: a header block (From, To, Cc, Date, Subject), the body (HTML when the message has it, with cid: images inlined), and the names of the attachments. Attachment contents are not converted. eml->pdf and epub->pdf: the HTML is printed by Chromium on Letter pages with 0.5 in margins. Remote requests and JavaScript are blocked, so remote images in an email do not load. pdf->txt/md: the PDF text layer from pdftotext, in reading order. A scanned PDF without a text layer yields no text. pdf->docx/odt/pptx: LibreOffice's PDF import. Pages become positioned text boxes and drawings (one slide per page for pptx), not reflowed paragraphs. xlsx/ods->csv: the first sheet only, UTF-8, comma-separated. csv->xlsx/ods: the CSV is read as UTF-8 with commas. ### GET /jobs/:id Status of a /convert, /extract or /extract-media job: {"ok", "jobId", "kind", "status", "attempt", "createdAt", "startedAt"?, "heartbeatAt"?, "finishedAt"?, "result"?}. status is queued | running | succeeded | failed. ok is true only when status is succeeded, so it is false while the job is queued or running. result is the same object a sync response returns. Records are durable, survive deploys, and are kept for 24 hours after the job finishes. Poll every 2 s. ### POST /extract One document in, a directory of readable artifacts out, plus a manifest. Same mode field and responses as /convert: sync waits up to 45 s and then answers 202; async answers 202 at once. Large documents take minutes, so use async and poll GET /jobs/:id. Body (JSON): sourceFormat string one of: pdf, docx, doc, dotx, odt, rtf, pptx, ppt, potx, odp, xlsx, xls, xlsm, xltx, ods, mp4, m4v, mov, webm, mkv, avi, mpg, mpeg, mp3, m4a, wav, aac, ogg, flac sourceUrl string presigned GET URL mode "sync" (default) | "async" destination object prefix string S3 key prefix ending in "/" that every artifact is written under post object presigned POST policy: {"url": "", "fields": {...}} The policy must allow keys with starts-with , any Content-Type (starts-with ""), and a byte ceiling. options object? renderPages "thin-text" (default) | "all" | "none" maxPages int 1..200 (default 8) rendered pages, from page 1 figures boolean (default true) extract embedded images maxFigures int 0..500 (default 40) Pipeline: 1. Office formats -> LibreOffice -> rendition.pdf. PDF sources skip this. 2. docx and odt -> pandoc -> text.md (GitHub Markdown; headings, lists and tables survive). Spreadsheets -> LibreOffice CSV per sheet -> text.md as one Markdown table per sheet ("## ", first row is the header, 2000 rows per sheet max). Other formats get text.md built from page text; slide formats use "## Slide N" headings. 3. The PDF -> pdfinfo (pageCount, title, pageSize), pdftotext -layout (per-page text), pdfimages -all (figures), pdftoppm (page renders). Figures: images >= 100px on both sides, with page anchors, kept as jpg or png in their native encoding (figures/fig-PPP-NNN.jpg). Skipped: near-blank JPEG layers, byte-identical repeats, and JPEG 2000, JBIG2 and CCITT images. When over maxFigures the kept set is spread across pages. A page with 2+ full-page images or 3+ images tiling 90% of it is a layered slide export: its figures are skipped and it is rendered. Page renders: PNG at 72 dpi, pages 1 to maxPages (pages/p-NN.png). renderPages "thin-text" renders when thinText is true or a page is layered; "all" always renders; "none" never renders. thinText is true when there is no text, wordsPerPage < 40, or the text layer is fragmented: mean token length < 3 (tools that place every glyph separately) or 3% or more of tokens are glued word fragments ("kingandLi"). A warning says which. Slide decks (pptx, ppt, potx, odp) also export their speaker notes through LibreOffice: pages[].notes per slide and notes.md ("## Slide N" sections) when any slide has notes. A failed export is a warning, not an error. 4. Every artifact is uploaded under prefix; manifest.json is written last. Video (mp4, m4v, mov, webm, mkv, avi, mpg, mpeg) -> ffprobe duration (durationS), scene-change timestamps (sceneChangesS), and JPEG keyframes: one every 10 s plus one per scene change, deduplicated by perceptual hash, at most 240, as files[] with role "keyframe" and timestampS (keyframes/keyframe-000123.0s.jpg). No text. options are ignored. Audio (mp3, m4a, wav, aac, ogg, flac) -> durationS only. Transcription is the caller's. For video and audio, pageCount, wordCount and wordsPerPage are 0, thinText is false and pages is empty. Response 200: {"ok": true, "manifest": {...}, "commands": [...]} manifest: format, kind ("pdf" | "document" | "slides" | "spreadsheet" | "video" | "audio"), prefix pageCount, wordCount, wordsPerPage, thinText (see step 3) title?, pageSize? {width, height} of the first page in PDF points converterCommit?: git commit of the converter build that ran the job durationS?: video and audio length in seconds sceneChangesS?: video only, seconds at which the picture cuts files[]: {role, path, contentType, bytes, page?, index?, width?, height?, timestampS?} role: "rendition" | "text" | "notes" | "figure" | "page" | "keyframe" | "manifest" path is relative to prefix; the object key is prefix + path pages[]: {index, text, notes?} warnings[]: non-fatal problems (pandoc fallback, skipped figures, figure caps, pages past maxPages left unrendered, missing speaker notes, render failures) Response 422: {"ok": false, "errorCode", "message", "commands"} ### POST /extract-media A page URL in (a video page, a playlist, or a direct video file), one or more MP4 videos out under a caller-owned prefix, plus a manifest. Always async: answers 202 {"ok": true, "jobId"} at once; poll GET /jobs/:id. Downloads and transcodes take minutes to an hour. Body (JSON): sourceUrl string public http(s) URL of the page; not presigned destination object same shape and policy rules as /extract. Mint the POST policy for 3 hours: it is used after the downloads finish. options object? maxVideos int 1..20 (default 5) videos taken from the page, in order maxBytesPerVideo int (default 2 GiB, max 4 GiB) larger videos are skipped maxDurationS int? longer videos are skipped maxVideos x maxBytesPerVideo must not exceed 12 GiB Refused before any download (422): unsupported_source YouTube (youtube.com, youtu.be, youtube-nocookie.com): YouTube blocks cloud addresses. Embed its player. blocked_address the host resolves to a private, loopback, link-local (cloud metadata), CGNAT or reserved address. yt-dlp follows redirects itself, so only the first hop is checked. Pipeline: 1. yt-dlp downloads up to maxVideos videos, preferring H.264/AAC at up to 1080p. One failed entry of several is a warning, not an error. 2. ffprobe each file. H.264 with AAC or no audio is remuxed; anything else is transcoded to H.264/AAC, at most 1920 px wide. Both put the index first (faststart). A file with no video stream is skipped with a warning. 3. Each video uploads as videos/NN.mp4; manifest.json is written last. Response (the job's result): {"ok": true, "manifest": {...}, "commands": [...]} manifest: sourceUrl, prefix, converterCommit? videos[]: {path, contentType "video/mp4", bytes, title?, pageUrl?, durationS?, width?, height?, transcoded} path is relative to prefix; pageUrl is the video's own page warnings[] {"ok": false, "errorCode", "message", "commands"}: unsupported_source, blocked_address, no_videos (the page has no downloadable video; message carries yt-dlp's error), or a code from the table below. ### GET /llms.txt This document. ## Error codes 400 bad_request malformed body; message says which field 401 unauthorized missing or unknown bearer token 404 not_found unknown job id 422 unsupported_pair /convert pair not in the matrix 422 unsupported_format /extract format not supported 422 missing_capability required engine not installed on the host 422 download_failed sourceUrl could not be fetched 422 workspace_failed temp workspace could not be created 422 engine_failed engine exited non-zero or could not start 422 engine_timeout engine exceeded its wall-clock budget and was killed 422 output_missing engine reported success but wrote nothing usable 422 unsupported_source /extract-media URL is YouTube or not http(s) 422 blocked_address /extract-media host resolves to a non-public address 422 no_videos /extract-media page has no downloadable video 422 upload_failed destination write failed 422 cleanup_failed workspace removal failed after the job finished 422 presigned_url_expired a source/destination URL or POST policy expired, either at submit or while the job waited in the queue 422 attempts_exhausted the job was interrupted 3 times (worker replaced mid-run) without finishing; resubmit 503 queue_unavailable the job could not be enqueued; retry with backoff ## Limits Queueing: jobs wait in a queue and workers scale on its depth, so there is no 503 busy; a burst means longer waits, not refusals. Per-step time budgets: LibreOffice and FFmpeg 10 min; Ghostscript, pdftotext, pdfimages and pdftoppm 5 min; ImageMagick, libarchive and the HTML renderer 3 min; pandoc and the Markdown renderer 2 min; pdfinfo, ffprobe and fontTools 1 min. An extraction runs several steps in a row. /extract-media: yt-dlp 45 min for the whole page, FFmpeg 30 min per video. Presigned URLs: mint them with a 60-minute lifetime. A job can wait in the queue and then run for about 30 minutes; a URL that expires before the worker starts fails the job with presigned_url_expired. Log output per step is truncated to 16 KB in "commands". ## Engines LibreOffice (office documents, PDF import, spreadsheets), pandoc (markup and docx/odt), poppler (pdfinfo, pdftotext, pdfimages, pdftoppm), Ghostscript, ImageMagick, FFmpeg, yt-dlp, libarchive/7-Zip, fontTools, postal-mime (email parsing), Chromium (HTML, Mermaid, email and ebook rendering).