Core (Node & framework-free)
kviewer/core is a framework-free entry point. It imports no Vue, Nuxt or DOM APIs, so it runs in plain Node (through the pdfjs-dist legacy build) as well as in the browser, React apps included.
The viewer's own search index is built with the same code. Coordinates your server computes match what the editor computes.
import { findTextRects } from 'kviewer/core'
const bytes = new Uint8Array(await readFile('contract.pdf'))
const [anchor] = await findTextRects(bytes, 'Unterschrift Kunde')
if (anchor) {
const [x1, y1, x2] = anchor.rect
// Place a 180×40 signature field just below the anchor text.
viewer.addFormField({
fieldType: 'signature',
pageNumber: anchor.page,
rect: [x1, y1 - 45, x1 + 180, y1 - 5],
})
}
Coordinates
Every rect is [x1, y1, x2, y2] in unrotated PDF user space: PDF points with the origin at the bottom left. Form-field rects use the same space, so you can pass the values straight through.
- A page's
/Rotatehas no effect on the numbers. - MediaBox and CropBox origins other than
0, 0are already included. Don't subtract them. - Rotated or skewed text yields the axis-aligned box that encloses it.
- A box goes from the baseline up to the font size. Descenders stick out below it.
findTextRects(pdf, query, options?)
Returns a Promise<TextMatch[]>, ordered by page and then by reading order.
A match can cover several pdf.js text items. pdf.js often splits a phrase such as "Unterschrift Kunde" into several items, or turns the space into an item of its own. kviewer matches against a normalised string for each page and then unions the boxes of the covered items. If a match covers only part of an item, that item's box is cut in proportion to the number of characters covered. This is an approximation because pdf.js doesn't expose the width of each glyph, but the browser and the server compute it the same way.
interface TextMatch {
page: number // 1-based
text: string // matched text in the normalised page string
rect: [number, number, number, number]
runs: TextRun[] // every run the match touches
}
| Option | Type | Default | Description |
|---|---|---|---|
pages | number[] | all | 1-based pages to search |
caseSensitive | boolean | false | Applies to string queries only. A RegExp keeps its own flags |
normalizeWhitespace | boolean | true | Collapse runs of whitespace and line breaks into a single space, in both the page text and the query |
batchSize | number | 5 | Pages extracted concurrently before yielding to the event loop |
signal | AbortSignal | — | Cancel a long extraction |
documentParams | DocumentInitParameters | — | Extra pdf.js getDocument options, used only when you pass bytes |
query can be a string or a RegExp, for example /Unterschrift (Kunde|Service)/.
extractTextRuns(pdf, options?)
Returns a Promise<TextRun[][]> with one array per requested page, in the order of options.pages (by default page 1 to n). It takes the same pages, batchSize, signal and documentParams options as findTextRects. A page with no text layer, such as a scanned page, gives an empty array.
interface TextRun {
page: number
str: string
rect: [number, number, number, number]
transform: [number, number, number, number, number, number] // pdf.js text matrix
width: number // advance along the baseline
height: number // font size
hasEOL: boolean // pdf.js saw a line break after this run
}
findMatchesInRuns(runs, query, options?) runs the matching step on runs you have already extracted. Use it to search one document for several anchors without parsing it again.
Input
pdf can be a Uint8Array, an ArrayBuffer or a pdf.js PDFDocumentProxy you have already loaded. If you pass bytes, kviewer opens the document, reads it and then closes it. Your buffer is copied first, so pdf.js does not detach it.
documentParams: { cMapUrl, cMapPacked: true }, pointing at pdfjs-dist/cmaps/. Latin text works without any extra setup.Anchored fields
Store a field relative to an anchor text once per document type, then place it on every new PDF of that type. The viewer builds these definitions (see Anchored fields) and resolves them with the same code.
import {
resolveAnchoredFields,
isResolvedAnchoredField,
toFormFieldDefinition,
writeFormFieldsToPdf,
} from 'kviewer/core'
import { PDFDocument } from 'pdf-lib'
const results = await resolveAnchoredFields(bytes, defs)
for (const r of results) {
if (!isResolvedAnchoredField(r)) console.warn(r.def.id, r.error) // 'anchor-not-found' | 'ambiguous'
else if (r.outsidePage) console.warn(r.def.id, 'reaches past the page')
}
// Optional: write them into the PDF as real AcroForm fields.
const pdf = await PDFDocument.load(bytes)
await writeFormFieldsToPdf(pdf, [], results.filter(isResolvedAnchoredField).map(toFormFieldDefinition))
interface AnchoredFieldDef {
id: string
fieldType: FormFieldType // 'signature' | 'text' | 'checkbox' | …
fieldName?: string // defaults to id
role?: string // free-form, e.g. 'kunde'
anchor: {
text: string
occurrence?: 'first' | 'last' | number // number is 1-based
pages?: number[]
caseSensitive?: boolean
}
offset: {
dx: number
dy: number
width: number
height: number
origin?: 'anchor-bottom-left' | 'anchor-top-left' // default bottom-left
}
meta?: Record<string, unknown>
}
The offset moves the chosen corner of the anchor box onto the same corner of the field. It is in PDF points and measured as the page is displayed: dx grows to the right and dy grows upwards. A field 40pt below its anchor stays below it on a page with /Rotate 90.
- An anchor that occurs more than once resolves to
ambiguous, unlessoccurrencepicks one. Occurrences are counted across the searched pages in reading order. - An
occurrencepast the last match resolves toanchor-not-found. - Fields are not clamped to the page.
outsidePage: trueflags the ones that reach past it.
placeFromAnchor(anchorRect, offset, page) and offsetFromAnchor(anchorRect, fieldRect, page, origin?) are the two halves of the offset math. page is a PageTransform (width, height, rotation, viewOffsetX, viewOffsetY).
writeFormFieldsToPdf embeds a Unicode fallback font only when a text field value leaves the WinAnsi range. In Node, provide the font bytes with setUnicodeFontLoader(() => readFile(…)) if you need that.