API

Core (Node & framework-free)

Find text and its position in a PDF, on the server or in any framework, with kviewer/core

kviewer/core is a framework-free entry point. It imports no Vue, Nuxt or DOM APIs, so it runs in plain Node (through the pdfjs-dist legacy build) as well as in the browser, React apps included.

The viewer's own search index is built with the same code. Coordinates your server computes match what the editor computes.

import { findTextRects } from 'kviewer/core'

const bytes = new Uint8Array(await readFile('contract.pdf'))
const [anchor] = await findTextRects(bytes, 'Unterschrift Kunde')

if (anchor) {
  const [x1, y1, x2] = anchor.rect
  // Place a 180×40 signature field just below the anchor text.
  viewer.addFormField({
    fieldType: 'signature',
    pageNumber: anchor.page,
    rect: [x1, y1 - 45, x1 + 180, y1 - 5],
  })
}

Coordinates

Every rect is [x1, y1, x2, y2] in unrotated PDF user space: PDF points with the origin at the bottom left. Form-field rects use the same space, so you can pass the values straight through.

  • A page's /Rotate has no effect on the numbers.
  • MediaBox and CropBox origins other than 0, 0 are already included. Don't subtract them.
  • Rotated or skewed text yields the axis-aligned box that encloses it.
  • A box goes from the baseline up to the font size. Descenders stick out below it.

findTextRects(pdf, query, options?)

Returns a Promise<TextMatch[]>, ordered by page and then by reading order.

A match can cover several pdf.js text items. pdf.js often splits a phrase such as "Unterschrift Kunde" into several items, or turns the space into an item of its own. kviewer matches against a normalised string for each page and then unions the boxes of the covered items. If a match covers only part of an item, that item's box is cut in proportion to the number of characters covered. This is an approximation because pdf.js doesn't expose the width of each glyph, but the browser and the server compute it the same way.

interface TextMatch {
  page: number                 // 1-based
  text: string                 // matched text in the normalised page string
  rect: [number, number, number, number]
  runs: TextRun[]              // every run the match touches
}
OptionTypeDefaultDescription
pagesnumber[]all1-based pages to search
caseSensitivebooleanfalseApplies to string queries only. A RegExp keeps its own flags
normalizeWhitespacebooleantrueCollapse runs of whitespace and line breaks into a single space, in both the page text and the query
batchSizenumber5Pages extracted concurrently before yielding to the event loop
signalAbortSignal—Cancel a long extraction
documentParamsDocumentInitParameters—Extra pdf.js getDocument options, used only when you pass bytes

query can be a string or a RegExp, for example /Unterschrift (Kunde|Service)/.

extractTextRuns(pdf, options?)

Returns a Promise<TextRun[][]> with one array per requested page, in the order of options.pages (by default page 1 to n). It takes the same pages, batchSize, signal and documentParams options as findTextRects. A page with no text layer, such as a scanned page, gives an empty array.

interface TextRun {
  page: number
  str: string
  rect: [number, number, number, number]
  transform: [number, number, number, number, number, number] // pdf.js text matrix
  width: number    // advance along the baseline
  height: number   // font size
  hasEOL: boolean  // pdf.js saw a line break after this run
}

findMatchesInRuns(runs, query, options?) runs the matching step on runs you have already extracted. Use it to search one document for several anchors without parsing it again.

Input

pdf can be a Uint8Array, an ArrayBuffer or a pdf.js PDFDocumentProxy you have already loaded. If you pass bytes, kviewer opens the document, reads it and then closes it. Your buffer is copied first, so pdf.js does not detach it.

Text in CJK fonts that rely on CMaps needs documentParams: { cMapUrl, cMapPacked: true }, pointing at pdfjs-dist/cmaps/. Latin text works without any extra setup.

Anchored fields

Store a field relative to an anchor text once per document type, then place it on every new PDF of that type. The viewer builds these definitions (see Anchored fields) and resolves them with the same code.

import {
  resolveAnchoredFields,
  isResolvedAnchoredField,
  toFormFieldDefinition,
  writeFormFieldsToPdf,
} from 'kviewer/core'
import { PDFDocument } from 'pdf-lib'

const results = await resolveAnchoredFields(bytes, defs)

for (const r of results) {
  if (!isResolvedAnchoredField(r)) console.warn(r.def.id, r.error)   // 'anchor-not-found' | 'ambiguous'
  else if (r.outsidePage) console.warn(r.def.id, 'reaches past the page')
}

// Optional: write them into the PDF as real AcroForm fields.
const pdf = await PDFDocument.load(bytes)
await writeFormFieldsToPdf(pdf, [], results.filter(isResolvedAnchoredField).map(toFormFieldDefinition))
interface AnchoredFieldDef {
  id: string
  fieldType: FormFieldType            // 'signature' | 'text' | 'checkbox' | …
  fieldName?: string                  // defaults to id
  role?: string                       // free-form, e.g. 'kunde'
  anchor: {
    text: string
    occurrence?: 'first' | 'last' | number  // number is 1-based
    pages?: number[]
    caseSensitive?: boolean
  }
  offset: {
    dx: number
    dy: number
    width: number
    height: number
    origin?: 'anchor-bottom-left' | 'anchor-top-left'  // default bottom-left
  }
  meta?: Record<string, unknown>
}

The offset moves the chosen corner of the anchor box onto the same corner of the field. It is in PDF points and measured as the page is displayed: dx grows to the right and dy grows upwards. A field 40pt below its anchor stays below it on a page with /Rotate 90.

  • An anchor that occurs more than once resolves to ambiguous, unless occurrence picks one. Occurrences are counted across the searched pages in reading order.
  • An occurrence past the last match resolves to anchor-not-found.
  • Fields are not clamped to the page. outsidePage: true flags the ones that reach past it.

placeFromAnchor(anchorRect, offset, page) and offsetFromAnchor(anchorRect, fieldRect, page, origin?) are the two halves of the offset math. page is a PageTransform (width, height, rotation, viewOffsetX, viewOffsetY).

writeFormFieldsToPdf embeds a Unicode fallback font only when a text field value leaves the WinAnsi range. In Node, provide the font bytes with setUnicodeFontLoader(() => readFile(…)) if you need that.
Copyright © 2026