Skip to content

About

Extract text and UI elements from screenshots in Go, via OCR and icon-detection models run through ONNX Runtime

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

go-textshot

Read the text and the clickable elements on a screenshot so a Go program can wait for a known screen, then click or type. It is the screen-reading layer behind GUI-automation tools such as go-mackit, for the cases where there is no accessibility tree or DOM to ask: VM consoles, recovery and installer screens, remote desktops.

Under the hood it runs the PP-OCRv6 detection and recognition models, and on demand the OmniParser icon_detect_v3 element detector, through ONNX Runtime. The runtime library and the models are downloaded on first use, pinned by SHA-256, and cached. No system packages: only a C compiler at build time.

Quick start

go get github.com/devcell-sh/go-textshot
package main

import (
	"context"
	"fmt"
	"image"
	_ "image/png"
	"log"
	"os"

	"github.com/devcell-sh/go-textshot"
)

func main() {
	f, err := os.Open("screenshot.png")
	if err != nil {
		log.Fatal(err)
	}
	defer f.Close()
	img, _, err := image.Decode(f)
	if err != nil {
		log.Fatal(err)
	}

	ctx := context.Background()
	eng, err := textshot.New(ctx, textshot.Options{}) // first call downloads ~40 MB into the cache
	if err != nil {
		log.Fatal(err) // errors.Is(err, textshot.ErrUnavailable): no OCR on this build/host
	}
	defer eng.Close()

	results, err := eng.DetectText(ctx, img) // []TextResult{Box, Text, Score}, reading order
	if err != nil {
		log.Fatal(err)
	}
	for _, line := range textshot.Lines(results) {
		fmt.Println(line)
	}
}

The same program is in examples/detecttext. Elements and pixel state come from one more call:

screen, err := eng.DetectScreen(ctx, img)       // Screen{Size, Theme, Backdrop, Elements}
for _, e := range screen.Elements {             // Element{ID, Key, Kind, Box, Score, Text, Style, Value, UOM}
    fmt.Println(e.ID, e.Key, e.Box, e.Style.Disabled)   // 25 button:Continue (1160,554)-(1243,575) true
}

changed, fraction := textshot.Diff(prev, img)   // what moved since the last frame

The CLI does the same from a shell:

go install github.com/devcell-sh/go-textshot/cmd/textshot@latest
textshot screenshot.png            # one text line per output line
textshot -json screenshot.png      # JSON array of elements: file, id, key, kind, box, score, text, value, uom, style
textshot -prefetch                 # warm the cache (runtime + all models) without running

What it extracts

Every value below is measured from the pixels; nothing comes from the OS or the application.

Call Returns Fields
DetectText one TextResult per text line, in reading order Box in source-image pixels, Text, Score (mean per-character confidence, 0 to 1), Value and UOM when the line reports a quantity (10% complete is 10 percent, About 3 minutes remaining... is 3 minute)
DetectElements, DetectInteractables one Element per button, icon, text line or progress line ID (index in reading order for this frame), Key (role+name locator), Kind (button, icon, text, progress: a text line with a quantity), Box, Score, Text, Value and UOM on progress lines (and on a button or icon whose label carries one), Style
DetectScreen a Screen Size, Theme (light or dark), Backdrop (dominant colour), Elements with Style filled in
Style, on every element Background and Foreground colours, Contrast (WCAG ratio, 1 to 21: readable text is above 4.5, greyed-out below 3), Disabled, Highlighted (accent fill: a selected row, a default button), Focused and the Ring colour
Diff two frames compared the bounding box of what changed and the fraction of pixels that did (0 when settled, 1 when the sizes differ)

One element from the macOS Recovery fixture, as textshot -json and json.Marshal render it. Boxes flatten to x0,y0,x1,y1, colours are #rrggbb, text is omitted when empty and ring unless focused:

{"id":25,"key":"button:Continue","kind":"button","x0":1160,"y0":554,"x1":1243,"y1":575,"score":0.75,"text":"Continue",
 "style":{"background":"#f5f5f5","foreground":"#bababa","contrast":1.78,"disabled":true,"highlighted":false,"focused":false}}

The focused list icon on the same screen carries "focused":true,"ring":"#6fa2e7", and the Screen marshals as {"width":1920,"height":1080,"theme":"dark","backdrop":"#34343a","elements":[...]}. Runnable examples for each call are on pkg.go.dev.

How elements are labelled

DetectElements fuses OCR with the element detector: an interactable region with text inside is a button (push button, menu item, list row, link), one without is an icon (toolbar icon, window control, checkbox), and text outside any interactable is text, or progress when it reports a quantity: {"kind":"progress","key":"progress:10% complete", ...,"text":"10% complete","value":10,"uom":"percent"}. An icon with a title right next to it on the same row (a list row's icon) takes that title as its Text while staying an icon. ID is the element's index in reading order for that frame (a set-of-marks number). Key is a role+name style locator (button:Continue, text:Disk Utility, icon:Disk Utility for the row's icon, icon@700,310 when unlabelled, #2 on repeats) that stays the same across frames while the layout does not move. Find an element by Key, confirm it by Box, then act. The element model (232 MB) is downloaded and loaded on the first elements call, so engines that only read text never pay for it.

DetectScreen adds what the pixels say, with no model: Disabled is a button or text line whose contrast is below 3, Highlighted a saturated background, Focused a saturated 1 px ring around the box, and Theme follows the backdrop's luminance. Colours are measurements, not design tokens: exact hex values, font families and point sizes are not recoverable from a screenshot.

Typography

For callers that repaint text (masking, redaction, translation overlays):

faces := []textshot.Face{{Name: "Inter Medium", Family: "Inter", TTF: interTTF}, ...}
tf, _ := textshot.DetectTypeface(img, line.Box, line.Text, faces)
// tf.Face, tf.SizePx, tf.Bold, tf.Italic, tf.SlantDeg, tf.Stroke, tf.Similarity

DetectTypeface renders the line's text in every face you supply, scales it to the crop's ink height and keeps the best pixel overlap. It never names a font you did not pass: the answer is "which of mine reproduces these pixels best, and how well". Size, slant and stroke weight are measured on the pixels alone and are reported even when no face fits.

Pure helpers

For callers with their own pipeline or build constraints:

  • MergeElements(interactables, lines) does the button/icon/text fusion and keying from detector boxes and OCR lines; ComposeScreen(img, interactables, lines) adds Style and Theme on top. Neither loads a model.
  • Lines(results) is the text of each result, in order.
  • ParseQuantity(line) reads the number with a unit of measure a line reports ("About 3 minutes remaining..." is 3 minute, "10% complete" is 10 percent, "1 hour and 5 minutes" is 65 minute; units are second, minute, hour, day, percent). It is what fills Value/UOM on text results and elements. Quantities lists every such number on a line; ParseDuration and ParsePercent return a time.Duration or the percentage. Bare numbers and clock times are ignored.
  • Prefetch and PrefetchElements download the runtime and the models into the cache without loading them. They need no cgo, so a CGO_ENABLED=0 build can warm a cache for another.
  • errors.Is(err, textshot.ErrUnavailable) from New means OCR cannot run on this build (no cgo) or host (no ONNX Runtime release); the message names the remedy.
  • The pins are public for inspection: DetModel, RecModel, ElementModel and RuntimeAssetFor(goos, goarch) carry name, URL, SHA-256 and size; DefaultCacheDir() is where they land.

Options

Option CLI flag Env Default
CacheDir -cache TEXTSHOT_CACHE_DIR os.UserCacheDir()/textshot
LibraryPath -lib TEXTSHOT_ONNXRUNTIME_LIB downloaded into the cache
MaxSideLen -max-side 1920 px (downscale only)
Threads -threads runtime.GOMAXPROCS(0)
MinScore -min-score 0.5
ElementMaxSideLen -element-max-side 1280 px (downscale only)
ElementMinScore -element-min-score 0.1
Logger -q, -v discard

The CLI also has -t, which prints each image's line or element count, the screen theme and the elapsed time to stderr, and -prefetch.

Platforms: linux/amd64, linux/arm64, darwin/arm64 (ONNX Runtime 1.30.0), darwin/amd64 (1.23.2, the last Intel macOS release). Other platforms can point LibraryPath at a local libonnxruntime.

Per 1080p frame on two CPUs: OCR about 650 ms at full resolution, 430 ms at MaxSideLen: 1280; element detection about 4 s at ElementMaxSideLen: 1280 (the YOLOv9-E detector is heavy, so call it on demand rather than in a polling loop). The element session tries the platform GPU execution provider (CoreML on macOS, CUDA elsewhere) and falls back to the CPU when the loaded runtime does not have it.

Comparison

Alternative Use it when textshot instead when
Accessibility APIs, chromedp, rod, Playwright You control the app or browser and can ask it for its element tree: exact roles, names and states All you have is pixels: a VM console, a boot or recovery screen, a remote desktop
gosseract / Tesseract Document OCR with many languages, and a system libtesseract is acceptable Screen text at UI sizes, no system packages, and buttons and icons fused with the text
PaddleOCR, RapidOCR (Python) A Python pipeline that already has the same PP-OCR models A Go program, with the models downloaded and pinned by digest at first use
Cloud OCR (Google Vision, AWS Textract, Azure AI Vision) Highest accuracy on documents and handwriting, no local compute Offline, no per-call cost, and screenshots never leave the machine
OmniParser (Python) A GPU and a captioning model for icon descriptions The same icon_detect_v3 detector on CPU, fused with OCR into role+name keys, from Go

Limitations

Everything is derived from pixels, so some things are out of reach and some are known weak spots. Status and the package to change:

Limitation Status Where
A macOS menu bar OCRs as one line ("Terminal Shell Edit View Window Help"); the detector's per-item boxes carry no text open, the det probability map saturates across the bar at every threshold tried internal/ppocr
11 px monospace text can lose a thin leading glyph (-bash-3.2# reads as ash-3.2#) open internal/ppocr
Logos read as elements (the TianoCore logo is a Disabled button); low-score icons on terminal output are detector noise by design, raise ElementMinScore to 0.25 if they matter internal/uidet
Window and panel hierarchy out of scope for now: a flood-fill region detector was tried and dropped, nested panels break the fill-ratio rule internal/screen
Mouse cursor out of scope: the template depends on OS, theme and scale
Font family and size, exact design-token colours not recoverable from a screenshot; Style reports measured colours only
Checked state, element classes beyond button/icon/text out of scope: needs a multi-class model internal/uidet

The rules and thresholds behind these are in the pipeline section of AGENTS.md.

Development

task test           # hermetic
task test:e2e       # TEXTSHOT_E2E=1, downloads runtime + models once
task bench
task models:export  # re-export icon_detect_v3 to ONNX (needs uv; one-off, see below)
task models:check   # verify dist/models/*.sha256 are pinned in models.go

See CONTRIBUTING.md for setup, where things live and the design rules, and CHANGELOG.md for what changed. CI runs the hermetic tests on every push and pull request.

The element model ships upstream only as TorchScript. task models:export converts it with a pinned PyTorch environment (tools/export-icon-detect, driven by uv) and writes dist/models/icon_detect_v3.onnx plus its SHA-256. Pushing a models-v* tag runs the same export in CI (.github/workflows/release.yaml), checks the digest against models.go, and publishes the file as a release asset, which is the URL ElementModel downloads from.

See AGENTS.md for the package map, pinned artifacts and cache layout.

License

Apache-2.0. The embedded ppocrv6_dict.txt and the PP-OCRv6 models are Apache-2.0 from PaddlePaddle/PaddleOCR. The pipeline is a port of lib-x/ppocr-v6-go. The element detector is icon_detect_v3 from microsoft/OmniParser-v2.0, MIT (the YOLOv9-E checkpoint; the older YOLOv8 icon_detect is AGPL and is not used).

About

Extract text and UI elements from screenshots in Go, via OCR and icon-detection models run through ONNX Runtime

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages