Skip to content

Releases: jerryjliu/fly_ocr

Fly OCR: compound eyes, characters, and a fixed connectome

Choose a tag to compare

@jerryjliu jerryjliu released this 13 Sep 02:07

Fly OCR sends glyph pixels through a fixed MaleCNS connectome and trains a small decoder on downstream neural activity. This first public release includes calibrated letter recognition, original-input comparisons, numeric tables, an honest compound-eye replay, and a 13-page research report.

  • 166,700 retained neurons and 25,582,938 frozen connections.
  • Calibrated 68-class benchmark: 87.6% overall; 85.0% on letters. Small, font-dependent evaluation; not general OCR.
  • 2:56 compilation showing letters, numeric tables and formatting failures, with actual errors preserved.
  • 11 source recordings, compact trained checkpoints, a static local viewer, reproducible scripts, source licenses and a clean export manifest.
  • Local checks: 26 tests; 983 recorded events verified; 12 videos fully decoded; browser controls and isolated installed inference checked.

The two-eye view is a schematic atlas of recorded sampled light. It does not change recognition and is not a reconstruction of fly optics. Preprocessing and table/string assembly are conventional geometry outside the circuit. No biological literacy or practical OCR advantage is claimed.

Download the video, PDF report, or full source-and-assets archive below. The README includes local setup and reproduction steps. The large source connectome and local caches are downloaded separately, not included in the archive.

Fly OCR: a 40-second results-first teaser

Choose a tag to compare

@jerryjliu jerryjliu released this 13 Sep 03:39

A 40-second introduction to Fly OCR, presented in inverted-pyramid order: the accomplishment first, evidence next, a little context, then the code and research report.

  • Opens during final-model text recognition, with the recorded word Cash already visible.
  • Shows a successful table of values: 21 of 21 exact cells in the selected crop.
  • Introduces the fixed fly circuit and small trained decoder after demonstrating the result.
  • Keeps the animated 3D fly and actual recognition mistakes.
  • Leaves development history, input comparisons, ablations and diagnostic counters to the detailed report and longer walkthroughs.

The 1:45 social walkthrough and 2:56 original compilation are preserved. The README now leads with finished capabilities too.

MP4: 1920 × 1080, 24 fps, H.264, 960 frames, captioned and silent. Includes the updated audited repository archive. Rendered event traces, full video decode, source hashes and previous-video checksums verified; viewer build passed.

Rendering and edit details: docs/teaser-video.md.

Fly OCR social cut: a fly attempts to read a PDF

Choose a tag to compare

@jerryjliu jerryjliu released this 13 Sep 03:13

A separate 1:45 social cut opens on character recognition from the first frame and puts a rendered 3D fly on the PDF beside the recorded results.

  • All seven OCR examples remain: letters, heading, more labels, two numeric tables, and the 3-degree tilt failure.
  • The fly follows recorded glyph positions through illustrative animation. It does not control or alter recognition.
  • Raw predictions, errors, retinal samples, and neural activity remain from the original saved runs. No retraining.
  • 1920 × 1080, 24 fps, H.264, 2,520 frames; captioned, without audio.
  • Original v0.1.0 compilation and research report preserved.

Includes the MP4 and an updated, audited public repository archive. Rendering instructions and visual provenance are in docs/social-video.md.

Validation: full video decode; every glyph covered; original compilation checksum unchanged; all 11 inference recordings (983 events) verified; viewer build and replay checks passed.