AutoFileEmail
    IntegrationsSolutionsComparePricingFAQBlog
    Sign inConnect Drive
    1. Blog
    2. Why We Don't OCR Your Receipts (And Why You Shouldn't Want Us To)
    On this page
    Why OCR exists at allYou probably don't keep booksWhat OCR costs you in privacyWhat OCR costs in dollarsWhat we actually see"But I want my data extracted"The bottom line

    Why We Don't OCR Your Receipts (And Why You Shouldn't Want Us To)

    MMitchel Kelonye
    •
    May 15
    •
    Privacy
    Receipts
    Anti Ocr
    Data Security

    Studio Ghibli-inspired banner about privacy-first receipt filing for a calm, modern workspace

    Every receipt-management app on the market reads your receipts.

    Hubdoc OCRs them. Dext OCRs them. Receipt Bank, AutoEntry, Expensify, Bill.com - all of them parse the line items, extract the totals, and store the structured data on their servers. That's the whole pitch: "we read your receipts so you don't have to."

    We don't do that. Not now, not ever.

    This article is the why. It's also the longest case I'll make for treating "we don't read your stuff" as a feature, not an absence of a feature.


    Table of Contents

    • Why OCR exists at all
    • You probably don't keep books
    • What OCR costs you in privacy
    • What OCR costs in dollars
    • What we actually see
    • "But I want my data extracted"
    • The bottom line

    Why OCR exists at all

    OCR on receipts isn't a feature someone invented because it was cool. It's a workflow tool for a specific job: bookkeeping.

    A bookkeeper takes a stack of receipts and turns them into a general ledger. That ledger has columns for date, vendor, account, debit, credit, memo. Each receipt becomes one or more journal entries. Every month, those entries get reconciled against bank statements.

    This is double-entry accounting. It's a real thing. Real businesses need it. The tools that came up around it - QuickBooks, Xero, Wave - all assume you're keeping books on a monthly cadence and need to reconcile.

    OCR shows up because the boring step of "type each receipt's vendor, date, and amount into a journal entry" is hideous data entry. You can pay an accountant $50 an hour to do it, or you can buy software that auto-extracts the fields and lets you review them. OCR is the cheaper-than-a-bookkeeper solution.

    Notice what just happened. OCR is a bookkeeper-replacement tool. It only matters if you're keeping books.

    Receipts and bookkeeping concept showing OCR exists as a workflow tool in Studio Ghibli style

    You probably don't keep books

    If you're a solo founder, freelancer, indie hacker, or one-person agency, here is what actually happens at tax time:

    You collect a year of receipts. You hand them to a CPA. The CPA categorizes them - what's software, what's office expense, what's professional development. The CPA totals each category. Those totals go on Schedule C. Tax return done.

    You did not do double-entry bookkeeping. You did not maintain a ledger. You did not reconcile against bank statements monthly. You did the once-a-year handoff that has been the freelancer norm since freelancing existed.

    Your CPA does not need OCR'd data from you. Your CPA needs the PDF. They have their own software (UltraTax, Drake, Lacerte) that has its own data extraction if they want it. What they want from you is organized files.

    The freelancer who buys Hubdoc is buying a forklift to move one box. Once a year. The other 364 days the forklift sits in their garage at $35 a month.

    Solo founder realizing they probably don't keep books

    What OCR costs you in privacy

    Let's be specific about what gets sent where.

    When Hubdoc or Dext OCRs your receipts, here's the data flow:

    1. The PDF arrives in their cloud (uploaded by you, or fetched from your inbox).
    2. The PDF gets converted to images.
    3. The images go through their OCR engine (often a third-party API like Google Document AI, Amazon Textract, or Microsoft Form Recognizer).
    4. The structured output - line items, vendor, date, total, tax, payment method - gets stored in their database, indexed and searchable.
    5. The original PDF stays in their storage.
    6. You get a dashboard view showing all your receipts as structured rows.

    Step 3 is the part most people don't think about. To OCR your receipt, the receipt has to be sent to an OCR engine. That engine, depending on the vendor, may keep the data, train models on it, or ship it to a third-party processor.

    What's in a receipt? Vendor name. Date. Amount. Payment method (often the last 4 digits of a card). Sometimes line items so detailed they reveal exactly what services or products you bought. Adobe knows you're on Creative Cloud. Stripe knows your monthly revenue. AWS knows your infra footprint. AWS knows it because they sent the invoice; the receipt-management company knows it because they parsed the invoice.

    Multiply by every solo founder with their entire vendor stack on the platform. Now there's a database that can answer "which SaaS tools is the solo-founder market using" with line-item granularity. That database has a price tag. Some company will eventually pay it.

    Maybe today's receipt apps are responsible. Maybe they're not. The point is: you have no control once the data is parsed and structured. You can delete the PDF; the extracted data lives separately.

    We avoid all of that by not parsing in the first place. Your PDF lands in your Drive. Drive sees it. Google sees it. We see the metadata we need to file it (sender, date, filename) and nothing inside the file.

    The full privacy comparison versus competitors lives at /vs/hubdoc if you want the side-by-side.

    What OCR costs in dollars

    Receipt apps charge what they charge because OCR isn't free.

    Modern OCR runs on GPU compute. A single PDF takes a few cents of GPU time to parse, plus storage, plus the indexing step. At scale, with one user uploading 500 receipts a year, that's a few dollars in raw compute per user per year.

    Add the OCR vendor markup, the SaaS company's gross margin target (75% plus), and the marketing payback period, and you get pricing like:

    • Hubdoc: bundled in QuickBooks Online plans at $35+/mo
    • Dext: $26/mo for Solo, more per document tier
    • AutoEntry: $12/mo for the smallest plan, then per-document
    • Expensify: $5/user/mo for the cheap plan, but with submission caps

    If you're a solo founder, you're paying $144 to $420 a year so a SaaS company can read every line of every receipt you get. For a once-a-year tax handoff. Where your CPA throws away the OCR data and reads the PDF anyway.

    We don't have a GPU bill. We have a "list directory in Drive, save attachment, increment a counter" bill. That bill is small enough that we can offer the free tier free for one inbox forever. Not free for 30 days. Not free until you hit 100 receipts. Free.

    The Backfill Packs ($29 Tax Year, $59 3-Year, $99 Lifetime) cover the one-time cost of pulling old emails out of your inbox and filing them. Forward filing - the day-to-day "new email arrives, file it" - stays free.

    We don't charge per receipt because we don't pay per receipt. The math composes correctly.

    Cost of OCR in dollars represented with receipts and price tags

    What we actually see

    For full transparency, here is the entire metadata we store per filed document:

    • Sender email address (to derive the folder name)
    • Date of the email
    • Subject line (so the dashboard can show "Stripe payout - April 14")
    • Original filename of the attachment
    • A pointer to where the file lives in your Drive
    • A unique message ID from your mail provider (so we don't double-file)

    That's it. We don't store the PDF on our servers. We don't open it to look inside. We move it from your inbox to your Drive, log the move, and stop. The file lives in your Drive, under your account, with the drive.file scope which means we can only see folders we created and never the rest of your Drive.

    The detailed access scopes are listed in our privacy policy. The TL;DR: we have less access to your data than your IDE does.

    "But I want my data extracted"

    Some people legitimately want OCR. If you're a small business with $500k revenue, doing real bookkeeping monthly, with a real bookkeeper, the OCR tools earn their keep. Buy them.

    Our user is the one-person shop, the indie hacker, the moonlight freelancer. The person who wants their PDFs in Drive and their CPA happy. The kind of person whose entire tax workflow fits in three folders. That's it. If your needs are bigger, we're not the right tool, and we're never going to add OCR to compete - it would change the privacy contract and the price.

    This is also why we'll never build per-email metering. The cost structure that lets us not OCR is the same cost structure that lets us not meter. They're the same decision wearing two hats.

    The bottom line

    OCR is a bookkeeper-replacement tool. Most solo founders aren't keeping books. So OCR is solving a problem they don't have, at a price they shouldn't pay, with a privacy cost they didn't agree to.

    Filing - just moving the PDF to a sensible folder - is the part that's actually missing. We do that. Nothing else. By design.

    Try the version that doesn't read your stuff. Free for one inbox at autofile.email. Two-minute setup. Connect Gmail, connect Drive, walk away.

    If you want to compare us head-to-head with the OCR-everything tools, see how we stack up against Hubdoc. The tradeoffs are explicit.

    We don't read your receipts. We just file them. That's the whole thing.

    The last time you'll dread tax season.

    Connect Gmail and Drive, watch the 30-day preview file itself, and never think about new email attachments again. Forward filing is free, forever. When tax season comes, grab a Backfill Pack and we'll sweep the rest of your history.

    Connect Drive — free See pricing

    Thanks for reading! If you want to see future content, subscribe to our RSS feed.

    ← Older
    Auto-File iCloud Mail Attachments via IMAP
    Newer →
    Gmail Attachments Are a Tax-Time Landmine (Here's the Fix)
    AutoFileEmail

    We don't read your documents. We just file them. Receipts, invoices, statements — sorted into your cloud, automatically.

    Product
    • Pricing
    • FAQ
    • Blog
    • About
    Compare
    • AutoFileEmail vs. Receiptor AI
    • AutoFileEmail vs. Hubdoc
    • AutoFileEmail vs. Dext
    • AutoFileEmail vs. Zapier
    • AutoFileEmail vs. CloudHQ Save Emails to Drive
    Integrations
    • Auto-save Gmail attachments to Google Drive
    • Auto-save Gmail attachments to Dropbox
    • Auto-save Outlook attachments to OneDrive
    • Auto-save Outlook attachments to SharePoint
    • IMAP + Google Drive Integration
    Solutions
    • AutoFileEmail for Freelancers
    • AutoFileEmail for Bookkeepers
    • AutoFileEmail for Landlords — Schedule E ready by April
    • AutoFileEmail for Consultants
    • Tax-Time Receipt Organization
    Legal
    • Privacy policy
    • Terms
    • Security
    © 2026AutoFileEmail · We don't read your documents.PrivacyTerms