Most organizations store scanned PDFs, receipts, and image files in SharePoint — but without OCR, Copilot and Microsoft Search can’t read or return them.
In this video, Microsoft MVP Drew Madelung joins me to walk through how Optical Character Recognition (OCR) works in Microsoft 365. You’ll learn:
✔️What OCR does and how it enhances SharePoint search and Purview
✔️Which files are supported (and what won’t work)
✔️How to estimate costs using the built-in OCR estimator
✔️Tips for enabling OCR in both SharePoint and Microsoft Purview
✔️How OCR integrates with Copilot and DLP policies
🧠 If you want your files to be searchable — and secure — this is a must-watch.
🔗 Learn how to enable Pay-As-You-Go billing in this episode
📺 Watch the full SharePoint Content AI series here
Watch my 100+ courses on Pluralsight
Video Summary
- OCR in Microsoft 365 makes your images and scanned PDFs searchable, unlocking their value for Microsoft Search, Copilot, and compliance tools like Purview. It extracts both printed and handwritten text and indexes it automatically.
- You need to enable OCR manually—it’s not on by default. You can turn it on in Microsoft Purview and even target specific SharePoint or OneDrive sites to control scope and cost.
- OCR enhances compliance and security by allowing DLP and other Purview policies to detect sensitive information in images. Just note: embedded images in emails or documents won’t be scanned unless they’re added as actual image files.
- It’s cost-effective, at around $1 per 1,000 items (each image or PDF page counts as one). There’s built-in deduplication, and you can use the OCR Estimator in Purview to preview your costs before enabling it.
- Pro tips for power users: The extracted text is stored in a column called
MediaServiceOCR, OCR only works on files under 50MB, and it runs in the cloud—not locally—so plan for network traffic. Additionally, OCR in communication compliance (such as Viva Engage) utilizes a separate model.
For more information, read the transcript blog below, or watch the video above!
Transcript
SharePoint has always been great at storing files, but many organizations have scanned PDFs of documents, contracts, invoices, or maybe even have them as images. And don’t get me wrong—SharePoint can store them—but because they’re not searchable, you lose a ton of value. Not only can people not search for them, but AI tools such as Copilot are also unlikely to be able to return them. But what if I told you that with just a few clicks, you can fix all that? In this video, I’m joined by Microsoft MVP Drew Madelung, who will help us deep dive into enabling Optical Character Recognition or OCR in SharePoint.
Drew, thank you so much for joining me. Thank you, Vlad, for having me. Today, I’m going to discuss OCR inside Microsoft 365. My name is Drew Madelung, and I’m glad to talk about this today. What we’re going to be going through is first an intro to what OCR in Microsoft 365 is, talking about where it works inside of the platform as a whole, how to estimate your costs—which is one of the main things we need to understand, as this is a consumption-based scenario—and some tips that I’ve seen in working with OCR that you need to understand as you move into enabling and working with OCR inside of your organization.
And for those of you who have been following the series, this is the eighth episode in the series. And of course, if you just got here as the first video, I’m sure you’ll enjoy it. Make sure you check out the other episodes in the series. We’ll have a link to the playlist in the description below.
The first thing we’re going to start out talking about is what OCR is and how it works. Optical Character Recognition is a service built into Microsoft 365. You might have thought that it had always been there, but it hasn’t. So, you could be uploading your files or images today and thinking they’re going to be indexed—that they would be available to you in things like Copilot—but they’re not in the way that you would expect.
Now, with OCR inside of SharePoint and Microsoft 365, it allows you to actually extract printed and handwritten text from the images and documents and files that you have inside of SharePoint. What happens is that, let’s say I upload a file, the actual text or the handwritten text is extracted from that information and then indexed in search, which then makes it findable inside your platform. That indexed data is not just available in search. One of the best things that you get as part of this is that that’s how you enables advanced Purview features that use OCR.
So, if I’m sending an email that has a receipt with PII in it or proprietary information and you have policies set up that would allow you to detect that sensitive information—when you enable OCR, the Purview features can now be enlightened by that information to allow you to find that information in your DLP policies to help protect your organization more. When this is set up, there’s actually going to be a new metadata column added to the areas in the site that have OCR enabled. There’s a column called the “extracted text” column, so you can visually see what is happening as part of OCR.
Now, other files might not use that information or don’t use that column, but they are still indexed as part of the Purview functionality. One concern you might have is that you have a ton of files—maybe millions—and wonder about duplication. If I upload the same receipt 10 times, you’re not going to be charged from an OCR perspective for all of that, as there’s native deduplication as part of the OCR service to ensure unique images are scanned. So, you don’t have repeated scans and costs as you’re working through this.
From a Purview point of view, this can be so important for organizations because if you don’t have that enabled and end up not giving or retaining the right information, that can have a huge financial impact. So, it’s a must. What I’m seeing the most is that people don’t know that it doesn’t exist. By enabling this, you gain a base understanding of what’s possible. If they had ever run any DLP tests or policy validations, they wouldn’t know that it’s not working. You do need to set this up, and I’ll go through how to set it up in both locations. It’s not just a single setup to empower you for every scenario.
Also, this is available in over 150 languages, so it will be able to extract text not just in English but in multiple languages, depending on the area you’re in. It’s not a language issue like we might see in other content AI solutions today.
So let’s take a look at what OCR looks like. I’m going to come into SharePoint Online, and I have a site called Content Processing Operations. You can see I have a collection of images—some golf courses, background images, and a picture of myself. Let’s look at one down here that has a sign sitting on a golf course. I’m going into a different view—a list view—and you’re going to see these files that are in this library have a new field called “extracted text” and “image tags.” These two columns contain text extracted from the image I uploaded. This column is just like any other SharePoint column—it’s indexed, it’s a crawled property, and it’s what SharePoint and other solutions will use to render and return values in it.
So I’ll take a receipt—a standard receipt with information. It’s a JPEG file, and OCR will extract the text. It may take a few minutes, especially the first time you configure it, but within a set time, it will extract it. That one was fast—it extracted “online receipt,” “east repair,” and other details. You don’t have control over what is extracted, but even partial extraction is better than nothing.
Now, let’s say I take a document with a receipt image inside it. The “extracted text” column is not filled in for that one, but the data is still indexed. OCR still happens on that file, but Office documents now inject the extracted text directly into the index. That means Copilot and Microsoft Search can still use it—you just won’t see it through the SharePoint UI today. That’s something you noticed the first time, too—you expect it to show up visually, but it doesn’t always. The biggest thing is to trust the system. If you want to test it, run a standard Microsoft Search for keywords in that image—it will return a match.
OCR has supported file types. Not every file type is supported, but it hits the core ones—PDF being the biggest. If you’re doing any autoscan or scan capture with content AI, you can stack those solutions together. It also supports core image formats like JPEGs and PNGs. Purview workloads around Teams, Exchange, and Windows have a smaller supported file type subset. Your main ones for end-user search are in SharePoint and OneDrive.
The images have to be added, not just embedded. If you embed an image in an email, it won’t be OCR’d. It needs to be added from your device. Images also need to be under 50MB, and the maximum pixel resolution is 16,000 x 16,000. The bigger risk is actually with very small images, not scanning well.
Images are only scanned after OCR is turned on. It won’t go back and scan all your old PNGs. It requires some interaction, like modifying or uploading the file. There’s no manual trigger like with autofill columns. But from a Purview perspective, scanning happens on interaction or in transit (like sending or sharing).
As it would be an action on the file, be aware that files sitting at rest won’t get scanned until modified. The key is to turn it on now so it starts indexing files as you go. Before we cover the most important part—how to turn it on—you mentioned that images must be under 50MB, but documents can be over that size. Documents can exceed that as long as each image inside is under 50MB. For PDFs, the whole file must be under 50MB.
Now, about cost and setup. There is an OCR Estimator tool in preview. You can’t run it after OCR is enabled. You need to go into Purview, go to Settings, and click OCR. If you’re there for the first time, you’ll see “Try for Free.” It scans your organization and provides an estimated OCR cost based on a set timeframe. You can download and compare reports by rerunning the estimator. You can choose to scan Exchange, SharePoint, OneDrive, Teams, and enrolled devices.
Setting this up is part of Microsoft 365 Pay-As-You-Go services. What’s unique about OCR is that you can pick specific SharePoint and OneDrive sites to enable OCR. You don’t have to do the whole tenant. You can upload a CSV to define your target sites. There’s no dynamic way to maintain it—it’s manual. This helps control cost and lets you pilot it on certain sites.
You cannot do a subset of Teams—it’s SharePoint-specific. For those who haven’t watched the full series, there’s a separate video on how to enable pay-as-you-go services and link them to your Azure subscription. That video covers best practices and billing setup.
Lastly, to set up OCR in Purview: After running the estimator and enabling pay-as-you-go, you still need to configure Purview OCR separately. Just enabling OCR in SharePoint doesn’t activate Purview OCR. Go back into Purview, pick the locations you want to enable it for, and your existing Purview policies will inherit the OCR capability. You don’t need to reconfigure DLP policies—they will automatically use the OCR settings.
Sensitive Information Types and Trainable Classifiers will also inherit the OCR config. For example, if you’ve built a custom type for Social Security Numbers, it will now apply to OCR-extracted content.
To recap, there’s a SharePoint-only experience and a Purview experience. If you want Copilot to find content in SharePoint, you don’t need to configure Purview. But if you want compliance value, you need to do both. Often, compliance teams drive this from the Purview side, and SharePoint admins benefit from the added visibility.
The cost is about $1 per 1,000 items. Each page of a PDF or each image is one item, and each unique image is only scanned once. While it’s low-cost overall, it can add up at scale. Use the estimator to plan accordingly.
It’s also worth noting that OCR doesn’t support DLM or Records Management labels on endpoint devices. So, if you have images on devices and want to apply retention labels, that won’t work—only SharePoint and OneDrive are supported.
Some final tips: the internal name of the extracted text column is “MediaServiceOCR.” If you’re using custom search solutions, you’ll need to map that to a refinable string. You can add the extracted text column to your document library views, but it won’t show up until after OCR is enabled.
Both incoming and outgoing emails are subject to OCR. However, Exchange DLP policy tips don’t support OCR, so you won’t see those in Outlook or OWA. The OCR scan happens in transit. OCR processing doesn’t happen locally—it’s done in the cloud. So be mindful of network usage, especially on enrolled devices.
Communication compliance tools like Viva Engage have their own separate OCR engine. This setup doesn’t affect that—it’s a separate model already in use.
Lastly, OCR requires Pay-As-You-Go. Please refer back to the other video in the series for billing setup details.
Drew, that was really—thank you so much for this amazing presentation. And for everybody else, if you enjoyed this presentation, please make sure to check out all the other videos in the series. You’ll have a link to the playlist appearing on your screen right about now. If you want to connect with Drew, all his LinkedIn and social media profiles are in the description below. Drew, thank you so much again for taking part in this video.