Master SharePoint AI: Extract Data from Unstructured Documents (Invoices, Contracts & More!)
Struggling to extract data from documents that don’t follow the same template? Single Class Document Extraction in SharePoint Content AI has you covered!
In this episode of the Ultimate SharePoint Content AI Deep Dive Series, Vlad and Microsoft MVP Gokan Ozcifci guide you through extracting metadata from invoices, contracts, proposals, and other unstructured documents. Learn how to set up, train, and deploy AI models directly in SharePoint—without needing a separate module for every document type.
What you’ll learn:
✔️ Setting up Single Class Document Extraction
✔️ Training models to classify and extract metadata
✔️ Best practices for accurate data extraction
✔️ Common scenarios like invoices and expense reports
📌 This is part of our Ultimate SharePoint Content AI series — check out the full playlist to go even deeper.
Watch my 100+ courses on Pluralsight
Video Summary
- Single-class document extraction in SharePoint lets you train AI to extract metadata from unstructured documents like invoices or contracts—no fixed templates needed.
- It’s built into SharePoint Premium and uses a pay-as-you-go model (about half a cent per page per value), making it a great option if you don’t have AI Builder credits.
- You train the model by uploading sample documents, labeling them as “good” or “bad,” and then defining what data to extract using either custom rules or Microsoft’s templates.
- Once trained, the model can be reused across multiple document libraries and can even leverage existing content types and columns—great for information architects.
- After applying the model, SharePoint automatically processes new uploads, extracts the data, and displays it in your library view—fast, accurate, and scalable.
For more information, read the transcript blog below, or watch the video above!
Transcript
When all your documents follow the same exact structure, extracting data from them using AI can be easy. But what if you have a dozen types of documents—whether it’s invoices, contracts, or proposals—that do not follow the same template? This is where single-class document extraction comes to the rescue. It helps make sure that you can extract metadata from unstructured content without having to create a separate module for each one.
In this video, I’m joined by Microsoft MVP Gokan Ozcifci, who will deep dive into single-class document extraction. “Gokan, thank you so much for joining me.” “Thanks for having me, Vlad. And yes, you pretty amazingly explained this. I guess people already know me, but hey, my name is Gokan, and I’m going to be part of the SharePoint Content AI series, explaining a few modules. We also have a bunch of other videos, as you probably mentioned in other videos. Please see and check them.”
“Do you want to add something to that slide, Vlad?” “Sure. Yeah. So this is part five of the Ultimate SharePoint Content AI series. If you’ve been watching the series, I hope you’ve been enjoying it so far. If this is the first video that you joined us on, after you’re done with this video, make sure you check out all the other videos so you learn not only how to use extraction for unstructured data but all the other options you have with SharePoint Content AI as well.”
“Absolutely. So if I take it over now—and thank you so much, Vlad, for the nice introduction—we have multiple models today with SharePoint Content AI, and we will see all of them. But today’s focus is going to be on the unstructured, on a single-class model, which is in the middle. SharePoint Content AI, or SharePoint Premium, offers you a bunch of models that you could use in order to extract data, in order to extract information from your data and use that in multiple other forms. And today’s is the most complex one, and that’s what I tell my customers. The single-class is the most complex model that you could use because it’s the SharePoint-based one. It’s up to you to define the extractors. It’s up to you to define the knowledge. It’s up to you to define the logic in your model in order to get all that data. We will talk about the semi-structured and the structured in another video, but today it’s all going to be on the SharePoint-based one. We also refer to this as the teaching model. Basically, you will extract sentences or a specific region in a document. So we will go in our document, in our model, and tell it, ‘This is the piece of info you need to extract for me and bring that into my column.’ Models are created in content centers and are applied to the document library where you create them. So you can either create the model in a document library and apply it to that document library, or you can easily go to a content center and then create that model in that content center and then apply it to any document library you have access to.”
“The biggest thing that I think is phenomenal with those models is that it can be associated with a content type. For those who have been using SharePoint for a long time and you’ve been doing that right from an information architecture perspective, it’s phenomenal to be able to reuse the content types, the site columns you already have in your site, and not always create new content types for each model you create. Two words that you need to learn are the classifier and the extractor. The classifier is used to determine the type of the document, and the extractor is the data that we want to pull out and bring into our document library. And the good thing is that we are not limited by the AI Builder credits, which is on the structured side of the story.”
“Awesome, right?” “It’s pretty awesome. I don’t really like slides. I only have one more slide, and then we’ll deep dive into the demo.” “Knowing you, I’m surprised you have that many already.” “Visually, it’s like Gokan hates slides—only a few slides. And I also prepared for you a comparison slide between the unstructured and the structured. We will now look at the unstructured side of the story. You’ll see that the model creation is done in SharePoint. The classification type is the classifiers and extractors, as explained. The supported file types—we have to give it four good and one bad option to select from. And then you’ll see all the metadata support, supported regions, the cost, capacity, and so on and so on, which will be easier in a demo.”
“Is that good, Vlad? Shall we go for the demo?” “I can’t wait to see it live.” “There you go. Fantastic.” Before we go into the creation of the model, I want to highlight something really important. When you go as an admin to the admin.microsoft.com website, under the Setup section on the left side, you’ll see an option at the bottom called “Activate Pay-As-You-Go Services.” When you click on it, you can enable pay-as-you-go for two different things. The first one is the Syntex services. It might be a bit slow to load, but once it does, you’ll see options like OCR, document translation, autofill—all the Content AI features we want to show in these videos. Then there are agents in SharePoint. If you click on Syntex services, you’ll see that you can link it to your Azure subscription. There’s already a video available on how to do that.
The most important part is the settings pane. Here, you can see all the services available for Content AI. You can enable or disable services for your SharePoint environment. Since we’re focusing on unstructured document processing, you’ll find that option here. If you click on it, you’ll see a few settings. You can allow people to create, train, and apply models to process files. You can enable or disable that option. On the right side, you can choose how many sites can use this functionality. You can edit this and say, for example, “I want up to 100 sites,” and then upload a CSV file listing the sites that can use unstructured document processing.
You’ll also see the content center. The downside is that while you can rename the content center, you can’t change its address from the UI. In my demo environment, if I click on it, you’ll see that it no longer exists. So, just bear in mind that you can’t change the default content center from the UI. You can always create more content centers, but the default one is fixed.
Let’s say we enable this and leave it open for everyone—anyone can create models on any site. I’m happy with that, so I’ll go ahead and create a new SharePoint site. I’ll click on “Create,” and it will ask if I want a team site or a communication site. Let’s go with a team site. Once you enable pay-as-you-go services, you’ll see two new templates: “Accounts Payable” and “Contracts Management” from Syntex. You can use those, which I like, but you can also go with a standard team site. I’ll choose a standard team site and name it “Vlad Talks Tech.”
For unstructured document processing, I’ll click next and then create the site. I’m doing this live to show that nothing is scripted—if it fails, it’s SharePoint’s fault, not Gokan’s! I’ll click “Finish,” and now I should see my brand-new site. I have my shared document library. If I click on it, give it a second—it’s taking its time. You said it was SharePoint’s fault, and now it’s taking revenge! There we go. You can see the shared document library. When you click on the three dots, you should see “Classify and Extract.” When you enable pay-as-you-go services and unstructured document processing in SharePoint, that “Classify and Extract” button should appear.
If you click on it, you’ll see two tabs: “Available” and “Applied.” All the models you’ve created will appear under “Available,” and you can see which ones are applied on the left side. I’m going to create a new model. It will ask if I want to activate it—yes, I do. It takes a couple of seconds, and then it says it’s been activated. Now I can create my model.
This is where we talk about single-class and structured forms. I’ll choose single-class—the most complex, SharePoint-based one. When you click on it, it gives you a bunch of examples: what you can do, supported file types, supported languages, and so on. Click “Next,” and then give it a name. Let’s call it “Gokan’s Travel Expenses.” Add a description. Under advanced settings, you can choose to create a new content type or select an existing one. This is really important for information and knowledge managers because they often already have content types and site columns prepared. You can create your content type, add your columns, and reference them here. But since this is a brand-new site, we don’t have any content types yet, so we’ll create a new one.
You can also apply sensitivity labels—internal use, public, encrypted, highly confidential, etc.—and retention labels. I have one called “Neoxy Customers – Never to Be Deleted.” You can apply that, and the data will never be deleted. Let’s create the model.
Now, within SharePoint, I’m getting a new tab within my site. I have that SharePoint Syntex-like experience with four main pillars: add example files, classify and train, create extractors, and apply the model. First, I’ll add files. I’ll upload some files from my sessions folder—let’s say five Uber invoices. I’ll upload them and click “Add.” Those are all valid. Remember, we need at least five good and one bad file. I’ll add more to make sure I meet the requirement. I’ll also upload a Word file to show that this works with both PDFs and Word documents. Now I have five good ones and one bad one.
Next step: train the classifier. I need to tell the AI model which documents are good and which are bad. The Word document is the bad one—it’s not a valid travel expense. I’ll mark it as negative. The others are correct, so I’ll mark them as positive. It’s a bit of monkey work, but it’s necessary. Now I’ll click “Train.”
Here’s the most important part: I need to tell the AI model how to recognize a valid invoice. Microsoft gives you a bunch of options. You can start from a blank explanation or use a template. I’ll go with blank. I’ll name it “Uber Invoices.” What’s the best keyword to search for in my invoice? Probably “Uber.” I’ll use that. Under explanation type, I’ll choose “Words.” I’ll click on “oo,” and the word “Uber” appears. I’ll accept it, save, and train.
Under the evaluation tab, it checks all the documents to see if they contain “Uber.” In the upper right corner, you’ll see the accuracy score. I got 100%, which means it found “Uber” in all the right places. If you get less than 100%, that’s okay, but I always tell customers not to go below 95%. If you can’t get 95% on your examples, your real-world documents probably won’t perform well either. Add more explanations if needed—maybe “destination,” “American Express,” or your home address—whatever helps the model understand the document.
Now we’re done. I’ll click “Test,” but I’m happy with the accuracy score, so I’ll exit training. Next, we go to the extractor. I need to tell the model what info I want to extract. The most important one is the value—how much did it cost? I’ll create a new column called “Value.” Under advanced settings, you can create a new column or use an existing one. Again, this is great for information architects who want to reuse columns. In our case, we’ll create a new one.
You can choose the column type: single line of text, multiple lines, date and time, number, or URL. I’ll choose “Number.” Once created, you can refine the extracted info—keep the first value, remove duplicates, set a default value, etc. For example, if the address is empty, default to your home address. This is very powerful and not everyone knows it exists.
Now I’ll label the documents again—mark the Word file as “No Label,” then go through each file, select the value I want to extract, and click “Next.” I’ll do this for each document. Finally, I’ll save. Now you will see all the values that I want to extract and put into my column called “Value,” just under the label. I click on “Train” again. This time, instead of using a blank explanation, let’s go for templates. Microsoft provides a bunch of options for explanations—email addresses, email sender, email subject, credit cards, percentages, and more. What I always use are the “after label” and “before label” templates.
So I have my value. If there’s a pattern before and after the value, I can use that. For example, I click on “After Label,” and it already prepopulates something. Let me cancel and show you. Look—always “Gokan.” So I have my value, and there’s a pattern: “Value Gokan,” “Value” with a dollar sign in front, “Value” with a dollar sign in front and “Gokan” in the back. There’s always a certain pattern before and after the value.
I’ll use the “After Label” template and add “Gokan” as a temporary label. I’m happy with that. You can add more. I’ll add a second one just to be sure and go for “Before Label.” It shows the euro sign. I can add the dollar sign. Can I give it more options, like one of ten things? Yes, you can add a dollar sign, “Gokan,” or anything that makes sense and might appear before your label. You don’t have to rely 100% on what AI suggests. It’s great that it gives you ideas, especially if you’re new—it guides you in the right direction.
When you have no idea what to fill in, it already gives you two examples based on patterns in your data. That’s how I understood how this was working, because honestly, the documentation wasn’t very clear. I’ll remove “Gokan,” add the dollar sign, and click “Save.” Now I have it all set up. I click “Train” again. It checks all the explanations to see if there’s a match. You’ll see it’s green, green, green—and my accuracy score is again 100%. Fantastic.
I’ll test and exit the training. The only thing left to do is apply the model to a document library. We can either do that in the shared one or create a new document library specifically for this. But for now, I’ll just apply it to the shared document library. I’ll close this. I did that just to show you that we applied it to the shared document library, but we can always create a new one and apply the model to multiple document libraries in your site. That’s one of the cool things you can do.
Now I have my documents here. At first glance, nothing seems to have changed. But if you go to your view and switch to “Gokan’s Travel Expenses,” you’ll see a new view created with the same name as your model. Look what I have here: “Value.” Remember the column we created? It’s here. Now, if I upload my Uber invoices—all of them—you’ll see something new. Its content is AI-related. It’s gone now, but it basically tells you, “We’re analyzing your files.” We’ll pause the recording to show it. It tells you, “Hey, we’re analyzing your files,” and within a couple of seconds, you’ll see your value extracted.
What I love is that you can see a few data points that are phenomenal—the processed status and other information that helps you know if the processing failed or succeeded. It’s on the roadmap that Microsoft will add a processing status to your models. And there you go—it’s finished. You can see it straight away. Very fast. You can see the value. Let’s check if it’s correct: 36.29. I click on “Uber 003,” and let’s see—36.29. Phenomenal. Very quick.
Let me refresh to see if the others are finished. You’ve got to close your eyes when you refresh—it always works better. Oh no, still not. You’re peeking through—that’s why! So basically, that’s how it works. Very easy. You can upload any data, and you’ll see the status and the extracted value in your library.
The real advantage is that it doesn’t matter if the value is at the top left or bottom right. As long as there’s a pattern—maybe a dollar sign before or “Gokan” after—it will look throughout the document. As long as it finds that pattern, it will extract the value. It doesn’t matter if it’s at the start or end—it will still find it.
So that was all about unstructured document processing. Very quickly, we created a model. If you want to modify your model, just go to “Classify and Extract.” You’ll see the applied model. If you go to the model details, you can either remove it from the library or manage the model. That brings you back into the model builder, where you can add more extractors. If you want to add the customer name, VAT, or destination address, just click on “New Extractor,” give it a name like “Address,” and click “Create.” A new extractor will be built for you. Then you follow the same steps—mark good and bad examples, and train it using blank or template explanations.
“Gokan, I know I need to set up pay-as-you-go to make this happen.”
“Absolutely.”
“Is it free after that?”
“No. Well, I tried. You have to pay for it. We have another video about pricing, especially the pay-as-you-go part. What I tell my customers is to use the SharePoint model only if they don’t have AI Builder credits in their environment. This one doesn’t use AI Builder credits, so it’s very easy to use. It’s quite affordable. If I remember correctly, it’s half a cent per page.”
“Per processing?”
“Yes, that’s correct. Half a cent per page per value extracted.”
“So if I want the amount and also the VAT, that’s two values—one cent per page?”
“Exactly. It adds up if you have thousands of pages, but it’s still a great ROI compared to manual processing. We’ve had customers using other tools to extract just a few data points and paying thousands of euros. With SharePoint, it’s very accessible and easy to do.”
“I’m very glad with the data processing capabilities. I hope you enjoyed and learned a few things.”
“I learned a lot today, Gokan. Thank you so much for the amazing presentation. I really learned a lot.”
For those of you watching, go ahead and try it out. This is the second of three ways Gokan will teach us how to extract content from documents. Make sure to check out the other videos in the series. If you want to connect with Gokan, all the links to his X, LinkedIn, BlueSky, and other social networks are in the description below. Thank you so much, and we look forward to seeing you at the next one.
