SharePoint Premium Content Extraction: Single Class and Structured Content

Learn how to leverage SharePoint Premium: Single Class and Structured Content Extraction with Microsoft MVP Gokan Ozcifci. This 30-minute tutorial shows you how to train SharePoint to automatically identify, classify and extract key information from documents.

Discover the step-by-step process of implementing single class extraction, including:

  • Setting up content centers for storing models and training files
  • Training SharePoint with good and negative examples – Creating extractors to pull targeted information
  • – Applying trained models to document libraries

This functionality works with 20+ document types including Word, PDF, and PowerPoint, processing documents in seconds. The extracted information becomes metadata in your document libraries, ready for use in Power Apps, Data Verse, and other Microsoft tools.

Watch my 90+ courses on Pluralsight (opens in a new tab)

Video Summary

  • Creating a Content Center: Start by creating a Content Center in SharePoint, which will house all your models and training files. This is essential for organizing and managing your document extraction models.
  • Training the Model: Add example files to train your model. Ensure you have a mix of good and bad examples to help the model accurately classify and extract data. The more examples you provide, the better the model will perform.
  • Building Classifiers and Extractors: Create classifiers to identify the type of document and extractors to pull specific data from the documents. Use templates or custom explanations to guide the model in recognizing patterns and extracting the correct information.
  • Applying the Model: Once trained, apply the model to your document library. This will enable the automatic extraction of data from new documents added to the library, making the process efficient and accurate.
  • Leveraging Extracted Data: The extracted data can be used across various Microsoft tools like Power Apps and Dataverse. This integration allows for further manipulation and utilization of the data within the Microsoft ecosystem.

For more information, read the transcript blog below, or watch the video above!

Video Transcript

Are you tired of manually sorting through piles of documents in SharePoint? Well, with single class document extraction in SharePoint Premium, you can unlock smarter automation that saves time and reduces errors. In this video, I brought in Microsoft MVP and Regional Director Gokan Ozcifci, who will teach you how to train SharePoint to identify, classify, and extract key information from your documents like a pro. Whether you’re streamlining processes or improving compliance, this skill is a game-changer. Stick around to deep dive into this amazing functionality. Gokan, take it away.

Hello and welcome back everyone to the ultimate SharePoint Premium Deep Dive Series, where Vlad, Drew, and I are going to teach you every single artifact that you need to know about SharePoint Premium. In today’s video, we’re going to have a deep dive into the single class model, how you can use it, what it should do, and how you can set it up. So take a seat, fasten your seat belt, and let’s go for a 30-minute drive about the single class. But before that, let me introduce myself very quickly. My name is Gokan Ozcifci, and you’ve probably seen me around conferences, probably on this YouTube channel with Vlad, where we always talk about the latest technologies, especially with SharePoint Premium. Also, this is only one single video out of the 30-plus videos we want to cover for you, so make sure to subscribe so you can see all the videos about SharePoint Premium, contract extraction, SharePoint advanced management, and so on. All of that is going to be on this YouTube channel.

Now, if you ask me the question, “Gokan, what is the most powerful feature in SharePoint Premium?” I would tell you it’s content extraction. Why? Simply because in today’s era, in today’s organizations, we all have the same issue: it’s about the data. We have a lot of data. Many could be archived, many are probably obsolete, and many are even not used anymore, but they’re all there. We don’t want to delete them, and I do get the need why we don’t need to delete them. But finding information, and extracting the content from that data can actually be very problematic. Take the example that your manager tells you, “Hey Gokan, in that SharePoint document library, I have more than 10,000 contracts, and I need you to find me the ones where we worked with Vlad Catrinescu as an example. Find me all the contracts.” Well, good luck in finding them, right? So I know where the data is, but it hasn’t been touched for ages. You don’t know what’s inside the contract or in the legal agreement. It’s actually a real pain. So that’s the reason why content extraction can help you in extracting the content from any data or almost any data and bring them as a document library and manipulate that data.

When you enable the syntax services—yes, still syntax because Microsoft didn’t change the syntax to SharePoint Premium since a while ago; it should have been done like two years ago, but it’s still called syntax, so I’m going to call it syntax to avoid any confusion—once you enable the syntax services, you can go in a document library, and from there, you can start building your models. Once you click on that button on the ribbon in your document library, Microsoft proposes multiple ways of doing your models. The one that we’re interested in today is the single class, the SharePoint way of building models, the ones that we called before the teaching method. Right now, it’s called a single class, and it also replaced the free form, which makes me think that Microsoft is heavily investing in AI Builder because the free form and the structured IT are using AI Builder capacity from the Power Platform to extract the information, while the single class is all about myself. I’m going to provide the documents, I’m going to show which data and which content I want to extract, and I’m going to explain also how we should know that this is exactly the information I want to get. With free form and structured IT from the left and the right, I just have to select, and AI Builder is going to do everything for me. But that’s for another video. For today, it’s all about the single class.

But before going to the single class, be also aware that Microsoft proposes pre-built models, which a video will follow. It’s all about consuming those models. You cannot modify them, you cannot add extra fields, you can just basically consume it and apply it to your document library. One golden rule, let me give it to you for any customer, is if there is a need to extract data, or extract content from your data, is to first start with the pre-built models because they’re cheap. It’s only like 0.001 cents per page. I’m not a licensing expert, and God knows how many different ways of licensing we have with SharePoint Premium, pay-as-you-go, AI Builder credits, some that were or SharePoint advanced management, which was $3 per user per month, which now is for free with one single Copilot license. And please don’t quote me on that, right? Multiple ways of licensing, and this is the cheapest one. If this doesn’t respond to your need, because remember I said you can only consume it unless it’s the example of an Uber invoice if you can extract the name, the price, the location, and so on and so on if it can do it for you very well. Well, if you miss a few fields, well, you cannot add more fields. It’s trained and prepared and all done by Microsoft. Therefore, you need to go to the custom models, and if the pre-built is not an option, always go for the single class first because it is still cheap, somehow cheap. The free form and structure always require the AI Builder credits, and that is something which is not cheap. So always pre-built first, single class second, and then the AI models, free form and structure, as a last resort to extract your data from your content.

The single class is very fast, and that’s one of the attention points I want to give you. So when you build your model, we’ll see that in a couple of seconds, it will extract the data in a couple of seconds, while the AI Builder ones, the semi-structured and structured, well, it can take up to 24 hours to bring that data. I’m not saying that it’s going to take 24 hours, but it can go up to 24 hours depending on the amount of data you want to extract, the amount of files you give to your model to analyze, and so on and so on. The single class is phenomenal because it goes almost in 21, 22 different document types: Word, PDF, PowerPoint, and so on and so on. All of them you can use as a source in order to get your content from the data, and it’s phenomenal to work with blocks of text. If you have a very long legal agreement, a very long contract with a bunch of text, well, you can use your model to say, “Whenever you see the word Gokan, I want you to extract me everything which comes after Gokan,” or “Hey, I know that this document is like 60 pages long, but I want you to focus only on this piece of block and extract me that kind of information from that piece of text.” So it’s very phenomenal with blocks of text.

The single class, and I see that I have a small issue with the design, but that’s okay. You can extract sentences, and words from a specific region or the whole document, and that model that extracts this can be created in your document library. So I can create that in my document library, or the owner, the SharePoint admin, can create a content center, assign myself as the owner, and I can build my models in my content center. A content center is basically a SharePoint site with a specific template, content center for syntax models, and that container, and I’m going to show you that in a couple of seconds, can be used to store all my models, can be used to store all my training files. So the question here is, how many content centers can an organization have? Well, there is no limit. I would suggest you have a minimum of one, but then the question is, what if I build 50 different models and people from HR, marketing, or junior developers come and break your models or just use them? And you know it has a certain cost, right? So from an information architecture perspective, I would say try to build content centers as much as you need per department, per region, or per business unit. It all depends on how you manage your models and the cost applied to that model. Know that every model you create can be applied to any document library and is associated with a content type. Again, from an information architecture perspective, we tell the SharePoint people to first create the content type and add your columns and fields to it. If you don’t know how to do it, it’s all okay because SharePoint Premium gives you the chance to build one from scratch. Two words that are phenomenal to understand and to literally learn by heart are the classifier because the classifier is the term that we’re going to use to determine the type of document. “Oh, this is a legal document. Oh, this is an Uber invoice,” and so on and so on. The extractor is the one that we’re going to use to pull the extracted data and put it into the document library. The whole advantage is, that as it’s a document library and the data is in my document library, I can use that data in any other tool within the Microsoft ecosystem. I’m not going to deep dive into the Power Platform side, but with Power Virtual Agents, well, you can bring that into Dataverse and build your Power App on top of that. You can use that data in any other app that you would build with SPFx. Go for SE Pages, you can apply some sensitivity and retention labels on top of it. It’s SharePoint, right? Is it in SharePoint? Well, the sky is the limit, let me call it that way.

To summarize, it can be created from the content center. The two words, and don’t forget, you always need, and I’m going to show that to you, multiple good examples, a minimum of four and up to 10-15 good invoices and one bad to show the difference. It’s the pay-as-you-go model and it uses almost all languages, English characters, and the Latin alphabet. You can apply some retention labels and sensitivity labels on top of that data.

Now, all that being said, let’s go and create our first single class. Here, when you are in the SharePoint admin center, you can ask your admin to create for you a brand new content center. They need to go to the active sites and create a new one, and here, when you enable the syntax services, you will see that a brand new option appears called the syntax content center. Just give it a name, and I’m going to call it like “Syntax” or the “SharePoint Premium Content Center” for it as an example. I’ll give it to a primary administrator and then just send it to that person so they can start building their models for it. If you want to go for HR, definitely I’m okay with HR too. If you say, “Okay, I want to have only one content center for my organization,” that’s all okay too. So you can build one with this. Once this is created, you will receive a website like this. If I go to the homepage just to show you, you will see that you have syntax-centric web parts where you can see how your models are performing in the last 30 days. Do you see this? I already had like Uber invoice. You can see some charts, and the most important ones are actually those two here above: the models as well as the training files. If you go to the models and take the example that Drew, Vlad, and myself work for the same company and we are in IT, we have our content center for IT. Well, we can create our models in the content center and each of us can see who is doing what, and we can get an exhaustive list of all the models we have and the training files. Here, we can see all the files I need to have to train my model, right? So Vlad can add more files here, Drew can add, I can add, we can create even more document libraries and play around with all of that in order to create our model.

Let’s go here to the syntax content center, to the homepage, and let me here create a new document library. I’m going to call this “Uber Invoices,” right? And I’m going to say show that navigation, that’s all good. I’m going to create that. Now I have a brand new document library, and as you can see, my models and my training files should be pretty empty. Like I have only two which have been created like last year, and I have only that kind of two kinds of invoices, one about Uber and the other one about Neo. Just here, like two kinds of data. If I come back to my Uber invoices, well, once you enable the syntax services, you will see the classify and extract button appearing on top of your ribbon. When you click on that, well, basically you can actually here within the document library create a model, right? Or apply the ones that already exist, you can also see if a document or a model has been applied to your document library. So that’s one option. From your document library, you can create a model and bring that to life, or from the content center under the model’s document library, well, you can create a model just here.

Here is where you will see the same screen as I had in my PowerPoint presentation. Remember the free form and the structure, which are the AI Builder ones? This is the one that we’re interested in, the single class. When you click on it, well, it will give you a bunch of information, examples, training details, and the ones that I already mentioned to you about the file types and the language. I’ll click on next here. I’ll have to give a name, so I’ll call it “SharePoint Premium Video Series Uber Invoices.” So this is the name of the model I’m going to create, and you can see it gives you already a visual representation of how this model can be used. If you have a block of text, and then you can scope information as a word or as a block text to extract in your model. I’m going to create that. No need to give a description and I’m going to create that just here. This is the SharePoint way of doing stuff, right? That’s what I said. And here, well, it asks you to have four actions or to respond to the four actions. The first one is to add example files, classify, create and train the extractors, and apply this to the libraries that you want to extract. Okay, let’s add a few invoices. What you will now see is by default, it’s coming to the training files. Remember in my content center, I had the training files just here. So basically, I have five Word and five PDF files. So if I come back here, this one I can close this, well, you will see I have the same files here under the training files. So I’m going to choose one, two, three, four, five good invoices, and I’m going to choose one bad invoice. Okay, well, you will see examples of the training. I’ll have them all here, so you can always add more. So if you’re not happy with the results you’re getting, and I’m going to show you how you cannot be happy, well, you can always add more good or bad invoices and also remove them, right? If you click on a file, well, you can remove that file from here, from that model too.

Now we have to create our classifier and say, “Okay, is this a good or a bad invoice?” Okay, I know that the Word document is a bad one, so I’m going to click on no, and you will see here it’s going to come with a negative. Okay, this is a good one because it’s an Uber invoice. This is a good one, this is a good one as well, and as well. So I have five positives and one negative. That’s what I really recommend as the minimum, but what you can see here is you need only five examples, so four positives and one negative are also good. But what I would suggest to you is if you have more, add more because it will help your model understand your invoicing and how it’s looking. So I’m doing only like five in total, but if you can add 10 or 15, it’s all good. The more, the better, let’s say it that way. Okay, now I have like all good and this is my bad one. Okay, so now I’m going to click here above on the train, and here is the tricky side. I have to tell my model how it should understand that it’s all about Uber invoices and not Amazon, Facebook, or any other big vendor. Therefore, you have to create an explanation. You can either go for a blank explanation or you can go for a template explanation. I’ll go for a blank, and for the second example, we’ll go for the template. I’ll go for the blank. I’ll give it a name, so we’ll know about the name because that’s obvious, right? If we want to go for Uber invoices, the best would be to look for the word Uber, so the name of the company. And here, Microsoft asks you, or the SharePoint team asks you, to describe it either through a phrase list, which could be words, phrases, patterns, a regular expression, or a proximity about, “Okay, you should search for it.” I know I’m looking for Uber, it’s a word, so I have to choose the phrase list here. Bear attention, I’m just clicking on the U of Uber, and it already recognized that I have a bunch of times the word Uber coming through, and it suggests me to use that. Well, basically, I can click on the tab and accept it. And here, under the advanced settings, I can also tell it, like, you can search for it at the beginning of the file, at the end of the file, or the custom range. So I can easily come here and say, “You know what? Please search for my word just here.” Okay, so it’s all up to you on how you want to build your model. I’m going to say anywhere in the document because it’s like a one-page invoice, and I’m going to say save and train.

Now you’ll see my model is busy testing and training. I have a match everywhere, so that’s good. And above on the right side, you see an accuracy. If it’s 100%, you did an amazing job. Congratulations. If you’re 98%, phenomenal. If you’re between 98% and 95%, I would suggest you add maybe one or two more explanations. If you’re below 95% up to 90%, more files, more explanations, more of everything. Below 90% is not good. I’m not sure that below 90% you would get all the results you want to. So 100% is great, 98% is phenomenal, below 98% to 95%, I would add more data and more explanation. But hey, in my case, I can see it’s all good, so I’m clicking on test and exit the training. Okay, so we’ve added our files, we said in our model, “Hey, this is how you can know that it’s an Uber invoice by searching for the word Uber,” and now we have to create the extractors. Remember what the extractors are? It’s the value you want to get from that specific invoice. The most obvious one would be the price, right? That’s the price I want to know, how much it cost me to travel from point A to point B. Under the advanced settings, I can create a brand new column which will be called “price,” or remember what I told you at the early stages of this video, from an information perspective, if you create your content type first, you should have created your column first here too, and then you can just choose your column and then fill in that column. But here, we didn’t do that. I’m a bad SharePoint person, I would say. I didn’t create my content type, so well, let’s create a brand new column, call it “price,” and then click on create. Now we have a brand new column, pretty empty. We should tell our model with what it should fill that brand new column.

Here, it’s a bad, a negative one, so this doesn’t have any label. I’m going to click on save and then go for the next one. Okay, here, this is my value. I want to extract the 31.11 and click on next. Here again, this is 36. I want to extract this. 37 is, I want to extract that. 37 is, I want to extract this. And finally, I have something like 43.12. I want to extract this. So basically, what I’m telling my model is I have a brand new column which is called “price,” and you have to fill in that column with all those brand new values. Okay, pretty easy, right? But that’s not all. We have to now go on the train and explain how it could find that value because now it’s easy because we just highlighted it, but what if my invoice is now 10 pages long or a half page or it’s in a different spot? Well, SharePoint or the single class model can never know that, so we have to explain how it could get it. So you could again go for blank, but now let’s go for a template. Microsoft proposes you a bunch of templates that you could use in order to scope your information. You could go for a credit card, currency, date, or a bunch of different explanations. The ones that I really like are the after label and the before label. Remember, we always have the same pattern: 31.11, we have a Euro sign or a dollar sign if you’re in the US in front of my value. This is a recurring pattern, so we can use that, and we can tell my model, “Everything before, well, before the value I selected, there is something,” and you can see that the AI behind the scenes already detected the pattern and says, “Oh, Gokan, in the values you’ve been selecting, I always can see a Euro sign.” So I can just leave the Euro sign, right? And then just click on save. Is that good? Yes, that could be enough, but I want to make sure. I also want to use the after label. Well, you can see after the label, that it detects also a common pattern, which is Gokan and a temporary one. So I’m going to add this. Well, you can see, yes, after the selected value, I have a Gokan. Oh, that’s a correct pattern, so I can leave that too as well. Then you can add as many explanations as you want in order to achieve 100% accuracy. I’m going to just leave it like that, and I’m going to just double-check. So before the label is the Euro sign after the label is Gokan. I’m going to hear Euro sign, Gokan, all good. It looks all great, Euro sign, Gokan. So my explanation should work. I’m going to train this again. It will take the before the label and the after label, match it with the label and try to get the predicted one. So now I have all the matches, all good. If I close this, I have like 100% accuracy score for my model, and that’s it basically. I’m going to test and quit right here.

You have the entity extractors. Remember, I only created the price. Well, from here, you can create more values you want to extract. It could be the Uber driver’s name, the street from, the street to, the date, and as much data as you want to capture from your content. Well, you can create extractors from here. One thing also to know is when you select an extractor, you can click on the refine extracted info, and SharePoint Premium or the syntax services give you away. Well, normally within the document library or the SharePoint list, we have a few of them which are free and come with the base license, but Premium gives you a default value. Like if you extract something and there is nothing, or you want to bring a default value, well, you can do that. Or you want to remove duplicate values, well, you can use this operation to bring that info up to date. So that also is quite handy and nice to have.

Now, the last option would be to apply this to our document library. Well, you can do it from here, or you can come back to that specific document library just to show you that you can do it in both ways. So I’m going to just refresh it and then say classify and extract. And here, do I see it? Uber invoices, create October, October, no applied, still nothing. Okay, I’m going to come here to my model and from here, I’m going to the syntax content center, going to the Uber invoices, and then apply this from the content center. And now you will see that it has been applied here below to the Uber invoices. Coming here again to the Uber invoices, doing a quiet refresh, click on classify and extract. Oh, now we see the application. Remember, it always took me to the available and applied was empty. Now you will see that this has been applied. How can you know this? Well, you will see if you go under the views, well, you have a brand new view which is called “S Video Series Uber Invoices,” and if you click on it, well, you’ll see the brand new column that we just added called “price,” the processed, the pending status, the modified, and so on and so on. So now if I upload a bunch of files, and let’s go for all my Uber invoices, I’m going to upload them here. What you will now see is the following: we are analyzing your files. You’ll see the extracted info in the columns of the library in a matter of seconds. And that’s what I’ve been telling you since early in this video. And now I’ll have to keep on clicking refresh while I have to find something to tell you to not have a big blank, right? So I’m going to click again, and hopefully within a few seconds, hopefully less than a minute, the value should be extracted. Okay, what I will now do is I will just come here and then under the model settings, talk to you about this, and then we’ll go back and see if the values have been extracted. You can always modify and add the description. You can add this to no site, so you can restrict it. You can say, “Okay, you know what? I don’t want to use that model because it has some issues or we don’t use it anymore,” or you can just select your site, or you can apply this to all sites. Be careful, don’t apply this to all sites. You could do that, right? If it’s something regular, you’re in the EU, you have to be compliant with something, you can do that. Just know that the option exists for doing that. I always go for the selected sites. And then for the compliance, right? Remember what I told you, you can always apply the sensitivity label. Those are my sensitivity labels as well as the retention labels as never-to-be-deleted labels to my documents. You can delete that model from here, and basically, that’s the only option you receive from here to do that.

Now let’s go to my model and click on refresh. And there you go, in a couple of seconds, well, less than a minute, let’s call it less than a minute, it finished. Well, you can see the processing time, it finished about a minute ago, and you can see that it extracted perfectly the values I needed from my Uber invoice. So if I click on 002, it should be 31.11, and that’s the same as what you get. Now, one thing that I always show as well when we do the single class is when you open your PDF, well, you can ask the same question to your Copilot, and it will generate the same, same answer, like, “What was my final price?” Right? So you can do that with Copilot within your PDF, and it will check into your document, and it will find it, right? You can see it, it’s 31.11, phenomenal, very quickly, in a couple of seconds. But it will not, not that I’m aware of, bring this into your document library as a metadata, as a column field, and you know what? You can now play around with that data. That data is now here, and you can extract that data, bring it into Power Apps, into Dataverse, into anything you want, and apply that to you. So that’s basically what you can do in a few minutes with the single class model with SharePoint Premium. Again, very quickly: create your content center, add the training files into this tab, go to models, create your first model, make sure how the model can understand what kind of data this is, and make sure to extract every information you need by bringing a lot of explanations, and finally, apply it to your model. You’ll see in a couple of seconds or minutes, the data is going to be extracted from the content to your document library.

Well, that’s me with the single class model. Don’t forget to subscribe and check out upcoming videos because we know that Drew, Vlad, and myself, are going to rock and roll together to bring all the awesome artifacts to you on this channel. Take care and see you very soon.

Thank you so much, Gokan, for this amazing deep dive. For all of you watching, I hope you enjoyed this deep dive video, part of the ultimate SharePoint Premium Deep Dive Series. Make sure you check out the playlist appearing on your screen right now to see all the released videos in the playlist. Subscribe to the channel to get notified as soon as the other videos come out. I can’t wait to see you in the next one.