Understanding Content Processing with SharePoint Premium

Get the content processing foundation you’ll need to get the most out of SharePoint Premium with the help of 3 Microsoft MVPs!

Watch my 90+ courses on Pluralsight (opens in a new tab)

Video Summary

  • Introduction to Content Processing in SharePoint Premium: Automates document classification, extraction, and management using advanced AI, making data handling more efficient.
  • AI-Driven Features: This includes tools like automatic metadata extraction, image processing, and OCR, which enhance document management and data discoverability.
  • Content Centers: Serve as central hubs for managing AI models and content processing tools, ensuring efficient distribution and cost management.
  • Licensing and Costs: Consumption-based pricing means you only pay for what you use, with minimal costs compared to the benefits, such as reduced manual labor and improved efficiency.
  • Practical Applications and ROI: Real-world examples show significant cost savings and efficiency gains using features like taxonomy tagging and pre-built AI models for tasks such as invoice processing and sensitive information detection.

For more information, read the transcript blog below, or watch the video above!

Video Transcript

Managing content in SharePoint just got a whole lot smarter. The content processing pillar in SharePoint Premium takes your organization to the next level by automating the way you classify, extract, and manage your documents. From advanced AI-driven tagging to seamless data extraction, it’s all about turning chaos into clarity. For this video, I brought in Microsoft MVPs Gokan and Drew to introduce this pillar that helps you save time, boost compliance, and stay ahead in the digital era. Gokan, Drew, take it away.

Hello, welcome everyone. This is part three of the Ultimate SharePoint Premium Deep Dive series, and we’ll be talking about the processing introduction. It’s not the first time I’m on the channel talking about this subject, and I’m with Drew, who is going to bring a new way of thinking and doing stuff with SharePoint Premium. I’m extremely excited to be with him, and we’ll introduce you to the content processing experience.

Now, when we look at the content processing experience in SharePoint Premium, we’ve defined multiple definitions around it. This is a phenomenal overview. Maybe we can start with that, Drew. What do you think about the bullet points we have on screen? Content processing as part of SharePoint Premium has seen the biggest change and growth around the AI conclusions part. We always could work with metadata and process documents as part of, I’d say, next-gen content services. Content processing brings us next-gen content services for AI-driven management and automatic metadata extraction.

When working with information, we want to find relevant data. We now have AI tools as part of SharePoint Premium content processing to extract that information and make automation better. What do we want to do with that information after the fact? It also includes much more advanced features around image processing autofill scenarios, and OCR, which we’ll talk about in future videos as we go deeper into each of these functions of content processing.

Absolutely. The whole idea is that you get all the information from your content and can actually use that data and bring that intelligence somewhere else. You’ve already mentioned it, right? We have a bunch of tools in our SharePoint that we can use to bring that intelligence, such as data from content extraction, content assembly, image tagging, taxonomy tagging, and so on. There are a lot of cool videos coming that we will discuss together. What’s the one that’s exciting you the most here? Mine is autofill. I love the fact that you can just write a prompt and get information from your content. What’s the thing that you say is a must-have for organizations when you see that list?

I’ve been working with content extraction since its early days in SharePoint Syntax and prior. That’s the one I’ve worked with customers the most and seen true value in. I think a big advantage of this series is that we’ll have individual videos for each of these features that will break them down in detail. Content extraction will have a lot because there’s so much there. This video series will be great to come back to and reference as far as what feature is most important to you at a specific point in time. For me, it’s been content extraction. It’ll be interesting as we go down this full series of autofill and how they relate to each other.

Absolutely. It’s funny you mentioned Syntax. Just as a reference for those who are new to the channel and have never worked with SharePoint Premium, Syntax is the name of SharePoint Premium. Before that, it was called Project Cortex, then it went to SharePoint Syntax, then Microsoft Syntax, and now it’s called SharePoint Premium. As we mentioned earlier in our videos, a lot of people still refer to it as Syntax, but that’s not a bad thing. It just shows that there is a need today to create those videos to educate people on the naming and the proper usage of the names.

Absolutely. As you start breaking names down further, it’s good to look at the feature set. Why is content processing important? What are the actual benefits outside of just these tools? We’ve talked about content extraction, but what content extraction provides is the ability to take a document a PDF, or multiple types of documents, upload them into SharePoint, extract information and metadata directly from that document, and then use that information to build certain things, whether it’s workflows or automation. We can now build brand new documents using metadata that we’ve extracted using content assembly, all as part of the content extraction family.

I know customers today extract that data and use that metadata in Power Apps. They have SharePoint as a data source, and they can grab that information and use it in their apps, which is a barrier-breaking way of thinking. Now all the data is easily accessible, and that data, as it resides in SharePoint, can go anywhere from Power Platform to search-driven pages. We can even apply some policies on top of it, which I think is the major benefit of content extraction.

Once you’ve extracted that, there are even more advanced benefits you get as part of content processing around new enhancements that you can take as part of the data. You can take that data and manipulate it as part of the content services lifecycle to add your information to it. You can start annotating and merging these documents and even do live translations of this information to give us a multilingual and multi-experience, which we did not have in the past.

Once we’ve done that, we’ve now extracted information, built new information, and modified that information. A lot of what content processing brings is the ability to find and discover it after the fact. As we introduce new AI scenarios such as Copilot, having more metadata allows better discoverability. That’s where we can start to bring in automatic tagging through managed metadata and autofill to give us a much cleaner experience in finding the content we’ve been looking for as part of processing.

Absolutely. We talk a lot about content extraction, and we have multiple ways of doing it. We have the teaching method, which we also refer to as unstructured content processing, which we’ll deep dive into in upcoming videos. We also have semi-structured and structured content processing, also referred to as freeform and layout. Finally, we have the pre-built models. If I had to explain this to someone who has never used SharePoint Premium before, I would say the teaching method is the most complex because you have to define your own patterns and bring your intelligence to your model since you are the one creating that AI model.

For semi-structured and structured content, we use AI Builder or the intelligence of AI Builder to get that data out. The pre-built models are the ones that Microsoft builds for us. You just apply that model to a document library, click next a couple of times, and hopefully, the pre-built model can then get the data you want from the documents.

When we build those models, we have multiple ways of doing that. We can either do this in a document library or build it in a Content Center. I don’t really talk about Content Centers in my sessions or even when we talked with Vlad on his channel a while ago, but I realized in the last six months that Content Centers are phenomenal in usage. It’s very important. Maybe you can explain why that is important and why everyone should adopt Content Centers and not build directly into a document library.

As part of content processing, when you think about it at a higher level, it’s about understanding what’s happening in our organization. We want to empower end users to do this, but they might not know how to work with some of the models that Gokan has talked about. A Content Center gives us a central source where we can store these models, store our next-gen content services scenarios, and distribute those back out to our enterprise. This way, we can manage the cost as we go through and have visibility into the usage of these modern tools and AI models within our organization. It gives us a central source to manage and view this information, which you don’t have in the sprawling world of content today. We talk about content sprawl a lot, and this is almost like model sprawl that we want to keep as part of our Content Center today.

One recommendation I give to my customers and clients is whenever you want to extract data from your content, always go pre-built first because it’s the easiest and cheapest option within the ecosystem. Secondly, go for the unstructured method. It may be complex, but it doesn’t require AI Builder credits like the semi-structured or structured content does. Do you agree with that? I’m always saying this, but I just want to hear it from you if this is the way to go.

Absolutely. We’ll go into detail in each of the videos, but at a high level, 100%. The pre-built model gives you a starting point where people who are nervous about building new models for organizing their content and extracting content can start. You can always step up from there if the content extraction you’re trying to pull out is not adequate or doesn’t have the right confidence level.

Fantastic. I’m so happy that you agree with me. That means I was right when doing that. To wrap up, we also have a few new slides about the differences between unstructured and structured modeling. One is created in SharePoint, the SharePoint way of doing things, and the structured model is the AI Builder way of doing things. One requires you to use PDF files, Office files, four positive and one negative, to build your model. The structured model requires up to 50 megabytes or up to 50,000 pages to build that model. Each has its advantages and drawbacks in building the models. One important thing to know is about the licensing. The unstructured model is the pay-as-you-go model, consumption-based. However, the structured model requires AI Builder credits. For example, 3,500 credits are included in each license per month, and 1 million credits allow you to process up to 2,000 file pages, which I think is a bit short. What do you think? Should people go to the AI Builder? What are the use cases for AI Builder? Maybe we can spend a few minutes discussing that.

What I’ve seen the most in this scenario is that unstructured information is really focused on document types that are truly unstructured and have very specific file types, often office-specific. Language restrictions can also be a significant factor. Unstructured processing does take more effort because you have to define classifiers and extractors to figure out how the content works. Structured processing may cost more, but it can scale at a level that unstructured processing can’t due to the types of files it works with. For example, if you’re working with very organized PDF files that have boxes around different levels of information, structured processing gives you a higher advantage, even though it has an extended cost with AI Builder included.

Absolutely, I agree with that. So, that’s basically the difference between unstructured and structured processing. If you don’t get it now, it’s okay because we will deep dive into multiple videos to explain everything in detail. The next slide shows that even Microsoft is working to provide pre-built models. They have models for contract processing, invoicing, and receipts, and the latest one for sensitive information processing. The only question I receive from customers and clients is whether Microsoft uses our data to build those models. Absolutely not. Microsoft does not use our information as part of this. These pre-built models are built by Microsoft using non-customer information that is most commonly used across the globe. The sensitive information type video will be very informative about the new functionality there. Microsoft does not train on our information, so please don’t be scared to use and try those models. Microsoft does not charge you if you test those models.

You can build an AI model to extract information, and if you don’t put it into production, you won’t be charged. It’s a fantastic opportunity for everyone who wants to extract data from content to use those models. Every organization needs that. If any organization tells me they don’t need it, it’s a big lie. Everyone needs to extract data from invoices, such as the amount, invoice number, or payment deadline. There was a big discussion about whether financial data should be in SharePoint. We can discuss this a lot, but if it’s in SharePoint, you can use these models to extract data and go far beyond, building your own applications, applying policies, and even checking if your data has sensitive information. For example, you can check if a credit card number is stored in your data and then remove that information from the document. The sky is the limit when it comes to imagination.

That’s what I want to say about the models, but that’s not all. When we talk about extracting content, we can also create content with SharePoint Premium. One feature that I really like is modern document completion in SharePoint Premium. Drew, maybe you can say a few words about that.

Absolutely. As part of content processing, we can extract information, but a big advantage of SharePoint Premium is taking that metadata and actually building content or documents using the metadata inside SharePoint. You don’t have to think about working manually. It’s about taking the information we have, using that knowledge, and building content automatically. This brings metadata back to life. We’ll have a much deeper video about content assembly, but think about content processing, it goes from content creation to content extraction and back to new content creation.

What I really like is not only the metadata but also the existing list items and metadata. You can combine that and create your document. Before we get into the details for the subsequent videos, we want to make sure we level-set the stage regarding how the licensing and cost work for this, as this is one of the most challenging aspects customers want to know about. How does the licensing work for content processing in SharePoint Premium?

It’s basically consumption-based, meaning you will only be charged for what you use. When we say that to customers who are not really into the AI era or the cloud era or have never used any consumption-based resources, they get scared. They worry about using a million pages or multiple times and the costs involved. I get the point; it’s scary when you don’t have the correct numbers. If you look at the screen, the numbers are very small. For example, the document processing for pre-built models is 0.01 cents per page, which is like a dollar for 100 pages. I think anyone around the globe working with M365 can afford that kind of transaction modeling. The only thing I think is a bit expensive is e-signature, but we will discuss this in another video in this series.

If you’re wondering if content processing in SharePoint Premium is right for you and you’re worried about the cost, what we’re trying to say is that there are cost implications, but we have not worked with any customers where cost has been a complete stopping point. When you look at this inside your company, consider the actual benefit of man-hours it would take to do this level of extraction manually. I’ve worked with organizations where people manually pull information out of content. You can always find ways to achieve ROI as part of content processing that you don’t have in your organization today. The cost here is usually minimal when you look at the benefits.

To give you a real use case, an organization digitally transformed by creating its SharePoint site and uploading all its data to SharePoint. They hired two people to manage their metadata full-time, adding the correct metadata and data into the columns. I showed them that taxonomy tagging could do that automatically. I’m not saying the answer was 100% always correct, but it was easily 99% accurate. They saved a lot of money by using that AI service, which we call taxonomy tagging. When you upload a file, it will automatically check the content and apply the metadata to the appropriate column. This was phenomenal, and I was so happy that SharePoint Premium had that feature, allowing them to save a lot of money.

That sounds like a very exciting feature, and I know we have a lot of videos to talk about as part of this series. Thank you for listening to this conversation. We’ll have a full collection of videos about all the features we talked about around content processing. Make sure you check out the rest of the videos in this series, which you’ll be able to see in the description below, along with all the other features we talked about as part of the Ultimate SharePoint Premium Deep Dive Series. This was the introduction. My name is Gokan, not Vlad. I try to be as fun and good as Vlad, and he tries to be the Gokan. So, I’m Gokan, and this is Drew. Stay tuned for the upcoming videos.

Thank you so much, Gokan and Drew, for this amazing deep dive. For all of you watching, I hope you enjoyed this content-processing introduction video, part of the Ultimate SharePoint Premium Deep Dive Series. Make sure you check out the playlist appearing on your screen right now to see all the released videos in the playlist. Subscribe to the channel to get notified as soon as the other videos come out. I can’t wait to see you in the next one.