<img height="1" width="1" style="display:none;" alt="" src="https://px.ads.linkedin.com/collect/?pid=4011258&amp;fmt=gif">

MICROSOFT FABRIC · AGENTIC DATA MANAGEMENT

What agentic data operations actually look like in a Microsoft data estate

Komatsu’s requirements offer a useful way into the subject: the work agents take on, the Microsoft services around them, and the decisions people still own.

Komatsu’s brief for AI agents included a very specific request: understand the company’s abbreviations and data dictionary. The agents also needed to work with Microsoft Fabric and reach people through Microsoft Teams or Jira.

Those requirements sat alongside a need to process 10 million records a day in the project we presented at FabCon. Volume was only part of the brief. Komatsu had multiple ERP systems, limited hands-on resources for the project, and a clear expectation that business users would take responsibility for their data.

An agent has to fit into that working environment. A technically plausible suggestion is of little use when it misreads a local abbreviation or arrives without enough context for someone to approve it.

Agentic data operations give agents defined, recurring data-management jobs, with access boundaries and an agreed route for putting their recommendations into effect. In CluedIn, that work sits within master data management: the ongoing care of the customer, product, supplier and other business records that downstream systems depend on.

Give the agent a proper brief

Consider an illustrative product record arriving from an ERP system. Its category is missing, its description contains a local abbreviation, and a similar record already exists elsewhere. Completing the category and deciding whether the records describe the same product are different jobs. They may need different evidence and different approvals.

A useful instruction would spell out those limits:

CluedIn’s job configuration makes the scope explicit: select the records and fields, choose the skill, provide instructions, then test the job on a sample. A job can also be configured to consider only records modified since its last successful run.

Our FabCon demonstration covered data mapping, validation, enrichment, classification and duplicate-resolution work. It also showed a conversational request for a client-validation rule, including checks on email and phone fields. These are recognisable stewardship tasks, with outputs that a business user can inspect.

CluedIn’s graph-native foundation holds entities and their relationships. Where the relevant relationships and business definitions have been modelled and supplied to a job, they provide context for interpreting a record. The agent still needs permission to read that data. A knowledge graph is not an invitation to browse the entire estate.

With an appropriate run schedule, that defined job can become part of the recurring operation. Someone still owns its instructions, reviews its performance and adjusts it when the business changes.

Where Fabric and Purview fit

In the Komatsu solution described at FabCon, responsibilities were shared across the stack. Fabric handled data movement and integration, Delta tables, analytics and reporting. CluedIn covered entity-level quality, matching and merging, enrichment, a business semantic layer and agent-assisted data management. Purview provided the governance and lineage context.

For teams designing a similar environment, the wider platform capabilities matter too:

Microsoft Fabric
Microsoft describes Fabric as an analytics platform spanning ingestion, transformation, data engineering and reporting, with OneLake as its shared data lake. The platform already includes AI assistance and governance capabilities.

Microsoft Purview
Its integration with Fabric supports data discovery, lineage, information protection and auditing. Microsoft also documents data-quality checks for Fabric data.

CluedIn
The specialist agentic master data management platform combines entity resolution, data quality, enrichment and governed agent workflows. It maintains the business records whose consistency matters across source systems, analytics and AI.

The connections are concrete. CluedIn documents an Open Mirroring connector for publishing data into Fabric. Its Purview integration can also synchronise deduplication projects as assets, exposing their place in the processing lineage. The integration path depends on the customer’s implementation.

Microsoft Fabric, Purview and CluedIn: responsibilities in the solution

Fabric, Purview and CluedIn: responsibilities in the solution

A simplified view of the roles described in the FabCon session. Integration paths depend on the implementation.

What happens before a record changes

CluedIn’s production guardrails start with read-only access to master data, controlled creation of proposed outputs, access control, attribution and human review. There is a reason to be this specific: generating a useful recommendation and being authorised to apply it are separate responsibilities.

CluedIn’s agent documentation describes agents as read-only. They inspect the data made available to them and prepare recommendations; the platform handles execution through the configured approval process.

For proposed data-quality fixes and rule changes, human review is the default. An administrator can enable automatic approval of suggestions, but the agent still does not acquire unrestricted write access or the ability to approve its own work. Decide which workflows are suitable before enabling that setting. The review documentation explains the approval behaviour in more detail.

Audit trails and data-quality metrics need to be deterministic. A model’s explanation is useful to a reviewer; the record of what changed, who authorised it and which job produced it needs to be captured by the platform.

There must be a way back, too. CluedIn documents how to review and revert agent-job changes to golden records, either for a job or an individual record. Any downstream consequences still need to be considered in the design. Reverting a master-data change should never be casually described as undoing every action another application has already taken.

Agentic Data Management - Inside an agent job

Inside an agent job

An agent job in CluedIn, with its configuration and review surfaces. Product illustration from the FabCon presentation, not a view of Komatsu’s data.

The approval has to reach somebody

Komatsu’s request to use Teams and Jira deserves as much attention as the processing requirement. Business ownership of data becomes difficult when every decision requires someone to visit an unfamiliar tool and reconstruct what happened.

CluedIn’s Power Automate integration can turn platform events into workflows and send approval requests or notifications to Outlook or the Approvals app in Teams. The FabCon session illustrated the kinds of work involved, including suggested quality fixes, possible duplicates and proposed validation-rule changes.

A useful review request should identify the record, explain the proposed action and give the reviewer enough evidence to decide. Routing matters as well: the colleague who understands a product classification may be quite different from the person who owns customer matching.

That is part of the implementation work. A workflow that sends every uncertain case to the same overloaded person has simply created a faster queue.

The result worth looking at

The closing update in our Komatsu session reported that the operation had expanded to bring in all its product data. It also described a substantial change in the effort needed to maintain the data-management engine.

For another organisation, the useful test is whether the operation requires less effort while producing data people can rely on. Measure time spent reviewing exceptions, the quality of accepted changes, rejected recommendations and the cost of running the jobs. Check whether corrected issues stay corrected.

Context-window limits still apply. A requirement to process millions of records does not mean placing millions of records in a single model prompt. Scope, processing design and testing remain engineering responsibilities.

CluedIn Agentic Data Management: Monitoring data-quality issues by business domain.

Keeping an eye on data quality

Monitoring data-quality issues by business domain. Illustrative CluedIn interface; the displayed figures are not Komatsu performance results.

Start with the work people keep having to repeat

For a team already investing in Fabric, a sensible first candidate is a recurring job with a clear owner and a result that can be checked. Missing product attributes, inconsistent classifications or a queue of possible duplicates give you something concrete to test.

Define the records in scope, supply the business terminology, and agree who can approve which changes. Run the job on a sample before widening it. During the trial, keep a record of what the reviewer accepted, what they rejected and how much work the process actually saved.

Pay particular attention to the awkward records. Ask the person who normally handles them to sit alongside the first review. Their reasons for rejecting a suggestion will tell you where the job needs tighter instructions, better context or a decision that should remain human.

Bring those records to the demo. They will tell you more than a clean sample ever could.