← Back to Podcast/Microsoft Fabric End-to-End: From Raw Data to Business Decisions with Amit Chandak [MVP]
Episode Transcript

Microsoft Fabric End-to-End: From Raw Data to Business Decisions with Amit Chandak [MVP]

Microsoft Fabric brings data engineering, analytics, business intelligence, governance and increasingly AI together in one platform. But what does an end-to-end Fabric architecture actually look like when you move beyond individual features and start connecting everything?In this episode of the M365 FM Podcast, Mirko Peters is joined by Amit Chandak [Microsoft Data Platform MVP] for a practical journey through Microsoft Fabric — starting with raw organizational data and ending with trusted information that business users can use to make decisions.

WHY MICROSOFT FABRIC?
Before Fabric, organizations could already build sophisticated analytics architectures using Azure, Power BI and other platforms. The problem wasn't a lack of technology. In many cases, it was the opposite: organizations had too many choices, separate storage technologies, different compute models and multiple copies of essentially the same data.Amit explains how Microsoft Fabric attempts to simplify this architecture by bringing workloads together around shared foundations such as OneLake, common Fabric capacity and the Delta format. Lakehouses, warehouses, Power BI and other Fabric experiences can therefore operate as parts of a broader platform instead of completely isolated services.

ONELAKE AS THE FOUNDATION
OneLake is one of the central concepts behind Fabric. Amit compares it conceptually to OneDrive: instead of every analytics workload creating completely independent storage environments, OneLake provides a virtualized storage foundation across the Fabric tenant.Organizations can still separate data through workspaces, Lakehouses, Warehouses and domains, but those resources exist within a common Fabric storage architecture. This becomes particularly important when organizations want to reduce unnecessary duplication while maintaining security and organizational boundaries.

CENTRALIZED DATA OR DATA MESH?
Fabric doesn't automatically mean putting everything into one giant centralized analytics environment.For smaller organizations, a centralized architecture may still work well. As organizations become larger, Amit sees increasing value in domain-oriented architectures where areas such as sales, finance and purchasing can have their own workspaces and responsibilities.IT can remain responsible for availability, governance and the technical foundation while business domains increasingly take ownership of how their data is analyzed and consumed.

SHORTCUTS INSTEAD OF COPYING DATA
One of the recurring themes throughout the conversation is avoiding unnecessary copies of data.Fabric Shortcuts allow teams to reference data stored elsewhere rather than physically copying it into every environment that needs it. That can apply both inside Fabric and to supported external storage.Amit also explains an interesting architectural benefit of shortcuts: they can help separate workloads across capacities. This can become important when organizations want Power BI consumption workloads isolated from intensive data engineering workloads while still working with the same underlying information.

LAKEHOUSE VS. WAREHOUSE
One of the biggest Fabric architecture questions remains: Should you use a Lakehouse or a Warehouse?A Lakehouse can work with structured and unstructured data and is naturally aligned with Spark. A Fabric Warehouse focuses on structured data and provides the familiar T-SQL experience.Both ultimately use Delta for structured data inside Fabric, which means the decision increasingly comes down to the type of data, preferred technologies and workloads.Organizations with strong SQL teams don't necessarily need to abandon their existing skills. Teams working with very large datasets, advanced engineering scenarios, unstructured information or extensive data science workloads may find the Lakehouse and Spark approach more attractive.

GETTING DATA INTO FABRIC
Once the archi

Welcome everybody to the new episode of the MC65FM podcast.

Today we are going to end to end with Microsoft Fabric.

Not looking at one isolated feature,

but at how the different parts of the platform

come together to create a modern analytics architecture.

How should organization think about leg house versus data warehouse,

where do spiced back belong?

When should we use data flow gen 2?

How do the mental models fit into the architecture?

And ultimately, how do we turn all this technology

into power BI solutions that business user can actually use

and make a better decision?

My guest today is Amit Shaktaq.

Microsoft data, Microsoft data, platform MVP and power BI community

super user, Amit has more than 22 years of experience

in data analytics and currently leads data engineering

and the leg solution at Canerica is experience response

Microsoft fabric power BI SQL databases, Tableau,

in quarter data engineering modeling and visualization.

Today we are going to build the Microsoft fabric analytics

story from the ground up from architecture

to ingesting through transformations, storage,

semantics models, security performance, and final power BI.

Amit, welcome to the MC65 podcast.

Thanks, thanks for inviting me for this podcast.

It's my pleasure to be part of this one, Peter.

Awesome.

Before we get deeply technical,

how did your journey into data and analytics begin?

Yeah, so I joined Oracle in 2003.

I selected Oracle as an act of campus and I became the part

of their BI team.

They were going to one of the transformation at that time.

We were building this tool which is known as the

BI daily business intelligence and that's where I started my journey.

After I left Oracle, I started a company along with one of my colleagues

and I ran a data analytics startup for 10 years.

We created our own tool very similar to Power BI directory.

We used to have the tool completely on the web.

At that time, you know, having a complete web

all thing was a challenge.

So that tool was completely authored on web, designing reports,

creating the semantic models and everything was on the web.

So complete web all thing was there.

And then I done that for 10 years and post that I decided to move on.

That's where I joined my current company, Connecticut.

I also moved to Microsoft technology around the same time.

I became a Power BI community super user around 2020 and then around 2022,

October when I became Microsoft MVP for data platform.

Awesome.

That's, yeah, that's the way it went.

22 years, it's a lot of time.

But yeah, that's direct jump into the fabric part.

So let's start with the big question.

What problem is Microsoft fabric actual trying to dissolve?

See prior to Microsoft Fabric, see, it's not that we are doing data

engineering something new with fabric.

The whatever components fabric has, all those components

were previously available at different places.

So if you look at Azure ecosystem, you could have, you know,

100 plus choices to do what fabric is doing today.

If I look at the larger ecosystem, if I include Azure,

AWS, even Google platform, I might have 1000 plus choices to do my own data engineering work.

But the challenge was bigger.

Now choice of technology was one challenge.

Then other than choice of technology, create a multiple copies of data.

What happens in the data analytics or BIA ecosystem?

We say single source of proof.

We always talk about it.

But our own ecosystem, because we sometime we need data lake,

another time we need data warehouse and then we need powerful tools like power

BIA, which for input would be our self creating copies of data.

And all these tools are working on different technologies is they have different kind of back ends.

Their storage were different.

They they were created at different time frames.

So all of them were, you know, built differently, storing differently.

And they were using different ways compute powers getting compute.

A Microsoft node challenge and I think that's where they start in the center apps where

they wanted to address this one.

A fabric name was even higher where they sorted out the entire stuff.

Whether you have the structure data and structure data, where you wanted to have warehouse,

where you wanted to have power BIA or even real time analytics, they were able to put everything

at one place.

And they starting, you know, with a few foundations like one lake where, you know, your entire

data assets, whether it's a metadata data can decide.

Then they solved the problem of compute by saying, okay, all the workloads in Microsoft

fabric can share one single compute.

Now you can have multiple computers per requirement, but that that is one of the biggest achievement

because think about your running spark workload.

How would you compare that with power BIA or CQVV?

They're all together.

But everything gets converted to compute units and getting chard within the same number

of computers which is available with you.

Let's say paradew per 30 seconds.

Then comes the challenge which is there with the storage.

Now I have storage, which is available for the greater work, but what happens when I

store the data, let's say for lake house, I might prefer some format.

Now where else doesn't understand that format and that's why I was copying data from lake house

to where else when I was doing previous term.

Similarly, power BIA was not comfortable with data warehouse formats.

So it was creating on copy.

Now whenever we store the structured data, whether it's a lake house, whether it's a warehouse,

it all getting stored into delta format, which created that uniform format, which lake

house understand, which spark understand, which T-SQL understand, which even the power

query engine has been modified to understand that.

So it means I am not unnecessary creating copies of data.

Definitely most of us follows the medallion architecture and in that we will have, you know,

bronze, silver and gold, but it doesn't mean that I'm going to create unnecessary copies

of my bronze data and unnecessary copies of my gold data.

So these are some of the challenges which existed while, you know, we have the ecosystems

available for last 15 to 20 years to do data analytics.

But these are some of the challenges which emerges because of the growth of the data, which

is there with us and Microsoft fabric has, you know, went ahead and addressed that.

And how much to do an organization's data practice need to be before a depth thing fabric?

See Microsoft fabric is basically complete end to end solution for your data analytics need.

So even if you are starting for data analytics, Microsoft fabric can be there.

So the only in the chart is needed at the source level that you understand your sources

and you understand what is needed for data analytics.

If you are ready to analyze your data, if your data is arranged, let's say, if your data

is completely unorganized, you are, you know, still working on some of those spreadsheets

a lot, some of the, let's say ERP or CRM implementations are still ongoing or on middle of it, then

it might not be the right time.

But if you are a data mature company in terms of capturing the data and wanted to do

entities or already there with some analytics and wanted to transform it to the model system,

I think in both case you are ready for data and it is on Microsoft fabric.

It provides you everything these days.

If you, if you say I just wanted to do simple, simple, create very simple report using

Powerway, that's also possible.

And you want to follow the complete architecture, you know, bringing data to the lake houses,

transforming it and create warehouse and then wanted to do it.

That is also possible in Microsoft fabric.

What, what did you think, how important is domain oriented architecture in fabric or

or short organizations create one large centralized fabric environment or distribute responsibility

across business domains?

I think we are in an era where we talk about business owning these stuff where we talk about

data mesh architecture where we talk about domains.

So in the modern architecture, why, you know, your IT could be custodian of your data,

but you need to create, you know, multiple domains in which the data is going to decide

finally.

So you can have a master data separately analyzed for the organization level because master

data management is key for all the analysis post that I think it is based on let's say

sales, finance, and purchase can have their own workspaces can they have their own domain

and the data can decide.

If you are a small to medium size organization having a centralized data warehouse, a works

means it can have anything, but if you are a medium to large enterprise or a very large

enterprise, it is always good to have, you know, data mesh architecture where you have

the domains and, you know, the ownership lies with business more than IT.

So it is the combined governance, while the data, entity and the data availability ensure

by IT, but the analysis part of it, distribution part of it and the management part of this is

governed by the business.

And I think the core term when we talk about Microsoft Fabric is it's one leg, what is

one leg in practical terms?

Like today what happens?

Let's talk about our own stuff when we are trying to go online.

One other thing what we do is we either we choose let's say one drive for ourself, we are

trying to put our word file, excel file, everything on the one drive and say, okay, my data is

you know all stored at one place in one one drive.

Now the, now in Microsoft Fabric world, one leg is the one drive whether it is my lakehouse

data warehouse data my metadata my power be a cementing model everything is getting stored

at one place.

And the good thing about this is this is one single virtualize storage for your tenant.

Now think about you kind of have a tenant which is, you know, single tenant which is across

geographies and everything.

You need to take care about that.

And that is where you know, Fabric does it smartly that at the workspace level you can

have those regional aspect taking care of.

So virtualized as one but yet still internally you can have the security aspect or the regional

aspect covered.

And what problem do is the one leg soft compared to organization creating I don't know multiple

independent data data lakes.

See when we say data lake understand inside Microsoft fabric again we can have multiple lakehouses.

And think about it this let's say if I wanted to store it at different different places

typical previous architecture which you talk about let's say multiple lakehouses see when

we store at different places we need different security keys.

We need different management.

We have different costs around it.

Now here what is happening you are being charged for what you are storing maybe it is stored

practically for purpose in different work workspaces and different lake houses these are

just containers folders.

This is like if I go my computer I will not store everything in one drive or one folder I

will have this set of folders.

So Microsoft fabric make it so easy inside your one lake that it is one virtualized disk for

you where you have different different folders.

Now the folder could be lake house or something below that you have the data in that aspect

you are storing it.

So the entire storage at one place you are getting charge at one place now the security is

entire SSO security or your item level security which comes in and play.

So you can secure it easily it is all getting secure using one set of you know way it is not

different experience at okay if I have a data stored in let's say ADLS storage or AWS storage

I need to have different different processes for that and how do I secure and still give

access to everyone that's not a challenge here it is the same common interface to you know

put the data in and secure it for everyone using that the workspace level security or item

level security or inside even item means let's say that is warehouse or lake house or item or

entities inside that further I can go and secure it individual table columns or even the

role of a data.

And how should organization think about ownership when multiple teams consume the same data

through one lake.

See so for the ownership that is why we have these concepts of work spaces where the actual

security process starts so what we typically do is we divide the content into different

different work spaces and from those work spaces for individual items we give access to

the other people who wanted to access like I will have a master data workspace where entire

my master data besides now the same master data would be consumed by say steam it is also

being consumed by finance team so I will give the access the read access to all the people

who are consuming that data to read it and then further they can you know go add and give

the access to the end users and consumers.

Okay and when we have this architecture how important are short cut in this architecture

and when should you use a shortcut instead of a physical copying data.

Okay so shortcut could be a shortcut on external now first way to work because because

it is a shortcut for multiple purposes. So external shortcut is something basically the

data is outside the Microsoft family ecosystem and you can still create a shortcut to it.

Basically you have something on AWS or some somewhere in Google and then you wanted to

bring that.

When you create shortcut the data is typically not getting copied here it is just treated

as the data at that place and then we can when as and when we need it get query definitely

there is a cache mechanism to bring it.

Now the shortcut could be in Azure which typically it is as native storage because everything

else is another unless you are in different region.

Now within Microsoft fabric system ecosystem also we create shortcuts.

We create shortcuts to another lake houses or warehouses and bring the things together.

So what happens then you have this multiple domain architecture and we firmly believe in

not copying the data so what we end up doing is creating shortcuts so finance will have

a shortcut to master data to read the master data also shortcuts provide two more things

you can you know have your security bifurcated because of the shortcuts and one more thing

which shortcut does is basically shortcut uses the capacity of the current workspace.

So it is practically possible that I have a workspace one which is working on capacity

one and workspace to which is working on capacity to and I created the shortcuts in capacity

to workspace then it can use capacity to for that workspace and that is needed for the

power be use case because what would happen the report consumption is something which is

done by you know executives and that is the place where if your data engineering workloads

are working you don't want them to consume your entire capacity and these executives are

seeing the slowness into the report.

So you would definitely want your power be a reports to run without being worried about

what is happening to my data engineering which is continuously running in some organization

but it is reading the same data because it is the same only power be is not copying the data

now.

So when it is not copying so data though I have separated out the workspace the workspace

to is power be a workspace one is data which is lake house and the movement it queries

it goes to the workspace one where my lake house is decided to start using that capacity

but the shortcuts ensures that it uses the workspace to so I create a lake house with shortcuts

in workspace to and now it is accessing by power be I which ensures that now I am using

the capacity of workspace to not the workspace one.

So many objective no copy of data isolation of you know your workspace capacities and even

the security aspect when you try to secure and you don't want to secure some of those things

that can also be handled with shortcuts so there are many uses and it is a really whenever

it possible we should use shortcuts.

Now only thing one thing we should remember when we go cross systems is cross cloud because

in that case at every time you send the query other than caching it may end up querying

the data and sometime the other system it gives out the data it may have a cost so we

can also see what is the cost we are going to incur when we take out the data and that

will decide whether we need to have shortcut or mirroring or copying of data.

Awesome.

That is what one thing in fabric it is I can choose the lake house and I also can choose

the data warehouse when I choose what and why it is the data warehouse still existing there.

Okay so let us start with the term in the images known as our warehouse.

Basically if you call about we talking about data warehouse for almost maybe around 30 years

so there is one in this data which is known as data warehouse which is basically because

single source of thought where typically what all of us say is that it is where our goal

data is at the data which is ready for consumption for bi but when we come to the ecosystem like

Microsoft fabric we have a lake house we have a warehouse.

Now are we talking warehouse as a single source of two no we are talking here warehouse is

one of the storage piece and how the data is going to store.

So let us say lake house can store both structure and structure data warehouse can only store

structure data so one call is decided okay if I have unstructured data files I need lake

house to be there.

Now the second thing is basically the data when stored as structure data both lake house

and warehouse save it as delta format.

The only difference is the lake house delta files are governed by spark you can say and

warehouse delta files are governed by t-SQL so who is the first technology who understand

the data changes so in case of lake house the first technology which understand the data

changes is spark and in case of warehouse the first technology which understand the data

changes is t-SQL so now technology is a choice let us say we are unstructured data now where

whether I would use spark or whether I would use t-SQL that can decide my choices.

Structured unstructured was another choice.

Now there is one more thing which we have in Microsoft fabric use ecosystem is basically

seqlDB.

Now seqlDB comes in place when we want it to have OLTP kind of a system is if I want to

have create let us say app which is pretty much possible in fabric like I can create you

know application which can take the data insert data quickly in such cases we use seqlDB

now seqlDB replicate data again in data format which is very similar to warehouse.

So live versus means OLTP versus OLAP in OLAP we have a choice versus technology or versus

basically the structure unstructured data now lake house can have unlimited historical

data also we call time travel so lake house can pretty much you can control whatever you

want.

Where house is also between a range two days to one twenty days is something which you can

control one twenty is the right of the highest limit but time travel is possible so time

travel is basically I went ahead and let us update the data today but I can still see

what was my data before that update so there is more flexibility when you use lake house

because it is spark govern and the underlying data technology is more wealth versus the spark

so that is why you get more flexibility when you use lake house also it is one of the

misconception that everybody has that we cannot work on warehouse using the pi spark we

can work using pi spark the only thing is there are special snaps libraries which are there

which you have to use to work with that again there are certain limitation because where

house is still T seql govern so not all the operations are supported but you can work

with the pi spark in the now pi spark the technology because it is supported by spark or

let us say scala on it say spark are or spark sequel they have a huge extendable it means

you can run you know petabytes of data transformation everything using spark while sequel is a great

technology up to a volume of data so that is where if your system is really really large

you will go for lake house so if you are really huge on data you need to you know really

save bring the data whether I should be able to run fast but I want to do it quickly spark

helps you now on the other end sequel extend is own way is it is a same traditional sequel

way where it extend it goes you take the power and you know do does the job but I can tell

you up to you are up to few millions you will not be able to differentiate between the performance

we have implementations which have been done on lake houses and warehouses and both equally

performs good and on millions of rows sequel also never let you down it does really good

performance it shows that you know work done within time with lesser CU consumption all these

things makes it really amazing so I think it is more about choice of technology by the organization

and definitely if it is unstructured data and you need to do a lot of data science where lake house

definitely wins compared to warehouse the the thing the lake house architect

sorry require the different mindset so what if an organization already has a strong SQL and

traditional data warehouse team see see ultimately the two concepts is same now the here here is

understood even if I implement complete lake house so let's say bronze silver and gold all our

lake houses so my so-called one source of truth lake could be a lake house also which is so-called

warehouse in the functional terminology and it will all be warehouse if it is all structured data

I'm happy to be with warehouse but yes if I need to do a lot of data science and when I say data

science is not the use of AI or lm that's pretty much possible even using data is in celebrity but I

need to let's say do develop my own algorithm plus string classifications and everything that is where

you know that scale is typically more easily possible with spark and that is where probably I would

like my data underlying data to be in lake house compared to warehouse while I already said that you

know you can access your data on the warehouse using price part but definitely there is an intermediate

intervention which is required while saving so I would just if I have to do a lot of data science

stuff or I have a lot of unstructured data which I have to be I profoundly you go at the lake house

and read those the data lake fit into the story I think a little bit or what other

almanches of the data tables inside the fabric okay so when we initially so the concept of the lake

which we call data lake is basically you can store any kind of data any format of data

structured or unstructured someplace so lake house fits in that story so basically when we say

lake house the lake house the structured data is basically delta which is same as warehouse in

the in fabric world because it is also delta format in the back end but the unstructured data is

something which you are storing the file so we have the file part which we can store any kind of data

so the concept of one lake or the data lake so the concept of data lake is getting replicated

by lake house while warehouse represents the traditional sequel warehouses with those warehouses

kind of a properties as it properties and everything and yeah definitely acid properties are

something which are not supported on lake house which are supported on the house.

So I think now we understand where the data shows the lift but we need also get them there

what are the main ingestions in fabric.

So injation of the thing we can have a look at the data if it was now the second one

consists of the lake there are few sources for which the mirroring is directly available and for

others you have custom mirroring. Now the advantage of mirroring which by organization like is

basically it is self managed everywhere where we have a good change data capture mechanism mirroring

works very easy. Now Microsoft doesn't charge you for the compute as well as storage up to a limit

for bringing that time so let's say if on in fabric I have on F2 capacity I will get two terabyte of

free storage and pretty compute to bring that to terabyte. I am if I'm on F64,

our F64 is one of the very common capacities and the reason for that is one is there are very

high limitations on the data set size especially for power BA and the power BA viewer license are free

after F64 onwards. So if I am on F64 capacity the mirroring will allow me 64 terabyte of data to

be stored for free and to compute free for bringing that so that is my second choice or you can say

first choice if I don't come short course then comes my ability to bring in data using pipelines which

has copy activity and copy job. Copy job again is having very good CDC a co-options available and copy

activity also now both are meant for the purpose. So if I want to bring the data these are my choices

but yes the pipeline copy activity and copy job has around 40-50 connectors and that is where data

flow wins because data flow gen 2 has more than 200 plus connectors it is it has power query which

can also do the transformations. So data flow gen 2 would be your fallback option most of the

cases when you are not able to connect and all of them has generic connector like RSTPI and we

also have the ODBC generic connector or we have OLEDB all these things are also available web

connector etc. So our choice would be finally if we are falling back we use data flow gen 2 to bring

in the data it has maximum number of connectors and it has also fast copy connector now which will

make sure that it is as good as your copy activity in most of the cases definitely transformation

choices are there in case you are bringing from data flow gen 2 you can transform the data also

there. So now because we have this path and Python inside the system if your data is available

online right now it is not supported for on-premise sources you do have now various connectors available

at the notebook level previously also we were able to bring the data in but we have to mention all

the credentials in the notebook but for now we have connections which is basically supported

in a manner that you don't have to show your credentials in the notebook and that again creates one

more way to bring in your data now when you bring in the notebook the handling of the data is

mostly in your end it is basically how you bring how much you bring in everything while in all

other cases it is managed by the system so you have these many choices to bring in data.

And what would you say how important is power-carry knowledge for working effectively with

the data flow gen 2? See power-carry is something which is data flow gen 2 is having basically it is

same as our power-base power-carry or your our data flow gen 1's query and if we know power

query it would be really helpful because I can tell you those who are migrating from power-bava

word their first choices data flow gen 2 those who are coming from Azure word their first choices

typically the data pipelines and let's say PISPAR transformations and everything so we will get very

familiar things whenever we are coming from a different you know world to the fabric world.

fabric is new so somebody from the power-bava world will come somebody from the let's say

the PISPAR engine world will come somebody will come from different migration kind of a

pools which is very similar to pipeline you will come so orchestration everything it happens so

we will find some very similar tools here in Microsoft fabric to start with.

Okay awesome I think is there performance on main

tiny ability limits that we should know about in that of low gen 2?

Okay so this is a really easy so teach gen 2 is very good for a small to medium size of data

but when you come to a larger scale the transformation in PISPAR can excel much faster than the

in data flow gen 2 while Microsoft is continuously investing and consistently improving the

performance of data flow gen 2 when it is large volume of data but it becomes really easy

because Spark is PISPARC or Python is like an open language you can have tons of things to do

and manage it and that is why you know if it is a really large scale data transformation we

prefer PISPARC. I don't have a data flow gen 2 versus PISPARC but data flow gen 1 versus

PISPARC performance could be having 80% gain both in terms of compute as well as in terms of time

that is a proven fact but we have observed at the customer level. So data flow if you had a data

flow gen 1 if you completely converted into PISPARC because save a lot on the compute units as well

as the time both you are saving it is not one we are saving and that is where PISPARC could be a

really good choice for data transformation going forward in the future for large scale data.

That is interesting for someone coming from the PowerBRS-L part when they actually need PISPARC.

So the data transformation part is something which is basically instead of SQL can happen on

the PISPARC and when I say PISPARC I have this habit of calling PISPARC but we can use these Spark

notebooks which could be in Scala which could also be in Spark SQL. So if you are from SQL

where you will find Spark SQL very familiar to you. If you are from let's say power query world

you will see PISPARC almost has similar kind of functions which you have let's say if I want to

join they are also have I have table dot let's say join something and here also I have the data

frame dot join. So very similar functions I have here also to do one one on one operations with

compared to power query. So definitely if we are coming from let's say SQL plus power query world

we can start on fabrics with those but as the data goes is the things grows slowly we have to start

moving some of our stuff to PISPARC or maybe we can say Spark notebooks. Yeah yeah I think that's

interesting how do notebooks change the way data engineer teams work actually. See we have to

understand that the notebooks especially running by Spark is made for scale.

So because we can decide you know how much compute they should use they should use 10 node

small cluster or a step small cluster or large cluster all those things we can do and that

scale comes up with Spark and because of that you know you might do need to do heavy lifting

and that heavy lifting you want it to do and you also want it to save time not only the volume

of data increasing but you also want to reduce the time that is where easily can happen Spark and

Spark is meant for the entire illusion of Spark happens that okay I have the huge data how do I

want to process that data and get the time in time processing in quickly and that's where the

Spark was entered that you know okay distribute this entire jobs on multiple nodes and then do it

and then bring the results in and that's where the Spark notebooks are providing you here in

Microsoft fabric. So for scale and for you know getting the results faster the Spark always helps you

that's so sounds interesting is there any performance mistake begin us to with Spark?

The thing which we have to do is when you start as a beginner don't try to play around with

these Spark settings of fabric also secondly instead of dealing with it in the notebook level

please create those different Spark clusters settings at the workspace level or capacity level and

use it because I remember a case where what happens one of the team and we have allowed the setting

of the Spark cluster at the workspace level one of the team is created such a big cluster

and they were able to complete their job in one hour they were really happy but they end up

consuming the entire capacity for all other performance so those could be there and again

there is a little bit of there are less of control sometimes required in the fabric spark because

some of the things are already handled by fabric so you might not to explicitly go and start Spark

it automatically creates sessions for you so the small unlearning here that you know you already

already have a Spark session you only to explicitly started you can start using it and then there

are only selective parameters which it takes if you take care of these things I think it's pretty

smooth journey and even I think for a beginner who doesn't understand as a Spark works it is much

easier journey because you are not taking care of where is my root folder how should I start the Spark

I just know that I have a Spark session with all those capabilities and I can use it.

Awesome another topic I think the biggest buzzer words is the it's yeah I don't know the mid-medallion

architecture and I think it's a little bit of that marketing that we call it Bronx the Servant Gold

have have these vocabulary are modern how how how important are the the medallion architecture

for for fabric see for see architecture should depend on you know organization and what kind of

data we have see if war organization is having a very strong master data management done at the

source level and data cleaning was a real big focus probably you will have very clean data which

is coming into the grass layer itself then you daily need silver or then can you directly go to

silver plus gold layer that is the call you have to take because I firmly be the number of layer should

you entire dependent on the kind of data instead of you know I want to follow medallion architecture

because when we initially started the data engineering if you remember we used our staging area and

the warehouse the one source of truth we we were you know doing that entire as a part of transformation

maybe we were creating the intermediate table not explicitly calling it silver now because

this has been a well defined structure of medallion architecture we have bronze silver and gold

so if you have a very strong master data foundation which is laid outside probably just think about

do you really need one silver copy because you might have clean data but yes if the master data

foundation is not so great then we need to bring in bronze create a really strong clean layer for

the consumption of data at silver layer and then create the gold layer so and I see most of the

organization at some stage or other having this challenges of master data because it's not managed

so greatly so probably if you can say 80% of the organization will fall in a category when they have

to follow the medallion architecture of bronze silver gold yes there are organization who has

really strong on you know data quality and data master data management and everything

probably for them silver is not too much other than you know having a following an architecture

probably they could have bronze and silver gold together while doing the transformation putting

it into the gold yeah that's interesting so so where should the the the business logic

live so should it be the logic primarily exist in transformation in the warehouse in the

thematic model or some well now this is a testing now there are so this is a little bit of

the business logic which are little complex and the lying level we cannot keep them at the

semantic model level ideally speaking I would be happy to have all my major able or KPI's or

performance indicator at the semantic model level because it's easy it's dynamic I don't need to

reload the data for doing it so anything which I can keep at that level but the problem is if the

calculations are at lying level if I have millions of rows where I need to go and do one calculation

and then do it up it is not going to work out at the semantic model level and in such cases we always

prefer to have the calculation done at the gold layer means basically when we load our data from

silver to board we will apply those business logic so business logic is shared between the gold

layer creation and the semantic model our preference is to have it at the semantic model level because

they can change easily we can change formula we can create new matrices but definitely we don't

want to do role level calculations for that so we have to take a cautious call to make it at the

two levels that's how it moves to a side awesome and that's a little bit move closer to the

business layer where are the semantic models so important and what makes the good semantic model

okay see if you actually would ask this question we could be couple of your banks

or three years back for a semantic model is for consumption of power VA isn't it so we have a

semantic model we create power VA reports on that but within fabrics the things have changed

especially after data agent ontology and fabric app is all changing this playground and creating

the bigger role for semantic model so if you ask me if AI brings a challenge to let's say power VA

as a visualization where power VA is not facing challenge is the semantic model because

understand one thing our all business logic are going to decide the relationship and the key

majors and their definitions are known for semantic model and if I work on a semantic model I don't

need to tell the definition what is my net sense semantic model knows it what is my gross is what is

my revenue definition everything semantic model knows it and let's say if somebody need to do and

go and do some transformation on directly on my lake house and warehouse they don't know this

definition now we have a semantically lab library which can enable you now our data agents also

work on semantic model it means the data agent wanted to take the advantage of the definitions which

are not similarly the ontology can be created on our semantic model it which is going to take the

advantage of already existing knowledge by the semantic models the relationship the KPIs

now we have fabric apps which I got wrong recently so one of the things which has happened because

of this AI explosion over the last few years that we came to something which we are doing let's say

maybe a 20 years back that I can have my custom created report which whatever I want

and then from there in that place we move to the tools because it is heavy maintenance because

I can create my report let's say on Java technology by connecting to database and do everything

and have slice and dice and everything but it was a huge maintenance I can't do it and that's why I

knew tools like Power BI tab you to just do it with drag and drop but what AI has done over

the last couple of years is you know you just go type in some command and your react dashboard is

ready you say okay I need all these features it's already and that's where you know we were coming

back okay why should I have you know executive only looking at you know reports which is

fixed in the boundaries of what the tool can create and that is where you home we have seen a lot

of people were developing these AI apps and the challenge of all those AI apps is they understand the

warehouse or lake house better because they are having sequel and points and you the most the most

common technology to be used by anyone is sequel or as a sequel and because of that developing here

but our logics for lying in semantic model who will transform those major or those

capyra definitions into the sequel and there's again work and that's where Microsoft came up with

fabric app enabled by Rafe in framework now it can create these those reports those it's so it

used typescript and you can create those reports which all the kind of UI you need the custom UI

the kind of slicers you need the visual out of the box visuals and everything you on your semantic

model and that's what we call fabric and then it is again getting published inside your fabric works

pieces so same security same everything so now the role of the semantic model has greatly increased

in terms of relationship and the definition it owns but yes the consumption could be power be

a visualization could be data agent could be ontology ontology driven data agents or maybe fabric apps

but yes we have a layer which store which knows our business which understand our business well

and it can be used for the downstream by anything

and and what would you say how should teams approach measure of firsts calculated columns

okay so when we say calculated columns probably I would like them at the gold layer other than

these calculated columns which are now which is very interesting thing which has very recently

happened in the power be a world is contracts driven calculated columns now contracts given calculated

columns means they change their value because of user climate it also help us in you know

some time to hide some data basically if I wanted to do data masking it also helps in that

and sometime it helped in changing the column basically I want to show different column so if it is

if your column is not a context driven column it should happen at the gold layer all the majors

anything which is a major should be going to happen at the semantic model level context even calculated

columns should happen at the model level semantic model level and rest of the calculation which are

line level or really complex which is which can slow down at the runtime should be moved back to the

bouldier. If I worked with fabric and probably iDux was the I don't know the main topic and I see

Aids awesome in generating DAX code now did you think DAX stays important or makes tends to learn DAX

anymore see I think in the new world what is happening because AI is doing everything like I am when

I created a let's say referring dashboard it uses typescript I am not sure what typescript is

how it is written what is the structure of that and everything similarly it is happening with all

the technology like probably you can generate a pi spark notebook using co-pilot or even AI

that is to the DAX also but we have to understand when the things does not work and when the things

need to be optimized and that is where expertise comes in place now some of us who are there in the

industry for some time we had gained that expertise over a period of time now for the new learners

it is very easy to get a DAX from AI so do I really need to have an expertise I think in a longer run

at least for next few years you need to have that expertise on the subject I know the technology

is really fast and maybe after some time it is more on the logics and less on the technology so

probably the language would be you know using your logics and explaining your logic to the AI

by our prompts not let's say if C Java DAX of our query it's just the thing which internally getting

created and getting debugged also by AI only but yes definitely knowing something in depth how it

works it's always really important because in that case you can go and fix something which is

actually not getting fixed by AI or which is getting not optimized by AI but as this intervention

increase probably up to me loving you know how the algorithm works how the logic works at different

places is more important than you know the language is what we are working use in future.

And I think another topic a lot of people aware of actually it's it's security I have all my

data now and well like so how should we start with security in fabric architecture?

See in terms of the security start from your workspace is from where you can start controlling

by rules whether you are a viewer you are a member you are a contributor or you are an

advent that's the first step now in the fabric what typically we are trying not to have the workspace

level access we try to have item level access like in a lake house or warehouse I can have a read data

read and write and all those permissions then if further goes down I can go to the item level now item

level permissions typically we will try to control using one-leg security wherever possible

and if we are not putting one-leg security we are going to have semantic model level security

a one-leg security has came I think around one and a half to years back and it is still evolving at some

places it may not provide you that option and that is where we go and put the security

semantic model level but one-leg security is something which we are going to use

and that is where we can secure the data means we can do both the object level security

OLS and RLS using one-leg security so what we are going to do here is basically where but possible

in terms of OLS and RLS if possible we are going to use one-leg security if we are unable to control

at one level one-leg level or we don't want to contain at one-leg level basically

lake house of warehouse we are going to do it at the semantic model level and definitely object

level security needs a set level security basically or report so do your warehouse on

lake house we can control them at the works is either the workspace level or individual item level

so from workspace to item to inside the item with one-leg security

awesome so the how how do is the role level security works with the semantic level correct

or what's the best practice see the role of a security basically now it can be done at two level

one is basically at the one-leg security which is can again you can go and define okay I for

this is the filter for that so in one-leg you decide the filter and then you create a role and

that can assign to the people so that filter will get applied so you will let's say only get north

I create a filter it is north and it will get applied at the end all the endpoints SQL endpoints the

the semantic model endpoint or even at the file endpoint now when I come to the semantic model the

semantic model security of power be a remain same so it is the same way we create a role and inside

the role we again define a dynamic security using a table where we have the email addresses using

user principle a model static security by like saying reason equal to north and then we once the

semantic model is published we go in the security model layer and add users or security groups

preferably we should use security groups and security groups should be created like you know

this is a viewer security group or this is a financial viewer security group this is says viewer security

group or says contributor security groups we should create security groups and assign those security

groups to the security layer or to the report or to the one-leg security that is always a best

practice so so we have the workspace permissions we have the data permissions and we have the semantic

model security but how how to gather when this when when an organization has hundreds of thousands of

users and with different requirements yeah see I think the why we have these multiple layers and

but all these are all you know really smooth and integrated so the reason is let's say if I want

to have few people to be admin and having access to everything I don't have to you know go and

do multiple permission I can deal with workspace level permission and they can access everything

in the workspace they want but there are certain set of user I don't want to give them you know

access to everything let's say we happened with us also there are report user who are report builder

as well as the report viewer but we don't want to expose them the let's say warehouse is or

lake houses so we control the security at the semantic model level we give them access and we make

sure that it happens in such a manner that they don't have to see other objects so all those

possibilities are covered and that is why you have this two three years now we can do some AI

automation for that to make sure it happens like basically the security when we create the security

group on a larger organization these are also not manually created you have let's say a service

request which I go to raise and you will get added to a security group the same request may add you

to let's say a RLS table also where you get an access to a region to a city or to a product also

and then those RLS tables are used in your semantic model to govern the security and similarly

the assignment also happens let's say I asked for sales viewer role so sales viewer role with let's say

city as new new york so the service request will go add me to the sales viewer security group it will

go mid go and add me to the new york it will also go and add me to the let's say workspace or to the

reports which I can see as the sales viewer so all these automations can help us out. And I have the

fabric admin center I also have I don't know it's power be either same name but I also have the

option in power be I and and I also can use use entra how show I had all these tools or what's

the best practice here. So, so we have it that level so basically entra all the security where

where basically your all log is getting created and security groups are also getting created at the

Azure level so that part happens there after that fabric and power be I is the workspace level

security same way to apply suppose that we give the assignment of the objects on the fabric

level to the security groups permissions are also decided security or sometime you will let's say

try let's say fabric app came is a new feature now I don't want everybody to go ahead and create those

reports so I will give permission to a set of users which are part of a security group to create

those reports same way I'll give the finance user access to let's say as a contributor or a member

or admin to you know build the content and give it and viewers of those let's say finance

report security group of finance viewers I'll add deadness the viewers and also add them to security

so this is how it happens so definitely entrap part of the Azure creation of your emails and everything

not going away so that will happen security group creation should also happen from there and

rest of the security will be managed at the fabric level. Yeah I think we have a lot from from

the engineering part but yeah ultimately the most business users don't care about spark data

levels or like house they care about answering questions how do we bridge these these get

so the fabric has lots of options basically if you ask me future could be talking to your data

agents right now we also see a lot of people started using data agents and in that data agent

ontology driven data agent is something really important because what happens when you create ontology

on top of your data it do understand data much better it goes look at the data and try to have the

data part match and it can provide better answer and it can sometimes work even better on the

snowflake models or the models where you know the typical star schema kind of stuff is not there

it also works better there so ontology driven data agents would be our first choice in the future

to communicate with our data and understand our data then we will have the reports definitely

I will not wanted to go and communicate and always ask questions I would definitely want some

numbers in front of me every day morning now those everyday morning numbers will either come from

a fabric app report or power be a report there are very high chances in the future that your data

agents might create a report for you which is basically could be a fabric app and a power be app

after you ask a set of question okay this is how I perform these are the KPI I look at it

this is the information I need Monday morning this is what I need to do morning so data agent probably

could go ahead and create finally based on your communication a report which works the VR way

and not you know the report created by a developer which gives you okay same set of reports for all

four executive same set of reports for you know thousands of users so probably data agent and the

data agent created reports would be something which we will be frontending in future yeah I think a

lot of power be I and business users like to have a self service be I it's a little bit distressed

us from our most organization what do successful service ourselves levels I look like

okay so now see understand we are talking about this town self service but when we use that

creating those reports the movement we go into the complex calculation we realize we need to go

back to the developer even in the power be in world where you know you have the self service knowledge

of decks was required but in this new world of AI driven you don't need that even if a major is not

there your AI agent is going to create so this is real self sir your model doesn't need to have all

the formula your model doesn't need to have all the logics it will on the flight created now definitely

realize okay if it creates on the flight takes lot of time it's better to improve my model to answer

some of those questions maybe better to have a major for that so I think two self surveys what we

are calling is on the plate using data agents and the AI driven development so where you know

users will just be so I'll tell you I was on the weekend I have created entire fabric from

ingestion of data to transformation creating a power be a model creating a report just by talking

to the GitHub co-pilot I have not went ahead in any of the UI other than the testing so we're reaching

to that level now understand what would be the future that you define your requirement well to

define your requirement well you have tools also which will define your requirement well then they

will break it and then you will only go ahead and say okay go ahead and create the tickets on my

DevOps and then you will come to GitHub co-pilot and say okay execute all these tickets on my one

and this getting created and you do so the knowledge of technology only needed for the debugging skills

for what is not working how to make it work but it's going to be a really self-serve environment like

you know in our generation we used to learn computers but if you look at the next generation they

they know computer by you know they they know is as a skill be it's just like a language skill for

that for future the AI would be cut that kind of skill where everybody knows how to you know build

those complex some of those systems using you know AI so that will become a primary skill for the

next set of people who are you know now learning those things so probably for them it would be just

okay I go or I type prompt I'll have my stuff ready for me awesome yeah thank you are we really

little bit running out of time so I have in every session of rapid fire rounds so I ask a short

question you give a short answer is it though are you ready out yeah okay let's also warehouse

lecels escuelor pi spark pi spark dataflow gen 2 or notebook not book back so escuel

decks most underrated fabric feature fabric apps the biggest power be I'm modeling mistake

many many relationships what is when such an gallia called you and say hey you get all the money

resources to build your dream feature what will you build probably I would want to improve

upon on data agents so just ask them okay this is my requirement have a dashboard ready every day

it should change based on some requirements I don't want to see same dashboard every day one

fabric skill every power be I professional show to low probably you should be knowing power query

index if nothing else is fabric apps in you power be I yes awesome so and we should I invite

next at what questions do I ask I think you can this in Rajin who's again Microsoft fabric MVP

and scheme more questions about ontology and fabric app I think he's one lot of working on

real-time analytics actually he's working a lot on real-time analytics so maybe you have a session

on fabric real-time with Rajindra the data platform MVP and if an organization wants to start with

fabric tomorrow what should it's first project be I think any data project where you think I need

some kind of analysis it just in a starting point it could be just you know knowing you know

how my employee performance is there or how my sales is ongoing if you think you need data analysis

fabric is there for you awesome so yeah I'm a thank you for for joining me and talking taking

us through the computer the Microsoft fabric analytics journey what I think in this conversion

makes clear is that fabric isn't simply another in real analytics product real opportunity comes

from connecting data in guest gen engineering lake house where our architectures and

semantic model security power bar into one core and platform and the technology alone isn't

enough the architecture ultimately has has to deliver trust information that people can use

to make better decisions so for all the people to listen to the podcast and interested to connect

and see how it works you can look at the show notes and you find this profile there and yeah

thank you again I'm it for for being a part of the show thank you thank you for inviting me

This transcript was automatically generated by the podcast creator and may contain errors. Aggregated via the PodcastIndex API.