Microsoft Fabric End-to-End: From Raw Data to Business Decisions with Amit Chandak [MVP]
Microsoft Fabric brings data engineering, analytics, business intelligence, governance and increasingly AI together in one platform. But what does an end-to-end Fabric architecture actually look like when you move beyond individual features and start connecting everything?In this episode of the M365 FM Podcast, Mirko Peters is joined by Amit Chandak [Microsoft Data Platform MVP] for a practical journey through Microsoft Fabric — starting with raw organizational data and ending with trusted information that business users can use to make decisions.
WHY MICROSOFT FABRIC?
Before Fabric, organizations could already build sophisticated analytics architectures using Azure, Power BI and other platforms. The problem wasn't a lack of technology. In many cases, it was the opposite: organizations had too many choices, separate storage technologies, different compute models and multiple copies of essentially the same data.Amit explains how Microsoft Fabric attempts to simplify this architecture by bringing workloads together around shared foundations such as OneLake, common Fabric capacity and the Delta format. Lakehouses, warehouses, Power BI and other Fabric experiences can therefore operate as parts of a broader platform instead of completely isolated services.
ONELAKE AS THE FOUNDATION
OneLake is one of the central concepts behind Fabric. Amit compares it conceptually to OneDrive: instead of every analytics workload creating completely independent storage environments, OneLake provides a virtualized storage foundation across the Fabric tenant.Organizations can still separate data through workspaces, Lakehouses, Warehouses and domains, but those resources exist within a common Fabric storage architecture. This becomes particularly important when organizations want to reduce unnecessary duplication while maintaining security and organizational boundaries.
CENTRALIZED DATA OR DATA MESH?
Fabric doesn't automatically mean putting everything into one giant centralized analytics environment.For smaller organizations, a centralized architecture may still work well. As organizations become larger, Amit sees increasing value in domain-oriented architectures where areas such as sales, finance and purchasing can have their own workspaces and responsibilities.IT can remain responsible for availability, governance and the technical foundation while business domains increasingly take ownership of how their data is analyzed and consumed.
SHORTCUTS INSTEAD OF COPYING DATA
One of the recurring themes throughout the conversation is avoiding unnecessary copies of data.Fabric Shortcuts allow teams to reference data stored elsewhere rather than physically copying it into every environment that needs it. That can apply both inside Fabric and to supported external storage.Amit also explains an interesting architectural benefit of shortcuts: they can help separate workloads across capacities. This can become important when organizations want Power BI consumption workloads isolated from intensive data engineering workloads while still working with the same underlying information.
LAKEHOUSE VS. WAREHOUSE
One of the biggest Fabric architecture questions remains: Should you use a Lakehouse or a Warehouse?A Lakehouse can work with structured and unstructured data and is naturally aligned with Spark. A Fabric Warehouse focuses on structured data and provides the familiar T-SQL experience.Both ultimately use Delta for structured data inside Fabric, which means the decision increasingly comes down to the type of data, preferred technologies and workloads.Organizations with strong SQL teams don't necessarily need to abandon their existing skills. Teams working with very large datasets, advanced engineering scenarios, unstructured information or extensive data science workloads may find the Lakehouse and Spark approach more attractive.
GETTING DATA INTO FABRIC
Once the archi
Welcome everybody to the new episode of the MC65FM podcast.
Today we are going to end to end with Microsoft Fabric.
Not looking at one isolated feature,
but at how the different parts of the platform
come together to create a modern analytics architecture.
How should organization think about leg house versus data warehouse,
where do spiced back belong?
When should we use data flow gen 2?
How do the mental models fit into the architecture?
And ultimately, how do we turn all this technology
into power BI solutions that business user can actually use
and make a better decision?
My guest today is Amit Shaktaq.
Microsoft data, Microsoft data, platform MVP and power BI community
super user, Amit has more than 22 years of experience
in data analytics and currently leads data engineering
and the leg solution at Canerica is experience response
Microsoft fabric power BI SQL databases, Tableau,
in quarter data engineering modeling and visualization.
Today we are going to build the Microsoft fabric analytics
story from the ground up from architecture
to ingesting through transformations, storage,
semantics models, security performance, and final power BI.
Amit, welcome to the MC65 podcast.
Thanks, thanks for inviting me for this podcast.
It's my pleasure to be part of this one, Peter.
Awesome.
Before we get deeply technical,
how did your journey into data and analytics begin?
Yeah, so I joined Oracle in 2003.
I selected Oracle as an act of campus and I became the part
of their BI team.
They were going to one of the transformation at that time.
We were building this tool which is known as the
BI daily business intelligence and that's where I started my journey.
After I left Oracle, I started a company along with one of my colleagues
and I ran a data analytics startup for 10 years.
We created our own tool very similar to Power BI directory.
We used to have the tool completely on the web.
At that time, you know, having a complete web
all thing was a challenge.
So that tool was completely authored on web, designing reports,
creating the semantic models and everything was on the web.
So complete web all thing was there.
And then I done that for 10 years and post that I decided to move on.
That's where I joined my current company, Connecticut.
I also moved to Microsoft technology around the same time.
I became a Power BI community super user around 2020 and then around 2022,
October when I became Microsoft MVP for data platform.
Awesome.
That's, yeah, that's the way it went.
22 years, it's a lot of time.
But yeah, that's direct jump into the fabric part.
So let's start with the big question.
What problem is Microsoft fabric actual trying to dissolve?
See prior to Microsoft Fabric, see, it's not that we are doing data
engineering something new with fabric.
The whatever components fabric has, all those components
were previously available at different places.
So if you look at Azure ecosystem, you could have, you know,
100 plus choices to do what fabric is doing today.
If I look at the larger ecosystem, if I include Azure,
AWS, even Google platform, I might have 1000 plus choices to do my own data engineering work.
But the challenge was bigger.
Now choice of technology was one challenge.
Then other than choice of technology, create a multiple copies of data.
What happens in the data analytics or BIA ecosystem?
We say single source of proof.
We always talk about it.
But our own ecosystem, because we sometime we need data lake,
another time we need data warehouse and then we need powerful tools like power
BIA, which for input would be our self creating copies of data.
And all these tools are working on different technologies is they have different kind of back ends.
Their storage were different.
They they were created at different time frames.
So all of them were, you know, built differently, storing differently.
And they were using different ways compute powers getting compute.
A Microsoft node challenge and I think that's where they start in the center apps where
they wanted to address this one.
A fabric name was even higher where they sorted out the entire stuff.
Whether you have the structure data and structure data, where you wanted to have warehouse,
where you wanted to have power BIA or even real time analytics, they were able to put everything
at one place.
And they starting, you know, with a few foundations like one lake where, you know, your entire
data assets, whether it's a metadata data can decide.
Then they solved the problem of compute by saying, okay, all the workloads in Microsoft
fabric can share one single compute.
Now you can have multiple computers per requirement, but that that is one of the biggest achievement
because think about your running spark workload.
How would you compare that with power BIA or CQVV?
They're all together.
But everything gets converted to compute units and getting chard within the same number
of computers which is available with you.
Let's say paradew per 30 seconds.
Then comes the challenge which is there with the storage.
Now I have storage, which is available for the greater work, but what happens when I
store the data, let's say for lake house, I might prefer some format.
Now where else doesn't understand that format and that's why I was copying data from lake house
to where else when I was doing previous term.
Similarly, power BIA was not comfortable with data warehouse formats.
So it was creating on copy.
Now whenever we store the structured data, whether it's a lake house, whether it's a warehouse,
it all getting stored into delta format, which created that uniform format, which lake
house understand, which spark understand, which T-SQL understand, which even the power
query engine has been modified to understand that.
So it means I am not unnecessary creating copies of data.
Definitely most of us follows the medallion architecture and in that we will have, you know,
bronze, silver and gold, but it doesn't mean that I'm going to create unnecessary copies
of my bronze data and unnecessary copies of my gold data.
So these are some of the challenges which existed while, you know, we have the ecosystems
available for last 15 to 20 years to do data analytics.
But these are some of the challenges which emerges because of the growth of the data, which
is there with us and Microsoft fabric has, you know, went ahead and addressed that.
And how much to do an organization's data practice need to be before a depth thing fabric?
See Microsoft fabric is basically complete end to end solution for your data analytics need.
So even if you are starting for data analytics, Microsoft fabric can be there.
So the only in the chart is needed at the source level that you understand your sources
and you understand what is needed for data analytics.
If you are ready to analyze your data, if your data is arranged, let's say, if your data
is completely unorganized, you are, you know, still working on some of those spreadsheets
a lot, some of the, let's say ERP or CRM implementations are still ongoing or on middle of it, then
it might not be the right time.
But if you are a data mature company in terms of capturing the data and wanted to do
entities or already there with some analytics and wanted to transform it to the model system,
I think in both case you are ready for data and it is on Microsoft fabric.
It provides you everything these days.
If you, if you say I just wanted to do simple, simple, create very simple report using
Powerway, that's also possible.
And you want to follow the complete architecture, you know, bringing data to the lake houses,
transforming it and create warehouse and then wanted to do it.
That is also possible in Microsoft fabric.
What, what did you think, how important is domain oriented architecture in fabric or
or short organizations create one large centralized fabric environment or distribute responsibility
across business domains?
I think we are in an era where we talk about business owning these stuff where we talk about
data mesh architecture where we talk about domains.
So in the modern architecture, why, you know, your IT could be custodian of your data,
but you need to create, you know, multiple domains in which the data is going to decide
finally.
So you can have a master data separately analyzed for the organization level because master
data management is key for all the analysis post that I think it is based on let's say
sales, finance, and purchase can have their own workspaces can they have their own domain
and the data can decide.
If you are a small to medium size organization having a centralized data warehouse, a works
means it can have anything, but if you are a medium to large enterprise or a very large
enterprise, it is always good to have, you know, data mesh architecture where you have
the domains and, you know, the ownership lies with business more than IT.
So it is the combined governance, while the data, entity and the data availability ensure
by IT, but the analysis part of it, distribution part of it and the management part of this is
governed by the business.
And I think the core term when we talk about Microsoft Fabric is it's one leg, what is
one leg in practical terms?
Like today what happens?
Let's talk about our own stuff when we are trying to go online.
One other thing what we do is we either we choose let's say one drive for ourself, we are
trying to put our word file, excel file, everything on the one drive and say, okay, my data is
you know all stored at one place in one one drive.
Now the, now in Microsoft Fabric world, one leg is the one drive whether it is my lakehouse
data warehouse data my metadata my power be a cementing model everything is getting stored
at one place.
And the good thing about this is this is one single virtualize storage for your tenant.
Now think about you kind of have a tenant which is, you know, single tenant which is across
geographies and everything.
You need to take care about that.
And that is where you know, Fabric does it smartly that at the workspace level you can
have those regional aspect taking care of.
So virtualized as one but yet still internally you can have the security aspect or the regional
aspect covered.
And what problem do is the one leg soft compared to organization creating I don't know multiple
independent data data lakes.
See when we say data lake understand inside Microsoft fabric again we can have multiple lakehouses.
And think about it this let's say if I wanted to store it at different different places
typical previous architecture which you talk about let's say multiple lakehouses see when
we store at different places we need different security keys.
We need different management.
We have different costs around it.
Now here what is happening you are being charged for what you are storing maybe it is stored
practically for purpose in different work workspaces and different lake houses these are
just containers folders.
This is like if I go my computer I will not store everything in one drive or one folder I
will have this set of folders.
So Microsoft fabric make it so easy inside your one lake that it is one virtualized disk for
you where you have different different folders.
Now the folder could be lake house or something below that you have the data in that aspect
you are storing it.
So the entire storage at one place you are getting charge at one place now the security is
entire SSO security or your item level security which comes in and play.
So you can secure it easily it is all getting secure using one set of you know way it is not
different experience at okay if I have a data stored in let's say ADLS storage or AWS storage
I need to have different different processes for that and how do I secure and still give
access to everyone that's not a challenge here it is the same common interface to you know
put the data in and secure it for everyone using that the workspace level security or item
level security or inside even item means let's say that is warehouse or lake house or item or
entities inside that further I can go and secure it individual table columns or even the
role of a data.
And how should organization think about ownership when multiple teams consume the same data
through one lake.
See so for the ownership that is why we have these concepts of work spaces where the actual
security process starts so what we typically do is we divide the content into different
different work spaces and from those work spaces for individual items we give access to
the other people who wanted to access like I will have a master data workspace where entire
my master data besides now the same master data would be consumed by say steam it is also
being consumed by finance team so I will give the access the read access to all the people
who are consuming that data to read it and then further they can you know go add and give
the access to the end users and consumers.
Okay and when we have this architecture how important are short cut in this architecture
and when should you use a shortcut instead of a physical copying data.
Okay so shortcut could be a shortcut on external now first way to work because because
it is a shortcut for multiple purposes. So external shortcut is something basically the
data is outside the Microsoft family ecosystem and you can still create a shortcut to it.
Basically you have something on AWS or some somewhere in Google and then you wanted to
bring that.
When you create shortcut the data is typically not getting copied here it is just treated
as the data at that place and then we can when as and when we need it get query definitely
there is a cache mechanism to bring it.
Now the shortcut could be in Azure which typically it is as native storage because everything
else is another unless you are in different region.
Now within Microsoft fabric system ecosystem also we create shortcuts.
We create shortcuts to another lake houses or warehouses and bring the things together.
So what happens then you have this multiple domain architecture and we firmly believe in
not copying the data so what we end up doing is creating shortcuts so finance will have
a shortcut to master data to read the master data also shortcuts provide two more things
you can you know have your security bifurcated because of the shortcuts and one more thing
which shortcut does is basically shortcut uses the capacity of the current workspace.
So it is practically possible that I have a workspace one which is working on capacity
one and workspace to which is working on capacity to and I created the shortcuts in capacity
to workspace then it can use capacity to for that workspace and that is needed for the
power be use case because what would happen the report consumption is something which is
done by you know executives and that is the place where if your data engineering workloads
are working you don't want them to consume your entire capacity and these executives are
seeing the slowness into the report.
So you would definitely want your power be a reports to run without being worried about
what is happening to my data engineering which is continuously running in some organization
but it is reading the same data because it is the same only power be is not copying the data
now.
So when it is not copying so data though I have separated out the workspace the workspace
to is power be a workspace one is data which is lake house and the movement it queries
it goes to the workspace one where my lake house is decided to start using that capacity
but the shortcuts ensures that it uses the workspace to so I create a lake house with shortcuts
in workspace to and now it is accessing by power be I which ensures that now I am using
the capacity of workspace to not the workspace one.
So many objective no copy of data isolation of you know your workspace capacities and even
the security aspect when you try to secure and you don't want to secure some of those things
that can also be handled with shortcuts so there are many uses and it is a really whenever
it possible we should use shortcuts.
Now only thing one thing we should remember when we go cross systems is cross cloud because
in that case at every time you send the query other than caching it may end up querying
the data and sometime the other system it gives out the data it may have a cost so we
can also see what is the cost we are going to incur when we take out the data and that
will decide whether we need to have shortcut or mirroring or copying of data.
Awesome.
That is what one thing in fabric it is I can choose the lake house and I also can choose
the data warehouse when I choose what and why it is the data warehouse still existing there.
Okay so let us start with the term in the images known as our warehouse.
Basically if you call about we talking about data warehouse for almost maybe around 30 years
so there is one in this data which is known as data warehouse which is basically because
single source of thought where typically what all of us say is that it is where our goal
data is at the data which is ready for consumption for bi but when we come to the ecosystem like
Microsoft fabric we have a lake house we have a warehouse.
Now are we talking warehouse as a single source of two no we are talking here warehouse is
one of the storage piece and how the data is going to store.
So let us say lake house can store both structure and structure data warehouse can only store
structure data so one call is decided okay if I have unstructured data files I need lake
house to be there.
Now the second thing is basically the data when stored as structure data both lake house
and warehouse save it as delta format.
The only difference is the lake house delta files are governed by spark you can say and
warehouse delta files are governed by t-SQL so who is the first technology who understand
the data changes so in case of lake house the first technology which understand the data
changes is spark and in case of warehouse the first technology which understand the data
changes is t-SQL so now technology is a choice let us say we are unstructured data now where
whether I would use spark or whether I would use t-SQL that can decide my choices.
Structured unstructured was another choice.
Now there is one more thing which we have in Microsoft fabric use ecosystem is basically
seqlDB.
Now seqlDB comes in place when we want it to have OLTP kind of a system is if I want to
have create let us say app which is pretty much possible in fabric like I can create you
know application which can take the data insert data quickly in such cases we use seqlDB
now seqlDB replicate data again in data format which is very similar to warehouse.
So live versus means OLTP versus OLAP in OLAP we have a choice versus technology or versus
basically the structure unstructured data now lake house can have unlimited historical
data also we call time travel so lake house can pretty much you can control whatever you
want.
Where house is also between a range two days to one twenty days is something which you can
control one twenty is the right of the highest limit but time travel is possible so time
travel is basically I went ahead and let us update the data today but I can still see
what was my data before that update so there is more flexibility when you use lake house
because it is spark govern and the underlying data technology is more wealth versus the spark
so that is why you get more flexibility when you use lake house also it is one of the
misconception that everybody has that we cannot work on warehouse using the pi spark we
can work using pi spark the only thing is there are special snaps libraries which are there
which you have to use to work with that again there are certain limitation because where
house is still T seql govern so not all the operations are supported but you can work
with the pi spark in the now pi spark the technology because it is supported by spark or
let us say scala on it say spark are or spark sequel they have a huge extendable it means
you can run you know petabytes of data transformation everything using spark while sequel is a great
technology up to a volume of data so that is where if your system is really really large
you will go for lake house so if you are really huge on data you need to you know really
save bring the data whether I should be able to run fast but I want to do it quickly spark
helps you now on the other end sequel extend is own way is it is a same traditional sequel
way where it extend it goes you take the power and you know do does the job but I can tell
you up to you are up to few millions you will not be able to differentiate between the performance
we have implementations which have been done on lake houses and warehouses and both equally
performs good and on millions of rows sequel also never let you down it does really good
performance it shows that you know work done within time with lesser CU consumption all these
things makes it really amazing so I think it is more about choice of technology by the organization
and definitely if it is unstructured data and you need to do a lot of data science where lake house
definitely wins compared to warehouse the the thing the lake house architect
sorry require the different mindset so what if an organization already has a strong SQL and
traditional data warehouse team see see ultimately the two concepts is same now the here here is
understood even if I implement complete lake house so let's say bronze silver and gold all our
lake houses so my so-called one source of truth lake could be a lake house also which is so-called
warehouse in the functional terminology and it will all be warehouse if it is all structured data
I'm happy to be with warehouse but yes if I need to do a lot of data science and when I say data
science is not the use of AI or lm that's pretty much possible even using data is in celebrity but I
need to let's say do develop my own algorithm plus string classifications and everything that is where
you know that scale is typically more easily possible with spark and that is where probably I would
like my data underlying data to be in lake house compared to warehouse while I already said that you
know you can access your data on the warehouse using price part but definitely there is an intermediate
intervention which is required while saving so I would just if I have to do a lot of data science
stuff or I have a lot of unstructured data which I have to be I profoundly you go at the lake house
and read those the data lake fit into the story I think a little bit or what other
almanches of the data tables inside the fabric okay so when we initially so the concept of the lake
which we call data lake is basically you can store any kind of data any format of data
structured or unstructured someplace so lake house fits in that story so basically when we say
lake house the lake house the structured data is basically delta which is same as warehouse in
the in fabric world because it is also delta format in the back end but the unstructured data is
something which you are storing the file so we have the file part which we can store any kind of data
so the concept of one lake or the data lake so the concept of data lake is getting replicated
by lake house while warehouse represents the traditional sequel warehouses with those warehouses
kind of a properties as it properties and everything and yeah definitely acid properties are
something which are not supported on lake house which are supported on the house.
So I think now we understand where the data shows the lift but we need also get them there
what are the main ingestions in fabric.
So injation of the thing we can have a look at the data if it was now the second one
consists of the lake there are few sources for which the mirroring is directly available and for
others you have custom mirroring. Now the advantage of mirroring which by organization like is
basically it is self managed everywhere where we have a good change data capture mechanism mirroring
works very easy. Now Microsoft doesn't charge you for the compute as well as storage up to a limit
for bringing that time so let's say if on in fabric I have on F2 capacity I will get two terabyte of
free storage and pretty compute to bring that to terabyte. I am if I'm on F64,
our F64 is one of the very common capacities and the reason for that is one is there are very
high limitations on the data set size especially for power BA and the power BA viewer license are free
after F64 onwards. So if I am on F64 capacity the mirroring will allow me 64 terabyte of data to
be stored for free and to compute free for bringing that so that is my second choice or you can say
first choice if I don't come short course then comes my ability to bring in data using pipelines which
has copy activity and copy job. Copy job again is having very good CDC a co-options available and copy
activity also now both are meant for the purpose. So if I want to bring the data these are my choices
but yes the pipeline copy activity and copy job has around 40-50 connectors and that is where data
flow wins because data flow gen 2 has more than 200 plus connectors it is it has power query which
can also do the transformations. So data flow gen 2 would be your fallback option most of the
cases when you are not able to connect and all of them has generic connector like RSTPI and we
also have the ODBC generic connector or we have OLEDB all these things are also available web
connector etc. So our choice would be finally if we are falling back we use data flow gen 2 to bring
in the data it has maximum number of connectors and it has also fast copy connector now which will
make sure that it is as good as your copy activity in most of the cases definitely transformation
choices are there in case you are bringing from data flow gen 2 you can transform the data also
there. So now because we have this path and Python inside the system if your data is available
online right now it is not supported for on-premise sources you do have now various connectors available
at the notebook level previously also we were able to bring the data in but we have to mention all
the credentials in the notebook but for now we have connections which is basically supported
in a manner that you don't have to show your credentials in the notebook and that again creates one
more way to bring in your data now when you bring in the notebook the handling of the data is
mostly in your end it is basically how you bring how much you bring in everything while in all
other cases it is managed by the system so you have these many choices to bring in data.
And what would you say how important is power-carry knowledge for working effectively with
the data flow gen 2? See power-carry is something which is data flow gen 2 is having basically it is
same as our power-base power-carry or your our data flow gen 1's query and if we know power
query it would be really helpful because I can tell you those who are migrating from power-bava
word their first choices data flow gen 2 those who are coming from Azure word their first choices
typically the data pipelines and let's say PISPAR transformations and everything so we will get very
familiar things whenever we are coming from a different you know world to the fabric world.
fabric is new so somebody from the power-bava world will come somebody from the let's say
the PISPAR engine world will come somebody will come from different migration kind of a
pools which is very similar to pipeline you will come so orchestration everything it happens so
we will find some very similar tools here in Microsoft fabric to start with.
Okay awesome I think is there performance on main
tiny ability limits that we should know about in that of low gen 2?
Okay so this is a really easy so teach gen 2 is very good for a small to medium size of data
but when you come to a larger scale the transformation in PISPAR can excel much faster than the
in data flow gen 2 while Microsoft is continuously investing and consistently improving the
performance of data flow gen 2 when it is large volume of data but it becomes really easy
because Spark is PISPARC or Python is like an open language you can have tons of things to do
and manage it and that is why you know if it is a really large scale data transformation we
prefer PISPARC. I don't have a data flow gen 2 versus PISPARC but data flow gen 1 versus
PISPARC performance could be having 80% gain both in terms of compute as well as in terms of time
that is a proven fact but we have observed at the customer level. So data flow if you had a data
flow gen 1 if you completely converted into PISPARC because save a lot on the compute units as well
as the time both you are saving it is not one we are saving and that is where PISPARC could be a
really good choice for data transformation going forward in the future for large scale data.
That is interesting for someone coming from the PowerBRS-L part when they actually need PISPARC.
So the data transformation part is something which is basically instead of SQL can happen on
the PISPARC and when I say PISPARC I have this habit of calling PISPARC but we can use these Spark
notebooks which could be in Scala which could also be in Spark SQL. So if you are from SQL
where you will find Spark SQL very familiar to you. If you are from let's say power query world
you will see PISPARC almost has similar kind of functions which you have let's say if I want to
join they are also have I have table dot let's say join something and here also I have the data
frame dot join. So very similar functions I have here also to do one one on one operations with
compared to power query. So definitely if we are coming from let's say SQL plus power query world
we can start on fabrics with those but as the data goes is the things grows slowly we have to start
moving some of our stuff to PISPARC or maybe we can say Spark notebooks. Yeah yeah I think that's
interesting how do notebooks change the way data engineer teams work actually. See we have to
understand that the notebooks especially running by Spark is made for scale.
So because we can decide you know how much compute they should use they should use 10 node
small cluster or a step small cluster or large cluster all those things we can do and that
scale comes up with Spark and because of that you know you might do need to do heavy lifting
and that heavy lifting you want it to do and you also want it to save time not only the volume
of data increasing but you also want to reduce the time that is where easily can happen Spark and
Spark is meant for the entire illusion of Spark happens that okay I have the huge data how do I
want to process that data and get the time in time processing in quickly and that's where the
Spark was entered that you know okay distribute this entire jobs on multiple nodes and then do it
and then bring the results in and that's where the Spark notebooks are providing you here in
Microsoft fabric. So for scale and for you know getting the results faster the Spark always helps you
that's so sounds interesting is there any performance mistake begin us to with Spark?
The thing which we have to do is when you start as a beginner don't try to play around with
these Spark settings of fabric also secondly instead of dealing with it in the notebook level
please create those different Spark clusters settings at the workspace level or capacity level and
use it because I remember a case where what happens one of the team and we have allowed the setting
of the Spark cluster at the workspace level one of the team is created such a big cluster
and they were able to complete their job in one hour they were really happy but they end up
consuming the entire capacity for all other performance so those could be there and again
there is a little bit of there are less of control sometimes required in the fabric spark because
some of the things are already handled by fabric so you might not to explicitly go and start Spark
it automatically creates sessions for you so the small unlearning here that you know you already
already have a Spark session you only to explicitly started you can start using it and then there
are only selective parameters which it takes if you take care of these things I think it's pretty
smooth journey and even I think for a beginner who doesn't understand as a Spark works it is much
easier journey because you are not taking care of where is my root folder how should I start the Spark
I just know that I have a Spark session with all those capabilities and I can use it.
Awesome another topic I think the biggest buzzer words is the it's yeah I don't know the mid-medallion
architecture and I think it's a little bit of that marketing that we call it Bronx the Servant Gold
have have these vocabulary are modern how how how important are the the medallion architecture
for for fabric see for see architecture should depend on you know organization and what kind of
data we have see if war organization is having a very strong master data management done at the
source level and data cleaning was a real big focus probably you will have very clean data which
is coming into the grass layer itself then you daily need silver or then can you directly go to
silver plus gold layer that is the call you have to take because I firmly be the number of layer should
you entire dependent on the kind of data instead of you know I want to follow medallion architecture
because when we initially started the data engineering if you remember we used our staging area and
the warehouse the one source of truth we we were you know doing that entire as a part of transformation
maybe we were creating the intermediate table not explicitly calling it silver now because
this has been a well defined structure of medallion architecture we have bronze silver and gold
so if you have a very strong master data foundation which is laid outside probably just think about
do you really need one silver copy because you might have clean data but yes if the master data
foundation is not so great then we need to bring in bronze create a really strong clean layer for
the consumption of data at silver layer and then create the gold layer so and I see most of the
organization at some stage or other having this challenges of master data because it's not managed
so greatly so probably if you can say 80% of the organization will fall in a category when they have
to follow the medallion architecture of bronze silver gold yes there are organization who has
really strong on you know data quality and data master data management and everything
probably for them silver is not too much other than you know having a following an architecture
probably they could have bronze and silver gold together while doing the transformation putting
it into the gold yeah that's interesting so so where should the the the business logic
live so should it be the logic primarily exist in transformation in the warehouse in the
thematic model or some well now this is a testing now there are so this is a little bit of
the business logic which are little complex and the lying level we cannot keep them at the
semantic model level ideally speaking I would be happy to have all my major able or KPI's or
performance indicator at the semantic model level because it's easy it's dynamic I don't need to
reload the data for doing it so anything which I can keep at that level but the problem is if the
calculations are at lying level if I have millions of rows where I need to go and do one calculation
and then do it up it is not going to work out at the semantic model level and in such cases we always
prefer to have the calculation done at the gold layer means basically when we load our data from
silver to board we will apply those business logic so business logic is shared between the gold
layer creation and the semantic model our preference is to have it at the semantic model level because
they can change easily we can change formula we can create new matrices but definitely we don't
want to do role level calculations for that so we have to take a cautious call to make it at the
two levels that's how it moves to a side awesome and that's a little bit move closer to the
business layer where are the semantic models so important and what makes the good semantic model
okay see if you actually would ask this question we could be couple of your banks
or three years back for a semantic model is for consumption of power VA isn't it so we have a
semantic model we create power VA reports on that but within fabrics the things have changed
especially after data agent ontology and fabric app is all changing this playground and creating
the bigger role for semantic model so if you ask me if AI brings a challenge to let's say power VA
as a visualization where power VA is not facing challenge is the semantic model because
understand one thing our all business logic are going to decide the relationship and the key
majors and their definitions are known for semantic model and if I work on a semantic model I don't
need to tell the definition what is my net sense semantic model knows it what is my gross is what is
my revenue definition everything semantic model knows it and let's say if somebody need to do and
go and do some transformation on directly on my lake house and warehouse they don't know this
definition now we have a semantically lab library which can enable you now our data agents also
work on semantic model it means the data agent wanted to take the advantage of the definitions which
are not similarly the ontology can be created on our semantic model it which is going to take the
advantage of already existing knowledge by the semantic models the relationship the KPIs
now we have fabric apps which I got wrong recently so one of the things which has happened because
of this AI explosion over the last few years that we came to something which we are doing let's say
maybe a 20 years back that I can have my custom created report which whatever I want
and then from there in that place we move to the tools because it is heavy maintenance because
I can create my report let's say on Java technology by connecting to database and do everything
and have slice and dice and everything but it was a huge maintenance I can't do it and that's why I
knew tools like Power BI tab you to just do it with drag and drop but what AI has done over
the last couple of years is you know you just go type in some command and your react dashboard is
ready you say okay I need all these features it's already and that's where you know we were coming
back okay why should I have you know executive only looking at you know reports which is
fixed in the boundaries of what the tool can create and that is where you home we have seen a lot
of people were developing these AI apps and the challenge of all those AI apps is they understand the
warehouse or lake house better because they are having sequel and points and you the most the most
common technology to be used by anyone is sequel or as a sequel and because of that developing here
but our logics for lying in semantic model who will transform those major or those
capyra definitions into the sequel and there's again work and that's where Microsoft came up with
fabric app enabled by Rafe in framework now it can create these those reports those it's so it
used typescript and you can create those reports which all the kind of UI you need the custom UI
the kind of slicers you need the visual out of the box visuals and everything you on your semantic
model and that's what we call fabric and then it is again getting published inside your fabric works
pieces so same security same everything so now the role of the semantic model has greatly increased
in terms of relationship and the definition it owns but yes the consumption could be power be
a visualization could be data agent could be ontology ontology driven data agents or maybe fabric apps
but yes we have a layer which store which knows our business which understand our business well
and it can be used for the downstream by anything
and and what would you say how should teams approach measure of firsts calculated columns
okay so when we say calculated columns probably I would like them at the gold layer other than
these calculated columns which are now which is very interesting thing which has very recently
happened in the power be a world is contracts driven calculated columns now contracts given calculated
columns means they change their value because of user climate it also help us in you know
some time to hide some data basically if I wanted to do data masking it also helps in that
and sometime it helped in changing the column basically I want to show different column so if it is
if your column is not a context driven column it should happen at the gold layer all the majors
anything which is a major should be going to happen at the semantic model level context even calculated
columns should happen at the model level semantic model level and rest of the calculation which are
line level or really complex which is which can slow down at the runtime should be moved back to the
bouldier. If I worked with fabric and probably iDux was the I don't know the main topic and I see
Aids awesome in generating DAX code now did you think DAX stays important or makes tends to learn DAX
anymore see I think in the new world what is happening because AI is doing everything like I am when
I created a let's say referring dashboard it uses typescript I am not sure what typescript is
how it is written what is the structure of that and everything similarly it is happening with all
the technology like probably you can generate a pi spark notebook using co-pilot or even AI
that is to the DAX also but we have to understand when the things does not work and when the things
need to be optimized and that is where expertise comes in place now some of us who are there in the
industry for some time we had gained that expertise over a period of time now for the new learners
it is very easy to get a DAX from AI so do I really need to have an expertise I think in a longer run
at least for next few years you need to have that expertise on the subject I know the technology
is really fast and maybe after some time it is more on the logics and less on the technology so
probably the language would be you know using your logics and explaining your logic to the AI
by our prompts not let's say if C Java DAX of our query it's just the thing which internally getting
created and getting debugged also by AI only but yes definitely knowing something in depth how it
works it's always really important because in that case you can go and fix something which is
actually not getting fixed by AI or which is getting not optimized by AI but as this intervention
increase probably up to me loving you know how the algorithm works how the logic works at different
places is more important than you know the language is what we are working use in future.
And I think another topic a lot of people aware of actually it's it's security I have all my
data now and well like so how should we start with security in fabric architecture?
See in terms of the security start from your workspace is from where you can start controlling
by rules whether you are a viewer you are a member you are a contributor or you are an
advent that's the first step now in the fabric what typically we are trying not to have the workspace
level access we try to have item level access like in a lake house or warehouse I can have a read data
read and write and all those permissions then if further goes down I can go to the item level now item
level permissions typically we will try to control using one-leg security wherever possible
and if we are not putting one-leg security we are going to have semantic model level security
a one-leg security has came I think around one and a half to years back and it is still evolving at some
places it may not provide you that option and that is where we go and put the security
semantic model level but one-leg security is something which we are going to use
and that is where we can secure the data means we can do both the object level security
OLS and RLS using one-leg security so what we are going to do here is basically where but possible
in terms of OLS and RLS if possible we are going to use one-leg security if we are unable to control
at one level one-leg level or we don't want to contain at one-leg level basically
lake house of warehouse we are going to do it at the semantic model level and definitely object
level security needs a set level security basically or report so do your warehouse on
lake house we can control them at the works is either the workspace level or individual item level
so from workspace to item to inside the item with one-leg security
awesome so the how how do is the role level security works with the semantic level correct
or what's the best practice see the role of a security basically now it can be done at two level
one is basically at the one-leg security which is can again you can go and define okay I for
this is the filter for that so in one-leg you decide the filter and then you create a role and
that can assign to the people so that filter will get applied so you will let's say only get north
I create a filter it is north and it will get applied at the end all the endpoints SQL endpoints the
the semantic model endpoint or even at the file endpoint now when I come to the semantic model the
semantic model security of power be a remain same so it is the same way we create a role and inside
the role we again define a dynamic security using a table where we have the email addresses using
user principle a model static security by like saying reason equal to north and then we once the
semantic model is published we go in the security model layer and add users or security groups
preferably we should use security groups and security groups should be created like you know
this is a viewer security group or this is a financial viewer security group this is says viewer security
group or says contributor security groups we should create security groups and assign those security
groups to the security layer or to the report or to the one-leg security that is always a best
practice so so we have the workspace permissions we have the data permissions and we have the semantic
model security but how how to gather when this when when an organization has hundreds of thousands of
users and with different requirements yeah see I think the why we have these multiple layers and
but all these are all you know really smooth and integrated so the reason is let's say if I want
to have few people to be admin and having access to everything I don't have to you know go and
do multiple permission I can deal with workspace level permission and they can access everything
in the workspace they want but there are certain set of user I don't want to give them you know
access to everything let's say we happened with us also there are report user who are report builder
as well as the report viewer but we don't want to expose them the let's say warehouse is or
lake houses so we control the security at the semantic model level we give them access and we make
sure that it happens in such a manner that they don't have to see other objects so all those
possibilities are covered and that is why you have this two three years now we can do some AI
automation for that to make sure it happens like basically the security when we create the security
group on a larger organization these are also not manually created you have let's say a service
request which I go to raise and you will get added to a security group the same request may add you
to let's say a RLS table also where you get an access to a region to a city or to a product also
and then those RLS tables are used in your semantic model to govern the security and similarly
the assignment also happens let's say I asked for sales viewer role so sales viewer role with let's say
city as new new york so the service request will go add me to the sales viewer security group it will
go mid go and add me to the new york it will also go and add me to the let's say workspace or to the
reports which I can see as the sales viewer so all these automations can help us out. And I have the
fabric admin center I also have I don't know it's power be either same name but I also have the
option in power be I and and I also can use use entra how show I had all these tools or what's
the best practice here. So, so we have it that level so basically entra all the security where
where basically your all log is getting created and security groups are also getting created at the
Azure level so that part happens there after that fabric and power be I is the workspace level
security same way to apply suppose that we give the assignment of the objects on the fabric
level to the security groups permissions are also decided security or sometime you will let's say
try let's say fabric app came is a new feature now I don't want everybody to go ahead and create those
reports so I will give permission to a set of users which are part of a security group to create
those reports same way I'll give the finance user access to let's say as a contributor or a member
or admin to you know build the content and give it and viewers of those let's say finance
report security group of finance viewers I'll add deadness the viewers and also add them to security
so this is how it happens so definitely entrap part of the Azure creation of your emails and everything
not going away so that will happen security group creation should also happen from there and
rest of the security will be managed at the fabric level. Yeah I think we have a lot from from
the engineering part but yeah ultimately the most business users don't care about spark data
levels or like house they care about answering questions how do we bridge these these get
so the fabric has lots of options basically if you ask me future could be talking to your data
agents right now we also see a lot of people started using data agents and in that data agent
ontology driven data agent is something really important because what happens when you create ontology
on top of your data it do understand data much better it goes look at the data and try to have the
data part match and it can provide better answer and it can sometimes work even better on the
snowflake models or the models where you know the typical star schema kind of stuff is not there
it also works better there so ontology driven data agents would be our first choice in the future
to communicate with our data and understand our data then we will have the reports definitely
I will not wanted to go and communicate and always ask questions I would definitely want some
numbers in front of me every day morning now those everyday morning numbers will either come from
a fabric app report or power be a report there are very high chances in the future that your data
agents might create a report for you which is basically could be a fabric app and a power be app
after you ask a set of question okay this is how I perform these are the KPI I look at it
this is the information I need Monday morning this is what I need to do morning so data agent probably
could go ahead and create finally based on your communication a report which works the VR way
and not you know the report created by a developer which gives you okay same set of reports for all
four executive same set of reports for you know thousands of users so probably data agent and the
data agent created reports would be something which we will be frontending in future yeah I think a
lot of power be I and business users like to have a self service be I it's a little bit distressed
us from our most organization what do successful service ourselves levels I look like
okay so now see understand we are talking about this town self service but when we use that
creating those reports the movement we go into the complex calculation we realize we need to go
back to the developer even in the power be in world where you know you have the self service knowledge
of decks was required but in this new world of AI driven you don't need that even if a major is not
there your AI agent is going to create so this is real self sir your model doesn't need to have all
the formula your model doesn't need to have all the logics it will on the flight created now definitely
realize okay if it creates on the flight takes lot of time it's better to improve my model to answer
some of those questions maybe better to have a major for that so I think two self surveys what we
are calling is on the plate using data agents and the AI driven development so where you know
users will just be so I'll tell you I was on the weekend I have created entire fabric from
ingestion of data to transformation creating a power be a model creating a report just by talking
to the GitHub co-pilot I have not went ahead in any of the UI other than the testing so we're reaching
to that level now understand what would be the future that you define your requirement well to
define your requirement well you have tools also which will define your requirement well then they
will break it and then you will only go ahead and say okay go ahead and create the tickets on my
DevOps and then you will come to GitHub co-pilot and say okay execute all these tickets on my one
and this getting created and you do so the knowledge of technology only needed for the debugging skills
for what is not working how to make it work but it's going to be a really self-serve environment like
you know in our generation we used to learn computers but if you look at the next generation they
they know computer by you know they they know is as a skill be it's just like a language skill for
that for future the AI would be cut that kind of skill where everybody knows how to you know build
those complex some of those systems using you know AI so that will become a primary skill for the
next set of people who are you know now learning those things so probably for them it would be just
okay I go or I type prompt I'll have my stuff ready for me awesome yeah thank you are we really
little bit running out of time so I have in every session of rapid fire rounds so I ask a short
question you give a short answer is it though are you ready out yeah okay let's also warehouse
lecels escuelor pi spark pi spark dataflow gen 2 or notebook not book back so escuel
decks most underrated fabric feature fabric apps the biggest power be I'm modeling mistake
many many relationships what is when such an gallia called you and say hey you get all the money
resources to build your dream feature what will you build probably I would want to improve
upon on data agents so just ask them okay this is my requirement have a dashboard ready every day
it should change based on some requirements I don't want to see same dashboard every day one
fabric skill every power be I professional show to low probably you should be knowing power query
index if nothing else is fabric apps in you power be I yes awesome so and we should I invite
next at what questions do I ask I think you can this in Rajin who's again Microsoft fabric MVP
and scheme more questions about ontology and fabric app I think he's one lot of working on
real-time analytics actually he's working a lot on real-time analytics so maybe you have a session
on fabric real-time with Rajindra the data platform MVP and if an organization wants to start with
fabric tomorrow what should it's first project be I think any data project where you think I need
some kind of analysis it just in a starting point it could be just you know knowing you know
how my employee performance is there or how my sales is ongoing if you think you need data analysis
fabric is there for you awesome so yeah I'm a thank you for for joining me and talking taking
us through the computer the Microsoft fabric analytics journey what I think in this conversion
makes clear is that fabric isn't simply another in real analytics product real opportunity comes
from connecting data in guest gen engineering lake house where our architectures and
semantic model security power bar into one core and platform and the technology alone isn't
enough the architecture ultimately has has to deliver trust information that people can use
to make better decisions so for all the people to listen to the podcast and interested to connect
and see how it works you can look at the show notes and you find this profile there and yeah
thank you again I'm it for for being a part of the show thank you thank you for inviting me