hi i'm adrian cooper i'm chief product officer for um torrentia uh hello again if you saw my um earlier talk um this in this talk which is called when data becomes the asset i'm talking talking about data management in the age of agile integration give me a second so i want to drill into um into sort of below the systems level to the data or the metadata itself and the importance of data management data management for for digital transformation so you could say that the great digital innovators have recognized the necessity uh to connect data and systems not just for the front end or customer facing elements but also to join these together with the middle or back office systems they understand this this need to align online and offline so it's not really digital transformation in my view if we're only changing one part of it and organizations that only consider changing the public facing elements are likely to struggle to deliver value over time and i think certainly in my own experience [Music] you know a lot of museums have actually largely focused on that customer facing or public uh uh public uh or audience driven part um digital teams and services have been established but mostly you know ultimately another silo and the rest of the of the organization has been business as usual of course at least until the the pandemic um so i think collectively i mean there are obviously great examples uh and but we need to work towards what's what's often referred to as the connected enterprise or or in this case the connected museum in the last talk i introduced this concept of services as a way to help organizations to change their mindset um services as a macro level uh as a way to kickstart some of the thinking about key problems and develop what i call connected processes um so a service that doesn't respect or follow you know an internal organization and so if you're thinking about developing developing something there's this then an increased need for organizations to find smarter and easier ways to combine the data that's traditionally held in in these separate systems internally modern cloud native applications are generally designed with this microservices approach and microservices are a way to enable organizations to orchestrate and bring together and process and deliver efficient experiences and transform workflows so a microsoft sorry a micro services approach means that although data and information sources become more distributed across the organization the data itself is brought together or connected as needed to support the task the workflow um or the overall the overall service and this is what we mean by um agile integration so we can we can we can orchestrate and bring together data from a range of different microservices which each do something specific in order to present to the user an interface that allows them to do what they need to do for the set of tasks that's part of their service nothing more nothing less so if we recognize this need to make the alignment between the internal and the external uh if you like the workplace and the digital space my data is connected merged or enhanced through a range of new digital services what does this mean to the way we think about and manage our data or metadata so really so really what is the challenge for for data in that context so we could say that that data is the currency of digital transformation because you can't do digital digital transformation without transforming your approach to that to the data to metadata so as we start to plan any migration to new cloud native platforms with smart connections we need to think about how we can ensure that the data becomes or remains a key asset so it won't be smart unless you can unless you can connect it so for anything to be an asset it needs to have value and and therefore data needs to have value to be treated um and seen as an and be an asset and so i think there are some key challenges or key things to think about in terms of what as a profile we might think about uh you know in in terms of that data in this decentralized digital ecosystem so we need to be able to share data clearly we need to be able to trust it uh it needs to have meaning and and context um so obviously that's where taxonomies and and and controls uh come in um it needs to be consistent comprehensive um and accurate it needs to be accessible um controlled uh secure uh if you think about right rights and licensing and all of those things we only only the people that need to have access uh we'll get we'll get access and of course last but not least it needs to be machine readable so we need to we need to have uh computers to be able to look at data interpret it know what it means and to be able to do something logical with it to process it in order to to make these make these connections so it's clear that we need a more flexible and adaptive data management strategy in order to deliver this agile and trusted data which then begs the questions you know who's who's responsible or accountable how do we improve or maintain the quality and and how do we how do we make these smart uh connections or automate them and in term in terms of what we call the the data supply chain the joining together of the data as needed to support the different processes that make up uh an overall service so i think the ability to manage data as an asset across the whole organization depends on a number of things technology organizational culture governance and and sort of staff employee accountability you know everyone's got to do their bit and i think um deliberately in this slide i've made the organizational culture and and governance circles bigger than the technology and uh because i think that you know whilst a lot of other people assume that by upgrading their technology then the data is going to take care of itself sadly that's not that's not the case and and really it's about organizational culture and governance and thinking about things across the whole organization that's really going to going to make the difference of course technology uh will and can provide provide some help in terms of of cleaning up and sorting data to you know assuming that there are that there are rules to follow um you know data transformation services can be uh established and scripts written to try to clean up poor data and you can certainly use ai and machine learning to to to process um you know data and and and try and um you know using using machine learning and ocr and things like that to actually improve the the tags and and other metadata about about the about the assets um and and we can also use machine learning to to attempt to create links to uh linked open data sources um and you know wiki data and things like that to it to improve the uh the accountability and and add to the weight of of that content so where are we where are we now i think there are kind of two uh two types of barriers there's organizational barriers and there's tech there's technical barriers to this in terms of organizational barriers well really museums aren't any different to most other sectors and his data has been managed in roughly the same way in most most organizations across all sectors for this in the same way for the last 20 or 30 years and that tends to be uh that it's managed kind of independently within it within a departmental focus uh because in turn that that relates to the kind of systems that are that are put in place and so if you imagine saying in a museum context uh in typically museum data is designed and managed in these large scale business area specific applications such as a collections management system but but even within the same sort of core area as we say collections there is you know a lot of cases and certainly my own experience as a consultant a strong tendency towards separate management and control even within that that area so for example management of objects archives books all part of the overall collection and and that they're often handled in in independent uh systems and i know there are those who say well there are reasons why but i mean that's that's more historic i think than than relevant for for for the future but but having independent metadata structures having separate controlled vocabularies or taxonomies uh to to to to describe the data uh leads to issues where it's hard to harmonize any of that data um if if you you know right at the beginning i said of course services don't respect organizational culture so users don't need to know that your data is held as separately in in in a library collections or an archive system they don't care they just want to find uh what they're looking for by topic subject etc so we need to find ways in order to bring this data um together to aggregate it and to harmonize it and then on top of the kind of the thesauri problem as it were with multiple different separate systems we then got uh functions which which are perhaps duplicated across the organization so something like like loans might be managed by you know a number of systems independently in an organization or it could be that loans is only managed in one system and in order for say an archive object to be loaned it has to actually be added to the collection as a as a temporary record in order to then have a loan transaction uh carried out which so so these things you know are all examples that i've i witnessed as as a consultant and um this siloed approach as it were is perpetuated by some of course very basic and natural tendencies for teams to be possessive about their data and i think uh that's quite understandable and and there can be a reluctance to share data or to adapt cataloging practices to support the organizational wide needs and and it but in a lot of cases maybe the organization itself isn't isn't strong enough or clear enough to actually articulate what it wants in order for these departmental systems to come together which is then part of this you know service orientated thinking let's think more about the outcomes that we want rather than thinking in terms of systems and features and functions because that just really doesn't help going forward so um poor integrations and other technical limitations of course help to enforce that kind of departmental outlook so they be it becomes symbiotic organizational barriers uh are to the connected uh museum enterprise are in turn perpetuated by this this nature of legacy systems um some things legacy not not not because it's old um it um it it can be it can be legacy for for a number of reasons um but um you know they tend to offer this single interface single database you know single set of of of modules um and and often have data entry screens which are not terribly user friendly or difficult to use there's a lack of flexibility and so this leads to what's sometimes called field hijacking where users you know out of out of frustration or you know uh just just put data where they want to put it rather than necessary in the place that that is designated for it um so that incoming in combination perhaps with poor or lack of data validation tools uh you know if you're allowed to put a date in the text field or vice versa then then of course uh results in the data quality being uh significantly reduced that combined with with terminology control which may not be as effective as it could be all points towards lower quality and dirty data so you know how do we properly identify you know people places objects materials subjects events etc if the values that we're entering are are simply strings of characters that have no in intrinsic meaning that can be interpreted by by by other systems and where the you know the data's in in fields that means something different to what they were intended and that only confuses the situation so data managed through a sort of business specific application may well also require some form of interpretation or logic layer to have me to have meaning if the data is exported outside the system it can't necessarily be uh understood um and and and as you know it's been talked about a lot today you know if there are no apis that makes data hard to integrate and that results in of course time-consuming projects to build inflexible or one-off integrations between specific systems as as needed and there's certainly been a lot of debate about you know what's the right level of integration between the cms and the dems but from my perspective uh these point-to-point integrations are hard to maintain and any change to either point will result in the integration breaking and so we should be thinking more in terms of of cert of a service oriented view rather than this system to system type view because that that really is is an older way of thinking so if if given all of that we we look to compare kind of legacy data with our data data value profile that i that i put up earlier we can see that the legacy data uh comes with a number of number of issues often so you know it's it's inflexible it's it's possibly hard to share it doesn't necessarily have meaning um it could well be inconsistent because data isn't in the in the right fields or it's not being properly validated so which means that we can't trust it completely which means that its value is lower and and you know possibly you know again more importantly it isn't machine readable because we can't use any logic to interpret what it what it means so how can we move towards uh data being being seen as an asset and i think as we said if you know organizations have relied on this uh business or departmental uh areas creating and managing their own applications and data models and tax on taxonomies uh historically but i think now there's a need to change the approach to data management so that it works more effectively across the whole organization and that's that's going to be a big change for some and so in the connected organization data is the most significant and tangible asset it needs to become the fuel uh that that powers multiple use cases initiatives not not just designed for one particular use case or project or thing uh new digital services will only be as strong as the underlying data that that fuels fuels them and we need data to be structured and machine readable in order to have to have use and the value will come from the ability to collect and correlate data from different systems without need for this manual interpretation logic and automation so data that's locked inside systems that has no intelligent intelligible structure or meaning can't be connected um into a digital supply chain which is then where we have to fall back on manual integrations so as organizations shift from these large monolithic projects to more agile and service based approaches the concept of an organizational data catalog will become more and more important it's going to be important to build and maintain this shared data catalog in order to understand how key data types key entities gathered across the range of systems can be connected to support the overall data supply chain and and and support the needs of specific uh services so you can see here some examples of the kinds of entities i'm thinking about that are currently held across uh different systems whether that's objects or digital assets or products or locations or customers or rights you know you know again conscious that rights are handled separately in digital asset managements and collect management and collection management systems whereas the reality is that they need to be brought together need to be properly harmonized in in a in a museum context which makes you know currently makes workflow uh difficult for those who are trying to do uh do that role so we need to think about these key entities and look where this where this data is stored agree how we're going to you know share share on thing you know on things like taxonomies across the organization in order to be able to orchestrate the data in a way that it can come together using uh you know machines and services to deliver the kinds of services that you know uses uh and uh expect so in summary i i think there are you know there are a few a few key thing things to think about um and and so first of all is this rethinking the data supply chain making sure that that's done at an organizational level rather than necessarily a departmental or departmental by department level and i think the service based thinking where you're thinking about about specific problems and outcomes uh will will help with that there's certainly a need to increase the data modularity and and you know to break things down and move towards this micro services and focus on on the development of of apis uh to to uh to to support this this bringing together of the data um and and i need to think more as an organization about the data the data governments data is an asset that everyone has to manage well it's not it's not you know somebody over there's department uh you know to think to look after um you know everybody everybody has to have responsibility and of course we can use more modern systems to to improve that that and keep that keep that maintained as we as we go along and last but not least is the harmonization of metadata metadata models this is absolutely critical when we start to think about shared taxonomies not only between you know within an organization but actually between organ between organizations whilst you know we've been developing apis some of the bigger organizations have managed to do that there's certainly no uh um there's no standard around the way that they've done that so there are already issues in trying to draw draw data from different places together if they're if that if that's not uh it's self common so within the institution and then institution to institution as we start to collaborate more on on these on these projects um so so i hope that's given you uh some things to think about it's not always the most exciting part but i think if we don't get the data element right uh then a lot of the other things can't uh can't follow on um from that so thank you for for listening and good night