We've touched on what's involved on cultivating data-enabled employees in this blog now and again over the past few years, and in general what we refer to is making sure that people and teams have the skills and resources they need in order to make use of data in their work, whatever that entails. In some places this needs to be a nuanced understanding, since not all employees need the same data skills or to use the same tools, but lots of employees all across an organization can benefit from data in some capacity.
Going back a few more years, we observed a certain amount of pushback to the idea of a data-driven organization. A legitimate response might have been that really what *drives* an organization is mission, or profitability, or a founder's vision, whatever that might be, and really data is one of many tools or resources that can be deployed in the service of mission or vision. Of course, some of that resistance was simply an aversion to, or uncertainty about, using statistical or quantitative analysis to assess business operations, which to a large extent is what was meant by being data-driven. No doubt many of our readers experienced the first response trotted out largely to cover the second.
We always preferred the much less catchy "data-informed decision making," since it incorporated "being data-driven" without naively assuming that the data would speak for itself, it still necessitated some level of data enablement in the workforce, and it clarified expectations around how data and analytics would actually be used.
Using data to weigh pros and cons, or to review the effect of previous decisions, or even to make reasonable projections, is pretty standard today, which often goes unremarked but is, in its way, quite remarkable. A quarter century ago for sure, and in some cases probably as recently as a decade ago, it was far from uncommon to hear leaders express active opposition to the idea that data could be used to help make better decisions. (We'd entertain an argument that "decision by hunch," as one of our coworkers memorably called it, still relied on a certain amount of data, if often an incomplete, if not altogether flawed or biased, understanding of the relevant data. That's a topic for another day.)
Now, what it means to apply data to the decision-making process varies by organization, by the decision in question, by the data available for application, and many other factors. The situation in which there are two possible choices, and those choices are a 50/50 proposition, and the evidence from data overwhelmingly favors one choice over the other, strikes us as vanishingly rare, although of course not impossible. Much more likely, we suspect, are situations where "the data" can be used to provide further confirmation of what is generally believed, or to introduce disconfirming evidence, requiring further deliberations.
Ultimately, for organizations that want to be rigorous in their application of data to managerial processes, a few things have to be in place. Among those things are the will to be data-driven or data-informed, the ability to identify relevant data and understand it in context, the tools and skills to organize and analyze that data, and of course a reliable set of relevant data to work with.
You might call the discipline that brings all those things together something like "data management," which would encompass all kinds of interactions with data, from its initial collection all the way through to its application in decision-making situations. And you might call the framework that scaffolds and surrounds that discipline something like "data governance," don't you think?
So: we want to make better decisions. When we say a better decision, we mean, among other things, a decision that has been made after reviewing an appropriate amount of relevant evidence, as well as thinking the data needs that will be occasioned by the outcome of the decision (this part often gets short shrift but without it you're not really using data throughout the decision cycle, are you). What's an appropriate amount of data to apply? How do we determine which data are relevant or pertinent? What assures us that the relevant data have been fully understood and properly analyzed?
We are speculating backwards a bit here, we realize, but it seems to us that much of the discomfort with or resistance to becoming data-enabled as an organization stems from the fact that too often these questions have been answered with "solutions" like faster BI tools, bigger data warehouses, more data scientists on the payroll, and so on. All those tools and resources are great to have, if you can use them, but they don't now and didn't then get to the heart of the issue.
We suspect there are many organizations that invested in upgraded technology, believing that investment would in itself spur the growth and adoption of data-informed decision making. What we normally see after investments like these is that more data is available, more quickly, in a greater variety of formats. For some users, this is like manna from heaven. Finally, I can get all the data I'm looking for! At last, my reports come back in a few seconds, rather than minutes or hours. Eureka--the BI tool generates the graph or chart, and I don't have to port the data from one tool to another to create visualizations my self!
For a lot of users, more data mainly means more delays and more confusion. Sure, I wanted more data, but now I have to look at three different dashboards in two different locations to see it? Yes, the report I rely on runs faster, but now all the data elements have different names, and it's just different enough that I can't properly compare it to the older version. It's great that we put all this additional data in the lake, but now I'm swamped with access requests from users who just found out about it. Are we really saving any time? Are we any closer to making better decisions? How could we even tell?
Let's go back to our earlier set of questions. Which data is relevant, and how much do we need to look at for it to be meaningful? Which data is missing, and who knows why, or how to resolve that issue? For any given decision to be successful, what is the range of outcomes, and how we will measure them? Any particular individual probably does not know the answer to all these questions, or even where to take them in order to learn more. But, across the organization, there is deep and distributed knowledge about data, operations, history, and their interactions. The challenge accessing this knowledge? Too much is tacit.
Strictly speaking, tacit knowledge has generally been understood as knowledge that can't really be shared by writing or speaking, but rather it's acquired by observing, practicing, etc. Sometimes it's described as implicit rather than explicit knowledge. At the organizational level, people often call the knowledge acquired by group membership, by working with colleagues, reviewing historical artifacts, etc., as institutional knowledge. But, since a huge challenge with institutional knowledge is that for key aspects of it, only small parts of an institution are even aware of it, much less understand, the phrase is almost self-contradictory.
Whatever you want to call it, this kind of distributed, informal, and blinkered knowledge is a huge, and growing, source of technical debt. We take a data-first view of a lot of computing and information technology, and in our observation tacit knowledge is the secret sauce that keeps many organizations running. It's how people know, for example, to throw out a subset of data that was collected incorrectly, or that a certain spreadsheet needs to be refreshed at a certain time and pivoted a certain way before running the macro that generates the report. Tacit knowledge is also how people know to take certain questions to Jane in finance or Julie in human resources, even if the question is about some other aspect of the business. Maybe Pete can't write a perl script, but someone left behind a text file with instructions how to maintain and update the script that's been handling a key integration. Somewhere, if you're lucky, deep in a network share, there's a document with end-to-end mapping of the data in your reporting data store.
Our data intelligence, data catalog, and data quality solution, the Data Cookbook, and our data governance consulting practice, are both predicated in some ways on recognizing, capturing, formalizing, and sharing tacit knowledge. Longtime employees have developed substantial subject matter expertise, and know their way around wonky workarounds. Data stewards know why a given piece or set of data was collected, where it was stored, which regulations govern its use, and how it affects their work; the accumulated knowledge of data stewards almost certainly is enough to tell organizations which data is relevant, how much is required for a given decision or task, and what measurements of subsequent outcomes will be legitimate and meaningful.
Once you start to bring your organization's disparate data knowledge, lore, and practices into a repository, you'll notice duplications, conflicts, and gaps. (There's a good chance, by the way, that longtime data stewards, for example, got together in the past to share some tacit knowledge and solve a problem. We encourage you to seek out and publicize those examples--people can address data issues without always taking them to the top, and usually more quickly and quietly.) We can think of worse ways to describe data governance than eliminating duplicate work and data sets, resolving or setting aside data conflicts, and filling in data gaps. One added benefit of this understanding is that it includes everyone, not just "data people," and it orients users towards practical, visible, measurable efforts.
Will data governance, governance that suffuses every data operation and underpins every data interaction, in and of itself lead to better decisions? That seems like a lot to claim for, or to load onto the back of, data governance. But better data governance is almost certainly a precondition for data enablement. And what does a data-enabled organization look like? We think they seek to harness the data expertise of its employees and users, to cultivate the formalization and sharing of distributed institutional knowledge, and to foster agreement about the meaning, location, use, and provenance of data. Organizations like that will already bring data to bear at many points throughout the decision-making process.
We hope you found this blog post useful. Also check our our data governance spotlight resources located at https://www.datacookbook.com/spotlights. IData has a solution, the Data Cookbook, that can aid the employees and the organization in its data governance, data intelligence, data catalog, data stewardship and data quality initiatives. IData also has experts that can assist with data governance, reporting, integration and other technology services on an as needed basis. Feel free to contact us and let us know how we can assist.
Image Credit: StockSnap_UTEZRDTKPP_SeniorEmployee_TacitKnowledge_BP #1332