As always: you do need to secure your data, and users do need to work with it in ways that are permissible and appropriate, and there's nothing wrong with consequences for users or units that won't follow reasonable rules. After all, if your organization mishandles data in violation of a law or contract, there will probably be penalties of some kind, right?
Now, we hope we haven't introduced this post with weasel words. But it's legitimate to point out that "secure your data" may have a legal meaning, it may have a meaning that derives from voluntary participation in a contract, compact, consortium or similar arrangement, it may have an organizational meaning that is based in policy or norms or some combination thereof. Deciding what is permissible and appropriate use of data might be guided by legal requirement, by ethical principle, by specified job function, or any number of other guidelines and directives. And determining whether a rule is reasonable, and whether it's really a rule or more of a regulation or a standard procedure or even a recommendation, can involve differences of opinion and even some level of "I know it when I see it."
Ultimately we think it's a fantasy, and not a very productive one to indulge, that any organization can come up with the right combination of policies and regulations to provide immediate and clear answers to every single question about data access or usage. That doesn't have to mean you just throw up your hands and do nothing! In our view it speaks even more strongly to the value of a pragmatic approach to data governance that focuses on enabling safe and appropriate usage, that seeks to engage users at multiple points in the data life cycle, and that sees data issues when they arise as opportunities to increase understanding and to make marginal, incremental improvements.
We can almost hear you saying, "That's all well and good, but how do we actually govern data pragmatically, and what does user engagement really look like?" Pragmatism in a business context could include many aspects, but we might suggest focusing on being solution-oriented. First, you might make a list of all the data issues you're facing that improved or consistent governance practices might solve or at least ameliorate. Second, you might try to rank the issues by importance, or size, or visibility, whichever metric or metrics seem most appropriate. (This is probably where soliciting user feedback would be critical. Knowing which data issues really energize your users seems like a good yardstick for prioritizing.) Third, for the top X number of issues, you could name and describe which practices would be helpful, and what it would entail to implement them. Fourth, you might estimate the time, expense, drain on resources, and some form of ROI for your issues and their fixes.
At this point, it might be obvious that certain options are too expensive, or too time-intensive, or would draw attention/resources from other critical priorities, or that the payoff is too small. Much of the time, however, we expect that it's not immediately obvious where to start, or how deep to dig in. So another component to behaving pragmatically might involve being ready to act on opportunities when they present themselves.
One problem our clients nearly always face is some form of data sprawl. Units and departments are collecting, housing, and in some cases generating analytics using applications and products that no one else uses, and over which there is little if any central oversight. On the BI end of things, multiple reporting and analytics tools are in play, generating dozens if not hundreds or thousands of aging, imprecise, poorly documented data products. Integrations abound, crossing departmental boundaries, some of them fully automated, others manually executed on demand, utilizing multiple tools and techniques.
This isn't a new problem. Data has always been stored in multiple places, often somewhat haphazardly, and pipelines for moving, copying, or archiving data are well known for being thrown together at the last minute and used well beyond their original sell-by date. Someone in every organization has probably been sounding the alarm about this situation on and off for years. But maybe now this setup is impacting end users in new ways, or it's causing critical business projects to be delayed or avoided. Now key stakeholders are talking about problems with integrations, or not having access to key data because it's stored in a system they can't find. Maybe there's too much data in the lake, or too many data sources to track, much less verify.
For a long time, no one seemed to care about data sprawl. Now, it's something everyone's talking about. (In fairness, they might not be calling it data sprawl. Let's be realistic--not everything is going to fall exactly into place, even in a blog post.) And you can't solve sprawl without understanding its scope, right? Maybe you've long wanted to embark on a project to document all of the data systems, all of the integrations, all of the reports and dashboards and analyses, etc., under your corporate roof, so to speak. Suddenly, this work might meet users where they are, whereas in the past they would have had no interest, offered limited cooperation, and shown minimal enthusiasm.
We'd still say be pragmatic in your approach. Just because key users are engaged today doesn't mean that engagement will last eternally. You've still got to ask questions. Who's going to lead this effort, and what amount of labor will be involved to complete it? Who's going to benefit from this work, and how much benefit will they realize? Should there be a deadline for completion? If so, when? and why?
Start with the benefits or use cases--after all, if you can't name those, then any initiative, no matter how much you try to sell it, is already dead in the water. Many users would probably benefit from greater awareness of what data is being collected, who's responsible for it, how to gain access to additional data sets that could support their work, and so on. One office might be looking to license software that is already in use in the organization, or whose functionality is already available. In this case, the whole company could avoid spending excess money, or before spending it a more informed decision could be made. After all, maybe the new system would be a vast improvement over what's already in place, in which case it could be money well spent. A smaller number of users might recognize that some data, which is already in the organization's possession, would be a good candidate for pipelining into the data warehouse. (People who have to respond to audits, or worry about compliance failures, would see the benefit as well, although as we noted above this is unlikely to be a benefit that moves the needle among the rank and file.)
What kind of work will be involved in conducting this survey? Will it require (virtual) shoe leather, meeting with department representatives to ask for details or to follow up when a request for information goes unanswered? Are there machines or applications that could perform some of this work? Even if there are, somebody's time and effort are going to be drained for this work, which means some other work will have to put off, or perhaps ignored altogether. While a formal accounting might not be necessary, the sponsor of this work would do well to assess the actual cost of labor and potentially other tools, and to estimate the opportunity cost of doing this work instead of something else.
How comprehensive an inventory is really needed? There could well be units or offices using a system that isn't a duplicate, doesn't capture or risk exposing sensitive data, and requires essentially no IT support. Do those need to be included? Whether there's a deadline, and whether it's firm or soft, the scope of the work should be considered as well.
A pragmatic approach might say, the benefits of this data systems inventory are worth the expense, but the work cannot drag on forever, and so we are going to establish parameters from the beginning. Maybe we break things into tiers: in the first tier we include the systems that directly integrate with our ERP, and those integrations that we are internally responsible for; in the second tier we include independent systems that we know store sensitive PII, or have built-in analytics tools that are used to provide dashboards and data to the C-suite; the third tier is everything else. Tier 1 requires X amount of documentation, Tier 2 requires Y amount of documentation (which is less than X), Tier 2 requires Z amount (which is less than Y and might even include no documentation at all). We have six months to complete this work, and if we realize after four months that we're not going to finish Tier 1 by that deadline, then we reassess our previous assumptions. If we're in the middle of Tier 2 or Tier 3 by that point, we're stopping, and this effort goes back into the pool of potential data governance initiatives.
Pragmatism in this sense involves recognizing that not only can you not do all the things you might like to, even within the highest-priority areas you might not be able to do everything you'd like! It might also involve admitting that major data governance initiatives are going to be few and far between. Taking advantage of situations such as we describe above, where improved governance is the obvious solution, is a no-brainer. But there are undoubtedly many more occasions where a narrow, shallower action path could be effective.
Our flagship product, the Data Cookbook, contains numerous data intelligence features and methods for practicing and supporting data governance. We'd love it if all our clients used all its features! But even when, perhaps especially when, you're rolling out a tool like this one, we think you'll be more successful if you're clear about which problems you're addressing, how much effort you're willing to spend, whom you're looking to engage, and where you'll get the most return on investment (ROI).
We hope you found this blog post useful. Also check our data governance spotlight resources located at https://www.datacookbook.com/spotlights. IData has a solution, the Data Cookbook, that can aid the employees and the organization in their data governance, data intelligence, data stewardship, artificial intelligence, and data quality initiatives. IData also has experts that can assist with data governance, reporting, integration and other technology services on an as-needed basis. Feel free to contact us and let us know how we can assist.
Image Credit: StockSnap_SNIKX3KGKD_BikeRace_PragmaticDG #1333