As often happens in the computer industry, nomenclature is unwieldy and flexible as technologists, sales & marketing, and the rest of the world clash.
My case in point is the phrase "data loss prevention" or DLP. In other articles, I have talked about DLP as a technology -- in that it is used to analyze the content of a document or message, determine whether the content references a concept confidential or protected in nature, and uses rules or reporting to handle the content. As the concept of DLP was developed in the last decade, the industry struggled to find an appropriate phrase that defined it: phrases including content monitoring & filtering, content analysis, deep packet inspection, and others were used, but the industry and analysts settled on data loss prevention.
Many companies are marketing "data loss prevention" in relation to their technologies, but not in the context of analysis of document content. Instead, their approaches include building a wall around all corporate data (such as on a mobile device, or in a cloud-based document-sharing service), or providing some regular expression matching for message content. This is well and good, but I would suggest these technologies fall under the larger strategy of information protection rather than being specifically about "data loss prevention".
This goes to the heart of the matter: when we build true data loss prevention systems, the intent is to protect confidential information rather than just bits and bytes of raw data. Under the fundamentals of information theory, data is just bits and bytes, but information is found where there is entropy, or value, in the data. This is what distinguishes data loss prevention technology from other data protection technologies, and perhaps the better phrase for the technology would be "information loss protection."
Practically, though, we are probably stuck with the labels that have been adopted. So, I suppose we can accept a variety of technologies under the strategy of data loss prevention, including the technology of data loss prevention itself. Unfortunately, this will continue to be confusing to those inside and outside of the industry and troublesome for sales and marketing.
Showing posts with label DLP. Show all posts
Showing posts with label DLP. Show all posts
Monday, November 14, 2011
Monday, June 27, 2011
Fully-Functional Data Loss Prevention
Since Data Loss Prevention (DLP) became a known technology in the computer security arena a few years ago, a number of vendors of existing non-DLP security products added basic DLP-like features to enable detection of some common private or confidential information. However, a complete DLP implementation involves more than just regular expressions to match patterns in text in, say, email messages.
Certainly, email is a significant vector by which data loss occurs. More generally, the DLP industry terms data traversing the network as Data in Motion. However, there are many more protocols than just email, not the least of which include web-based email services, such as Google Mail, and social media services, such as Facebook, that could also be data loss vectors. A complete DLP implementation will likely be able to work with a number of common network protocols to manage Data in Motion.
DLP also manages data in two other important situations, Data in Use and Data at Rest. Data in Use DLP can manage data used on a workstation, such as monitoring data being copied to a USB flash drive. Data at Rest DLP can inventory and manage the private and confidential data stored on workstation and server's hard drives.
The ways in which most DLP systems are able to discover protected information extend far beyond basic regular expressions. Common approaches include pre-packaged sets of terms, database fingerprints, file fingerprints, special code to match data like credit card numbers, and more. I previously wrote an article on Classes of Protected Information and DLP that goes into much more detail on this topic.
In addition to managing protected data in the scenarios of Data in Motion, Use, and Rest, and using multiple approaches to finding protected data, DLP systems also offer sophisticated configuration, reporting, alerting, and case management services. There may be situations where certain groups of users are allowed to work with certain kinds of confidential information while others are not -- a DLP system might be configured to monitor such information use for the privileged users and block use by other users. The depth of reporting and alerting capabilities offered by a DLP system can make a DLP installation more useful by providing information ranging from summaries to detailed violation information as needed for management and compliance reports. Finally, DLP case management tools can enable rolling up multiple incidents into a consolidated case that can be managed as necessary to resolution.
In summary, a DLP system is a significant addition to an organization's data security arsenal.
Certainly, email is a significant vector by which data loss occurs. More generally, the DLP industry terms data traversing the network as Data in Motion. However, there are many more protocols than just email, not the least of which include web-based email services, such as Google Mail, and social media services, such as Facebook, that could also be data loss vectors. A complete DLP implementation will likely be able to work with a number of common network protocols to manage Data in Motion.
DLP also manages data in two other important situations, Data in Use and Data at Rest. Data in Use DLP can manage data used on a workstation, such as monitoring data being copied to a USB flash drive. Data at Rest DLP can inventory and manage the private and confidential data stored on workstation and server's hard drives.
The ways in which most DLP systems are able to discover protected information extend far beyond basic regular expressions. Common approaches include pre-packaged sets of terms, database fingerprints, file fingerprints, special code to match data like credit card numbers, and more. I previously wrote an article on Classes of Protected Information and DLP that goes into much more detail on this topic.
In addition to managing protected data in the scenarios of Data in Motion, Use, and Rest, and using multiple approaches to finding protected data, DLP systems also offer sophisticated configuration, reporting, alerting, and case management services. There may be situations where certain groups of users are allowed to work with certain kinds of confidential information while others are not -- a DLP system might be configured to monitor such information use for the privileged users and block use by other users. The depth of reporting and alerting capabilities offered by a DLP system can make a DLP installation more useful by providing information ranging from summaries to detailed violation information as needed for management and compliance reports. Finally, DLP case management tools can enable rolling up multiple incidents into a consolidated case that can be managed as necessary to resolution.
In summary, a DLP system is a significant addition to an organization's data security arsenal.
Tuesday, June 7, 2011
It's 10:00pm - Do You Know Where Your Data Is?
Data can be stored in so many places and be so vulnerable to loss or exposure. The obvious risk and probability of loss for protected data stored on devices like laptops often motivates security staff to make improvements in this area. Many people have an "a-ha moment" when they see how Data Loss Prevention (DLP) discovery agents can find and report confidential or protected data stored in unexpected places.
It's good practice to inventory where and how confidential / protected data is stored, create policy that defines where and how such data should be stored, then move towards the goal defined by the policy and monitor progress. (Helpful side benefits of this process include improving your backup and archive coverage of protected data, reducing duplication of data, and assisting your business continuity planning.)
The initial inventory of protected data can be overwhelming -- data can be dispersed over all the personal workstations and laptops in the entire company and in the oddest nooks and crannies of servers. But it's good to know where your organization stands with regard to protected data, and what your biggest points of risk might be. If you found confidential financial data being stored on laptops that don't have disk encryption, maybe that's your prime starting point. If you found multiple copies of confidential data stored on a server, maybe it's just a matter of consolidating the data and keeping employees better informed about what location to use on the server for that data.
When it comes to writing your protected data storage policies, keep flexibility in mind. Mobility is a big factor in employee computing use cases today, so if important data on laptops is common, then maybe a disk encryption solutions for laptops is needed rather than disrupting employees' work by requiring them not to keep data on laptops.
When your protected data storage policy is defined, then it's time to move toward it. Education will be important so employees understand why and how this process is happening. Some time & effort will be required to implement the changes, and perhaps some new software will be required for encryption.
As progress is made, DLP discovery software can be used to measure and monitor the progress, and watch for significant deviations from the policy that need to be addressed.
It's good practice to inventory where and how confidential / protected data is stored, create policy that defines where and how such data should be stored, then move towards the goal defined by the policy and monitor progress. (Helpful side benefits of this process include improving your backup and archive coverage of protected data, reducing duplication of data, and assisting your business continuity planning.)
The initial inventory of protected data can be overwhelming -- data can be dispersed over all the personal workstations and laptops in the entire company and in the oddest nooks and crannies of servers. But it's good to know where your organization stands with regard to protected data, and what your biggest points of risk might be. If you found confidential financial data being stored on laptops that don't have disk encryption, maybe that's your prime starting point. If you found multiple copies of confidential data stored on a server, maybe it's just a matter of consolidating the data and keeping employees better informed about what location to use on the server for that data.
When it comes to writing your protected data storage policies, keep flexibility in mind. Mobility is a big factor in employee computing use cases today, so if important data on laptops is common, then maybe a disk encryption solutions for laptops is needed rather than disrupting employees' work by requiring them not to keep data on laptops.
When your protected data storage policy is defined, then it's time to move toward it. Education will be important so employees understand why and how this process is happening. Some time & effort will be required to implement the changes, and perhaps some new software will be required for encryption.
As progress is made, DLP discovery software can be used to measure and monitor the progress, and watch for significant deviations from the policy that need to be addressed.
Friday, June 3, 2011
Cloud Computing and Protecting Confidential Information
A couple of months ago, I talked about the implementation of DLP in cloud computing environments. Since then, I have seen a few examples of how security-oriented firms are working with cloud computing vendors, such as Tripwire, enStratus, and others working with cloud vendors to provide internal compliance and validation.
Meanwhile, we have seen several large-scale data breaches, including numerous attacks on Sony, that involve attacks through web servers.
A significant use case for cloud computing is to provide scalable web services, so we have an interesting and significant security intersection between deployments of web servers (often with vulnerabilities) in the cloud, and the need for web application firewall (WAF), data loss prevention (DLP), and intrusion detection/prevention (IDS/IPS) to protect the web servers and the information to which they provide access.
There are some difficult problems with protecting outward-facing cloud-based web servers, though. It might not be feasible to scale WAF, DLP, and IDS/IPS systems alongside the web servers. It may be challenging to be able to monitor and/or intercept web traffic -- especially SSL web traffic -- to protect against attacks and data loss.
A solution to this problem might be to incorporate WAF, DLP, and IDS/IPS technology into the web servers themselves, so as the web servers are scaled, the protection automatically scales also.
Meanwhile, we have seen several large-scale data breaches, including numerous attacks on Sony, that involve attacks through web servers.
A significant use case for cloud computing is to provide scalable web services, so we have an interesting and significant security intersection between deployments of web servers (often with vulnerabilities) in the cloud, and the need for web application firewall (WAF), data loss prevention (DLP), and intrusion detection/prevention (IDS/IPS) to protect the web servers and the information to which they provide access.
There are some difficult problems with protecting outward-facing cloud-based web servers, though. It might not be feasible to scale WAF, DLP, and IDS/IPS systems alongside the web servers. It may be challenging to be able to monitor and/or intercept web traffic -- especially SSL web traffic -- to protect against attacks and data loss.
A solution to this problem might be to incorporate WAF, DLP, and IDS/IPS technology into the web servers themselves, so as the web servers are scaled, the protection automatically scales also.
Wednesday, May 25, 2011
Insidious Insiders: Bank of America
When I talk or write about inappropriate confidential information disclosure, I often point out that data loss prevention (DLP) systems most commonly help reduce the everyday mistakes by well-intentioned employees just trying to do their jobs. A DLP system also helps discover a malicious insider gathering or passing confidential information to outsiders. Regardless of intent, a good DLP system can help administrators notice a trend of confidential leaks and help build a case file for action with regard to a problematic insider.
A story I saw today about a problem at Bank of America that has been under investigation for a while where an apparently-malicious employee, who had access to "personally identifiable information such as names, addresses, Social Security numbers, phone numbers, bank account numbers, driver's license numbers, birth dates, e-mail addresses, family names, PINs and account balances," allegedly passed this information to criminals. The estimated resulting direct financial loss is $10 million. Indirect losses, including employee time spent investigating the problem, cost of credit report monitoring for affected customers, revisiting policies and controls, and diminished brand may be significant as well.
A DLP system is one of the best practices that a business can put into place to help track and prevent data breach events. If you have a DLP system in place, make sure it is correctly configured, installed in the correct locations in your network, servers, and clients, and make sure it is monitored. (It is highly likely that Bank of America has a DLP system in place, but I do not have any knowledge in regards to whether information from a DLP system helped with the investigation of this case.)
Other best practices for protection of information include:
A story I saw today about a problem at Bank of America that has been under investigation for a while where an apparently-malicious employee, who had access to "personally identifiable information such as names, addresses, Social Security numbers, phone numbers, bank account numbers, driver's license numbers, birth dates, e-mail addresses, family names, PINs and account balances," allegedly passed this information to criminals. The estimated resulting direct financial loss is $10 million. Indirect losses, including employee time spent investigating the problem, cost of credit report monitoring for affected customers, revisiting policies and controls, and diminished brand may be significant as well.
A DLP system is one of the best practices that a business can put into place to help track and prevent data breach events. If you have a DLP system in place, make sure it is correctly configured, installed in the correct locations in your network, servers, and clients, and make sure it is monitored. (It is highly likely that Bank of America has a DLP system in place, but I do not have any knowledge in regards to whether information from a DLP system helped with the investigation of this case.)
Other best practices for protection of information include:
- Limiting the amount and scope of information available to employees to that necessary to do their jobs. Often, employees are given increasing access to information over their tenure, and it's a good idea to review access to make sure potential for problems is limited.
- Logging information access and reviewing the logs for unusual patterns. A Security Event Manager (SEM, also known as SIEM) can help with this by making it possible to centrally manage and review information from servers.
- Limit network access for workstations and servers. Servers should generally not be using protocols like Internet Relay Chat or accessing random web sites. A network protocol manager or firewall can be configured to prevent unexpected network use. Unexpected use of web sites or network protocols from servers might be indicative of an intrusion that should be investigated.
Friday, May 20, 2011
Classes of Protected Information and DLP
Data Loss Prevention (DLP) systems have to deal with a variety of formats of data and identify protected data in those formats. In general, protected information falls into these formats:
For corporate proprietary information, document fingerprinting is the predominant approach to identifying parts or complete copies of proprietary documents. This requires the administrator to register proprietary documents with the DLP system, and then the DLP system can match fragments or wholesale copies of the proprietary documents.
Another approach that can be used for proprietary documents is to embed tags in the documents, such as "Company Confidential", and then add a simple rule to the DLP system to watch for that tag. However, this depends on corporate users applying the correct tags to the documents, and is easy for a malicious insider to circumvent, for example, by simply removing the tag before transmitting the document to an unauthorized recipient.
For data like personal health information (PHI) or personal financial information (PFI), several approaches (or a combination of approaches) are typically used. A combination of search terms can be used to determine if data contains information referring to a particular individuals or group of individuals, plus whether the data contains significant information about those individuals. For example, an email message from a bank containing the customer's account number, name, and account balance, it might be considered to be information protected under the Gramm-Leach-Bliley Act (GLBA).
Another approach to PHI and PFI is to use information from a corporate database, such as account numbers and customer names, in the DLP system to search for matches. If an account number and associated customer name turns up in an email message, the message might be considered to contain information protected under GLBA.
A third approach, specific to personal financial information, is to look for credit card information. Credit card numbers use a standard format and are assigned in specific ways, so it is possible to look at a sixteen-digit number and determine with a high degree of accuracy whether that number is probably a VISA or MasterCard credit card number.
For personal identifying information, an approach is to look for national identification numbers, state driver's license numbers, or account numbers. In the United States, the Social Security Number (SSN) is often used (and abused) for purposes of identification and authentication for financial and health purposes, and as such has gained status as a protected piece of information. Unfortunately, the format of the SSN was developed without the concept of check digits or embedded validators, so it is easy for a DLP system to mistake a number in the form 123-45-6789 as an SSN.
As for structured data, DLP systems can identify protected contents in a couple of ways. One is to write rules for the DLP system that match the format of data typically used in a company, such as forms that are often used for things like customer orders. Another approach is to use information from a corporate database, such as account numbers and customer names, in the DLP system to search for matches.
These formats cover the majority of ways I have seen protected information stored and transmitted in ways that DLP systems can help identify and protect the data.
- Unstructured text - as found in text documents - including various types of information:
- Corporate proprietary information or trade secrets
- Personal health records
- Personal financial records
- Personal identifying information
- Structured data - as found in spreadsheets, tables, database output, and CSV files
For corporate proprietary information, document fingerprinting is the predominant approach to identifying parts or complete copies of proprietary documents. This requires the administrator to register proprietary documents with the DLP system, and then the DLP system can match fragments or wholesale copies of the proprietary documents.
Another approach that can be used for proprietary documents is to embed tags in the documents, such as "Company Confidential", and then add a simple rule to the DLP system to watch for that tag. However, this depends on corporate users applying the correct tags to the documents, and is easy for a malicious insider to circumvent, for example, by simply removing the tag before transmitting the document to an unauthorized recipient.
For data like personal health information (PHI) or personal financial information (PFI), several approaches (or a combination of approaches) are typically used. A combination of search terms can be used to determine if data contains information referring to a particular individuals or group of individuals, plus whether the data contains significant information about those individuals. For example, an email message from a bank containing the customer's account number, name, and account balance, it might be considered to be information protected under the Gramm-Leach-Bliley Act (GLBA).
Another approach to PHI and PFI is to use information from a corporate database, such as account numbers and customer names, in the DLP system to search for matches. If an account number and associated customer name turns up in an email message, the message might be considered to contain information protected under GLBA.
A third approach, specific to personal financial information, is to look for credit card information. Credit card numbers use a standard format and are assigned in specific ways, so it is possible to look at a sixteen-digit number and determine with a high degree of accuracy whether that number is probably a VISA or MasterCard credit card number.
For personal identifying information, an approach is to look for national identification numbers, state driver's license numbers, or account numbers. In the United States, the Social Security Number (SSN) is often used (and abused) for purposes of identification and authentication for financial and health purposes, and as such has gained status as a protected piece of information. Unfortunately, the format of the SSN was developed without the concept of check digits or embedded validators, so it is easy for a DLP system to mistake a number in the form 123-45-6789 as an SSN.
As for structured data, DLP systems can identify protected contents in a couple of ways. One is to write rules for the DLP system that match the format of data typically used in a company, such as forms that are often used for things like customer orders. Another approach is to use information from a corporate database, such as account numbers and customer names, in the DLP system to search for matches.
These formats cover the majority of ways I have seen protected information stored and transmitted in ways that DLP systems can help identify and protect the data.
Wednesday, May 4, 2011
Data Loss Prevention and Mobility
At Palisade we are often asked how to protect data from loss when your employees and/or partners all have access to your corporate private/privileged data through handy little gadgets like iPhones.
The problem we are finding is that gadget vendors have not provided hooks into the devices so we can do DLP on the gadgets directly. In fact, software on iOS devices is intended to be quite isolated to prevent any application accessing information that belongs to another application, such as email messages or stored PDFs.
Enter some pretty cool software from Whisper Systems for Android systems. WhisperCore looks very intriguing:
The problem we are finding is that gadget vendors have not provided hooks into the devices so we can do DLP on the gadgets directly. In fact, software on iOS devices is intended to be quite isolated to prevent any application accessing information that belongs to another application, such as email messages or stored PDFs.
Enter some pretty cool software from Whisper Systems for Android systems. WhisperCore looks very intriguing:
WhisperCore integrates with the underlying Android OS to protect everything you keep on your phone. This initial beta features full disk encryption and basic platform management tools for Nexus S phones. WhisperCore presents a simple and unobstrusive interface to users, while providing powerful security and management APIs for developers.Will be looking into this more deeply :-) Maybe this would encourage Apple to provide hooks for similar software into iOS.
Wednesday, April 27, 2011
Surprising Data Loss Vectors
With the 2011 Verizon Business Data Breach Investigation Report and breach after breach after breach recently in the news, you might be thinking that information security is all about malevolent actors right now. The "black hats" seem to have become very good at targeting, infiltrating, and extracting valued data from desirable targets.
An article yesterday by Ellen Messmer at Network World spotlights another important issue in information security today: business partners sharing information insecurely. At Lutheran Life Communities (LLC), when they installed Palisade Data Loss Prevention (DLP) systems , it was found that business partners were transmitting personal health information (PHI) insecurely to LLC. LLC has chosen a practical response by warning business partners of the problem.
In my experience, this is not an isolated problem. In the past, the DLP vendor community has highlighted the "insider problem" where employees -- usually just trying to do their jobs -- end up using poor business practices and cause frequent exposures of personal financial information (PFI) and/or personal health information, the two most highly-regulated types of personal identifying information (PII). However, in numerous DLP installations I've observed, I have seen data inbound into organizations containing PFI and PHI violations, such as unwary customers sending credit card information in unsecured email messages into companies to request purchases. I have also seen medical facilities where PHI was unexpectedly being transferred insecurely in and out of the organization, just as LLC noted in Ellen's article.
We have become aware of the risks of data loss. Governments have begun enforcing data protection requirements. We have developed policies and tools that have significantly raised the standards for protecting confidential information. Let's put these tools and policies to good use.
An article yesterday by Ellen Messmer at Network World spotlights another important issue in information security today: business partners sharing information insecurely. At Lutheran Life Communities (LLC), when they installed Palisade Data Loss Prevention (DLP) systems , it was found that business partners were transmitting personal health information (PHI) insecurely to LLC. LLC has chosen a practical response by warning business partners of the problem.
In my experience, this is not an isolated problem. In the past, the DLP vendor community has highlighted the "insider problem" where employees -- usually just trying to do their jobs -- end up using poor business practices and cause frequent exposures of personal financial information (PFI) and/or personal health information, the two most highly-regulated types of personal identifying information (PII). However, in numerous DLP installations I've observed, I have seen data inbound into organizations containing PFI and PHI violations, such as unwary customers sending credit card information in unsecured email messages into companies to request purchases. I have also seen medical facilities where PHI was unexpectedly being transferred insecurely in and out of the organization, just as LLC noted in Ellen's article.
We have become aware of the risks of data loss. Governments have begun enforcing data protection requirements. We have developed policies and tools that have significantly raised the standards for protecting confidential information. Let's put these tools and policies to good use.
Thursday, March 17, 2011
Cloud Computing and Data Loss Prevention Implementation
I have been studying the Cloud Security Alliance's Security Guidance and other resources for the past few weeks with a focus on placement of Data Loss Prevention (DLP) capabilities.
At this time, it seems that the typical position for DLP in public clouds is in conjunction with other resources in private or public Infrastructure as a Service (IaaS). Data-in-motion DLP systems in a cloud can be positioned logically adjacent to servers deployed in a cloud to monitor and protect information on those servers. Data-at-rest and data-in-use DLP agents can be deployed on servers in the cloud to catalog and protect data on those servers. However, there is nothing significantly new or better in these DLP implementation approaches than is currently available traditional servers.
What would be useful in cloud implementations is an API or specification to allow DLP interaction in all three major service models: Infrastructure as a Service (IaaS), Platform as a Service (PaaS) and Software as a Service (SaaS). VMware's vShield family of products look like a step in this direction for the IaaS model: the vShield endpoint product looks like it potentially could enable data-at-rest DLP, but the network-oriented vShield products do not appear to provide direct access to network data streams to enable data-in-motion DLP.
I am looking forward to engaging with cloud computing vendors to see if we can create a generalized specification for access to cloud systems for DLP management and remove some of the fuzzy haze enveloping the data.
Updated: Christopher Hoff blogged about the lack of security functionality in cloud services (IDS/IDP, WAF, DLP, etc) about a year and a half ago. Do we dare hope that cloud providers are becoming any more interested in security than they were then?
At this time, it seems that the typical position for DLP in public clouds is in conjunction with other resources in private or public Infrastructure as a Service (IaaS). Data-in-motion DLP systems in a cloud can be positioned logically adjacent to servers deployed in a cloud to monitor and protect information on those servers. Data-at-rest and data-in-use DLP agents can be deployed on servers in the cloud to catalog and protect data on those servers. However, there is nothing significantly new or better in these DLP implementation approaches than is currently available traditional servers.
What would be useful in cloud implementations is an API or specification to allow DLP interaction in all three major service models: Infrastructure as a Service (IaaS), Platform as a Service (PaaS) and Software as a Service (SaaS). VMware's vShield family of products look like a step in this direction for the IaaS model: the vShield endpoint product looks like it potentially could enable data-at-rest DLP, but the network-oriented vShield products do not appear to provide direct access to network data streams to enable data-in-motion DLP.
I am looking forward to engaging with cloud computing vendors to see if we can create a generalized specification for access to cloud systems for DLP management and remove some of the fuzzy haze enveloping the data.
Updated: Christopher Hoff blogged about the lack of security functionality in cloud services (IDS/IDP, WAF, DLP, etc) about a year and a half ago. Do we dare hope that cloud providers are becoming any more interested in security than they were then?
Subscribe to:
Posts (Atom)