Showing posts with label Data mining. Show all posts
Showing posts with label Data mining. Show all posts

Friday, December 9, 2011

Can Software help Health care?


Apps, apps and more apps. Software is everything and everything runs on software.


Almost every industry in the U.S. has been disrupted by software. The health care field is not one of them.

Easily accessible consumer information makes everyone a little bit doctor. Emerging portable diagnostic devices will strengthen the transition. Are we up to it?

Not yet.

A large majority of people want to own their health information. Many want to store it online and have better control over it. Yet, most people don't want any extra work associated with updating and maintaining it. As public health record (PHR) expert Jim Tate said:
My 'dream PHR' continues to evolve. What I want now is a elegant interface which gives me a real time dynamic look into my record located somewhere in the stratosphere. I don’t want to have to do anything. Please don’t ask me to input anything or make more than 2 or 3 decisions. Make it simple, intuitive, powerful, and available on the internet and I will use it. Maybe."
Yes, we are inherently lazy, always trying to find shortcuts and reduce the amount of work to get a task done. Why spend time creating and maintaining our own records when a doctor can do it for us? Or even better, why not just live and enjoy life before we get sick?

Problem is, most of us at various stages throughout life suffer from subtle conditions like food sensitivities or allergies that doctors can't easily diagnose. They're relatively minor in severity, but if managed properly our lives would be a lot better off. So maybe all we need is a doctor who's just a mobile app away, always ready to answer our questions for free.

But will these legions of online doctors have enough insight into our everyday lives to know what we eat, what we breath, and what it is we're not saying to form an expert opinion?

Not likely. Even if we could wear mobile devices - always on, always connected, counting our steps, cataloging our night sweats, and equipped with miniature cameras to photograph what we eat - would the doctors be able to process all that information to form a useful diagnosis?

Aurametrix is an advanced analysis tool that correlates our symptoms, reactions and feelings based on what we enter into the system about our diet, exercise and conditions. Results from early usage of the tool show that even occasional sparse information - entered on days we feel better or worse than average - if properly evaluated can provide a snapshot of our health with sufficient insight to connect the dots to better health.  It's a form of collective intelligence that's already providing interesting discoveries without the need for us to know all the details. For example, it already knows what our foods consist of, how our daily activities or feelings align with past events, and that there are commonalities among many different things.

The future is already here but are we ready for the future?

REFERENCES

Archer N, Fevrier-Thomas U, Lokker C, McKibbon KA, & Straus SE (2011). Personal health records: a scoping review. Journal of the American Medical Informatics Association : JAMIA, 18 (4), 515-22 PMID: 21672914

Kim J, Bates DW. Analysis of the definition and utility of personal health records using q methodology. J Med Internet Res. 2011 Nov 29;13(4):e106.

Geissbuhler A, Kimura M, Kulikowski CA, Murray PJ, Ohno-Machado L, Park HA, Haux R.
Confluence of disciplines in health informatics: an international perspective. Methods Inf Med. 2011 Dec 6;50(6):545-55.

Macedo LG, Maher CH, Latimer J, McAuley JH. Feasibility of using Short Message Service (SMS) to collect pain outcomes in a low back pain clinical trial. Spine (Phila Pa 1976). 2011 Dec 3.

Lo Piparo E, Worth A, Manibusan M, Yang C, Schilter B, Mazzatorta P, Jacobs MN, Steinkellner H, Mohimont L. Use of computational tools in the field of food safety. Regul Toxicol Pharmacol. 2011 Aug;60(3):354-62. Epub 2011 May 12.

Benito PJ, Neiva C, González-Quijano PS, Cupeiro R, Morencos E, Peinado AB. Validation of the SenseWear armband in circuit resistance training with different loads. Eur J Appl Physiol. 2011 Dec 6.

Yoshiaki Sugawara,Chie Sugimoto, Sachiko Minabe, Yoshie Iura, Mai Okazaki, Natuki Nakagawa, Miwa Seto, Saki Maruyama, Miki Hirano and Ichiro Kitayama. Use of Human Senses as Sensors. Sensors. 2009, 9(5), 3184-3204; doi:10.3390/s90503184

Sunday, March 28, 2010

Health Data, Self-serve, Visualization, Semantic Analysis and Collective Intelligence

Notes from the Biomedical Data Mining Camp Session led by Dr. Irene Gabashvili and other health-related discussions at Data Mining Camp and Transparency Camp 2010.

The title for the session was "Biomedical Data Mining: Successes, Failures, and Challenges" (streamed online from Fireside C).
The topic stemmed from the similarly named last year's session - Biomedical Data Mining: Dimensionality, Noise, Applications - now split into several discussions including Bioinformatics & Genome Sequencing organized by Raymond McCauley, Dimensionality Reduction moderated by Luca Rigazio (HLDA / HDA; LDPP; Core Vector Machines; Sparse Proj SQ; Random Projection and Feature Selection).

Main sub-topics of the biomedical session were:
  • Reality Mining
  • Visualization
  • Imaging
  • Signal Processing
Reality mining is expected to improve public health and medicine. It was named as one of "10 emerging technologies that could change the world". This is not only about mining data pertaining to social behavior - although social factors do impact well-being and social data provides valuable health predictors. It's also about health data collected in real and near real time. Audience asked about ways to collect health data and their limitations. Some of the questions reflected earlier Q&A with the Data Mining Camp expert panel , especially Dr. Michael Walker, author of numerous FDA and CLIA-approved products to diagnose and treat disease. He commented on the need to have tools well beyond the data available to take cell samples every two minutes instead of laboratory testing every few months - in order to allow cellular simulations and obtain parameters for differential equations. Sampling frequencies don't come close to allow this kind of modeling. Irene Gabashvili agreed that first-principles cellular modeling for predicting health won't be possible (although there are engineers that believe a platform for real-time in-vivo measurements of most cells can be developed). Yet, good predictors could be and will be developed - based on sensors measuring macro-level observations and missing value estimators. Genetic information is not enough, we need to capture environmental risk factors. How can we separate genes and environment?, asked one of the participants. Aurametrix' initial focus is on chemicals in our food - and even though some may argue that our taste and satiety mechanisms are dictated by genes, food analytics provides insights into non-genetic components of our health. Other questions were on the time line for body sensor networks and data growth. Jeffrey Nick of EMC estimates that personal sensor data will balloon from 10% of all stored information to 90% within the next decade. Irene Gabashvili thinks that this will happen rather sooner than later, perhaps in the next two years.

Another interesting aspect of reality mining is crowdsourcing or collective intelligence - in order to get useful information from all the data (temporospatial location, GPS, activity, food, symptoms, behavior, communication content, proximity sensing), we need to analyze it not only on individual but also group level. We need to share more, without sacrificing privacy and security. Collective contributions can be reliable - Shamod Lacoui's answer to this is in selecting those who contribute, restricting inputs to domain. It would help to “filter out the dross”, while “saving the best”. It is needed to suppress noise, to infer intelligence from the collection of facts, clicks, steps, whatever one can contribute. This resonates with discussions at the Transparency Camp - one of the useful tools is SwiftRiver - free, open source software platform to validate and filter news. Swift relies on Natural Language Processing, Machine Learning and Veracity Algorithms to track and verify the accuracy of reports and suppress noise (like duplicate content, irrelevant cross-chatter and inaccuracies). Transparency Camp also posed a question on whether there is a need for an FDA-like institution to ensure information safety and healthy information consumption.

Self-serve was a topic of a smaller Data Mining Camp Session. Even though it was aimed at sales reps that need to go beyond Excel spreadsheets to mine private data of their interest, self-service is currently the only option for health care consumers. People need to analyze everyday life for health implications. They need better tools to not focus on metrics that are easy to collect instead of metrics we need to collect.

In order to mine high-dimensional health space, many disparate types of data should be mashed and validated, gaps should be bridged and structured metadata added to data. Randy Kerber talked about data formats and approaches to make it happen. Semantic web discussions involved NoSQL experts that mentioned limitations of gaining popularity technologies such as MongoDB, Cassandra and HBase. Another relevant session - on cloud computing - discussed its (sometimes over-rated ?) performance and Hadoop technologies.

Visualization techniques provide one of the most effective methods of extracting knowledge from health data. Remember who invented the pie chart? That's right, it was Florence Nightingale, a nurse who needed a way to better represent her data. one of the most famous examples of visualizing epidemiological data was Dr. John Snow's map of deaths from a cholera outbreak in London, Many other techniques and software tools exist, but maps remain popular - especially google maps API. One of popular tools for epidemiological data is Google Maps API. For example, it embeds Google Maps into healthmap.org with JavaScript.

One of the participants of Biomedical Session developed kidsdata.org (@kidsdata on twitter). It provides insights into geospatial autism statistics and visualizes trends and other useful health-related information.

Another way to display geospatial data is Dynamic Choropleth Maps. Complex networks can be also explored with alluvial diagrams and other approaches. More visualization techniques and tools can be applied to health data - to look at the data in new ways and gain useful insights.

Some of the questions from the audience were on the availability of data. Sources discussed included CDC (see, for example, NHANES laboratory files; eHealth metrics) and Entrez Life Sciences databases.

Signal Detection and Signal Processing for Mining Information was another discussion topic.
Questions were on data mining versus simple tracking and signal monitoring. It was agreed that data mining is the key to health management. Cardionet, body sensors (see posts on teletracking, M-health, Telemedicine: part 1; Telemedicine: part 2; Health 2.0 Software tools, Devices to keep you healthy), SNP detection, telemedicine applications, random and rare electrocardiographic events and other applications were also discussed.




See other materials from Data mining Camp 2010:


Reblog this post [with Zemanta]

Sunday, March 21, 2010

Mining Data Mining Camp Impressions


Data Mining Camp organized by Patricia Hoffman and San Francisco Bay Area Chapter of ACM, the Association for Computing Machinery, uses an Open Space Technology (OST) approach - no formal agenda beyond the overall data mining theme. Except the expert panel, sessions are compiled on-the-fly, based on real-time interest and participation.

Overal, the unconference - with a new location and almost doubled attendance - was a success - even if judging only by results of twitter sentiment analysis tools (subject of a not-so-successful data mining camp topic) - tweetfeel, twitrratr and twendz.

Obviously, completely ad-hoc sessions could be a bit chaotic - even though organizers briefly presented their topics and rooms were assigned adter counting a show of hands, there were surprises and unmet expectations. Here are sample quotes:
DataJunkie: OMFG Chaos trying to set up and plan which sessions to attend. Idea: have an online vote, and use sim annealing for scheduling. #DMCAMP

ihat: some sessions at #dmcamp have very low signal-to-noise... feature selection referenced tibshirani and boyd. and ppl butchered their work...

Many people preferred traditional formats to round table discussions - tutorials were the most attended sessions, while discussions were either the most or least liked sessions. The arrangement of chairs and the look of the room preset expectations of participants - some organizers did not really plan to present but had to come out with slides or tutorials. Great observation by Dominique Levin:

NextGenCMO
Room shape impacts success of un-conference: Circle of chairs works wonders to solicit audience participation at #dmcamp. Circle time!

Another interesting observation was that Linkedin turned to be the most efficient marketing tool for the conference. Twitter and other social networks did not seem to have an impact. The explanation could be very simple though - age group and education level of the target audience.

A winner of retweets was Chris Wensel -(interpreted as "influencer" by twitter data mining tools) - his message "Facebook dropped Cassandra for inbox search and hired HBase person to switch" was retwitted 18 times.

See also:
and last year's notes:


Reblog this post [with Zemanta]

Sunday, November 1, 2009

Biomedical Data Mining: Dimensionality, Noise, Applications

Health Prototype CandidatesImage by juhansonin via Flickr

ACM Silicon Valley Data Mining Camp on November 1, 2009
has attracted more than 200 people with different backgrounds and interests. It was held at Hacker’s Dojo, sponsored by REvolution Computing, KXEN (Knowledge Extraction Engines), and LinkedIn (See notes on this event by @Andraz of Zemanta, Ken's open source tools, and relevant #dmcamp twits) . Biomedical/Healthcare data mining topic was suggested by Junling Hu and Irene Gabashvili and supported by A.J. Chen, Greg Makowski, Sukanta Ganguly, and 40 other participants of the Data Mining Camp. Below is a brief transcript of the discussion. The session started from introductions, here are some of them:
  • Irene, with background in biophysics, medical informatics and CS, pursuing a personal health management venture, interested in data mining to advance personalized medicine;
  • Lawrence, with background in physics and software engineering and interest in health IT. He is the organizer of Google Wave meetup (you may know about Google Health Wave);
  • Hua, formerly with Kaiser, interested in medical scheduling and web development;
  • Liana, interested in Natural Language Processing for biomedical knowledge mining;
  • Maura, interested in Health IT, medical engineering and security;
  • Magnus, developing Medical Databases;
  • Kevin, interested in medical startups;
  • Watson, with background in genomics and machine learning;
  • Peters, working on medical devices and software embedded systems;
  • Steve, formerly of Applied Biosystems;
  • Roy of Codexis, focusing on data mining and pattern recognition in multivariate time series
  • Jima, with background in medical informatics;
  • Karsleep, interested in biomedical data mining;
  • Deena, scientific analyst interested in how data mining technologies could be applied to healthcare;
  • Junling Hu, scientist at Bosch, working on a device and software collecting and analyzing patients' information, based on daily questionnaires and other collected data.
There also were people interested in the subject merely as healthcare consumers (aren't we all?). Junling started the session from mentioning a recently published paper on computer technologies for healthcare determining strategic directions in the area. Irene also suggested to check the mHealth Summit focusing on mobile technologies to improve research data collection, healthcare delivery, and health outcomes. Junling described the project she was working on - inexpensive device collecting data and sending it to a "coaching" nurse that monitors stay-at-home patients. Next step is to mine the data automatically, thus reducing the load on healthcare professionals without sacrificing patients' well-being. Junling also mentioned some of the challenges such as compliance of participants who are typically not eager to fill out the 20-question surveys. This is especially bad for obesity studies. The data mining challenges mentioned during the session were: (1) Missing Data We are not talking about sparse data (discussed in one of the previous sessions on data mining with R), but actually missing data. Data is sparse if only a small fraction of the attributes are non-null - like the number of items we typically buy in a grocery store is much less than the number of products they offer. Data is missing if the values were never entered or the member combination is not meaningful (for example, obstetrics/gynecolgy values not meaningful for men) . One of the suggestions from the experts in the audience was to utilize "multiple imputation". Other suggestions included "once-a-week" questioning instead of daily surveys. Irene mentioned the 7D-PAR (Seven-Day Physical Activity Recall) , one of standardized questionnaires developed in the 80s (1,2) and other established methods. Questions and comments from the audience:
  • Data mining methods utilized for Chronic Disease Assesment and Elderly monitoring. Junling talked about unsupervised classification algorithms and two supervised learning methods she found to be most useful for her work - SVM and logistic regression. Both were equally good in predicting hospitalization events
  • Indicators of Goodness of Model Predictions. Suggested events were hospitalizations, mortality... It was noted that good indicators are yet to be found.
(2) Performance Criteria This was another health data mining challenge emphasized during the discussion. All standard methods can be applied such as accuracy, precision, recall, true positives, false positives and especially combinations of the last 2 measures. Junling mentioned breast cancer classifier developed by Siemens and other algorithms predicting emergency situations with 90% accuracy. Irene noted that one of the problems of digital mammography and other cancer predictors is a high rate of false positives. From 30 to 40% of cancers are overdiagnosed (3), thus increasing healthcare costs. This has to be changed. Several people in the audience emphasized that existing methods are averaging the population. Medicine needs to be truly personalized, we need better methods and more data. (3) Large Number of Input Features One of the main problems of health data mining is coping up with large number of input features. Obviously, a 20-question test is not sufficient. Should it rely on thousand questions or trillion inputs? And how to select a subset of relevant features to build robust learning models? Junling's preffered approaches are logistic regression and singular value decomposition. She would add features one by one and check if the overall accuracy for predictions remains good. Questions from the audience included:
  • A 3-5 year Vision for Health Data Mining: what do we expect to achieve? Participants expressed an optimistic outlook
  • Ray: Are most input variables discrete or continuous? The answer was: mixed
(4) Very-Large Scale Data The good thing about pattern recognition is that the more patterns you have, the better it performs. Google translator is a good proof of this assertion (although this translator needs even more patterns to do a decent job). Biomedical data sets such as CT, MRI, PET scans and other image data, gene expression, genetic variation are very large scale in nature. The challenge for data miners is to integrate and extract information from data of such scale. Questions from the audience:
  • What are the other large-scale studies trying ot mine patterns in health data, outside of US? Studies in China and Taiwan using similar devices and models; also in Europe
Adding to this interesting discussion that was unfortunately interrupted because of the lack of time, I'd like to mention a few other challenges facing biomedical data mining.
  • We should not underestimate the complexity of relationships between causative and effect variables in human health. Simplistic approaches are deemed to fail. Over-fitting could be a problem too
  • Integration between heterogeneous data sources and types,and putting content in context (semantic integration) remains a challenge.
  • Privacy Concerns associated with the Sharing of Individual Health Information.
References
  1. Blair S. How to assess exercise training habit and physical fitness. In: Behavioral Health, edited by Matarazzo JD. New York: Wiley, 1984, p. 424-447.
  2. Rauramaa R., Tuomainen P., Väisänen S., and Rankinen T. Physical activity and health- related fitness in middle-aged men. Med Sci Sports Exerc 27: 707-712, 1995.
  3. Gøtzsche, P.C., Jørgensen, K.J., Mæhlen, J. and Zahl, P.-H. Estimation of lead time and overdiagnosis in breast cancer screening. British Journal of Cancer (2009) 100, 219–219.
Reblog this post [with Zemanta]
blockquote { margin:1em 20px; background: #dfdfdf; padding: 8px 8px 8px 8px; font-style: italic; }