Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Friday, August 25, 2017

New Certification and Micro-Credentialing

In 2016 and continuing through the summer of 2017, a number of universities offering either a traditional face-to-face 2-year MBA, or an executive MBA, began confirming what many observed: enrollments started to drop, and students and employers commented that the cost had risen too high. Students and employers could not justify the cost due to a lack of return on investment.

People are turning to alternatives such as micro-credentialing which forms the core part of a competency-based learning program. Even Google is offering micro-credentials in its G-Suite for Education, which helps students develop skills using its cloud-based software.

Organizations are developing fast-track certification and micro-credentialing programs in response to quickly evolving industries and the need to obtain and demonstrate mastery with specific skills and knowledge.  Some of the emerging areas include new data analytics techniques, new areas of medical technology, home health care provider management, hospitality marketing, technology entrepreneurship, drones and UAV operation and analytics, urban organic farming, and more. 

Certification providers include companies with specialized experience and experts, colleges and universities, professional associations, and government agencies.
  •  Assessment to determine needs for new skills and knowledge
  •  Emerging needs aligned with certificates
  •  Situated learning: connect knowledge and skill to real-life setting
  •  Fast-Track Certification: Fewer courses, tighter timeline
  •  Characteristics of a “Fast-Track” program
  •  Digital badges used to motivate
  •  Content quality control to assure relevance of the content
  •  Assessment strategies to apply knowledge and skills in real-life situations
  •  Collaboration to encourage learning from each other
Mini-credentialing and certification programs appeal to individuals who need to expand their skills, and to do It quickly. Ideally, an individual should be able to complete their training within six months. In addition, the program should be affordable so that there is a very clear positive return on investment which more than pays for itself in increased income, expanded opportunities, and enhanced adaptability.

Big Data and Machine Learning: Susan Smith Nash and seismic lines for the Gulf of Mexico


Tuesday, April 11, 2017

Big Data and Deep Learning: Industry Downturn Means Uptick in New Analytics

From the Midland Register Times / April 2...
 Permian Basin operators are drilling deep and long — laterals — in order to recover more of the region’s crude and natural gas.

They’re also going deep — as in deep learning — as part of those efforts.

High-tech advances such as big data, deep learning and artificial intelligence are increasingly finding their ways into upstream exploration and production operations. For example, Exxon Mobil Corp. recently set a record for high performance computing for reservoir simulation.

Big data
Technological advances have created a wide spectrum of data for operators that goes far beyond well logs, seismic surveys and pressure readings.

“(It’s) massive amounts of data generated by different methods,” said Susan Nash, director of education and professional development with the American Association of Petroleum Geologists.
 “It’s so massive it’s contained in the cloud and other ways of organizing the data.”

That data can come in structured form, as in databases, or in unstructured forms, as in emails or PDFs, anything that can be digitized, she said.

To continue, click the link: http://www.mrt.com/business/oil/article/Industry-drills-deep-to-improve-production-11039830.php

Sunday, June 23, 2013

Data Mining: What You Might Not Know

The ultimate goal of data mining is not the acquisition of data, but the exploration and analysis  of massive amounts of data resulting in patterns, rules, and relationships. One of the key outcomes is the identification of reliable and meaningful patterns.

Meaningful patterns can do the following:

* Model typical behaviors
* Identify atypical behaviors
* Express possible cause and effect relationships
* Explain past behaviors
* Describe current conditions
* Develop a predictive model for the future

While developing patterns in data drawn from different types of data can show meaningful  relationships, time-series data mining can be used to postulate

* causality
* life-cycle behaviors
* impacts of proximal or distal relationships
* cluster formation and disaggregation

Characteristics of Data Mining Data

The data may represent changes of behavior and activities over time, or, alternatively, it could  represent the relationship of different types of data which have been collected at one point in  time.

* Same type of data, collected at different points in time
* Different types of data, collected at the same point in time
* Data streams, which involve ordered sequences of items that arrive over time

* offline: regular chunked arrivals
* online: continuous flow

Planning the Data Mining Process

The development of data mining can follow a fairly clear process:

1.  Determine the problem and definition
2.  Determine the characteristics of the data
3.  Develop a plan for data mining
4.  Review of similar data mining projects and algorithms
5.  Become familiar with data, data issues, potentially meaningful data subsets
6.  Data preparation and conditioning
7.  Model and algorithm development
8.  Evaluation of model / comparison with other models
9.  Implementation, which involves generating reports, or continuing to develop ongoing activities


Data Mining: Key Tasks

Key tasks in data mining include the following:

* Identification
* Eliminate unnecessary or distracting frequent item sets
* Definition and differentiation
* Optimize storage and recovery of streams
* Classification
* Cluster recognition
* Segmentation
* Discovery of motifs (sub-sequences)
* Detection of similar clusters or sets
* Detection of outliers and anomalies
* Create predictive models

Data mining that involves continuous streams of data presents unique challenges because of the  nature of the data and types of patterns that are meaningful, given the array of patterns that are  possible to develop. It is also challenging to integrate incoming data with existing databases in  order to qualitatively evaluate patterns in a timely way. It is also challenging to avoid the  "concept drifting" problem, which means that the usefulness and validity of the results will  degrade over time.

In general,

* Sensors and surveillance: networks, physical locations, manufacturing, transportation
* Performance monitoring: manufacturing, networks, controls
* Transaction / activity monitoring: retail, web performance, manufacturing

Algorithms

A literature review of algorithms suggests that data mining for data streams is generally  performed using three different major classifications of algorithms, and that they do not yield  the same results, which could be quite significant, depending on the application.

Landmark Window Based Data Mining

What is measured is the difference between a specific time-stamp (the landmark) and the present.

Pros: Complete comparision with an a priori property
Cons: The order in which information is considered and placed into sets can lead to errors

Damped Window

This approach privileges new data over old or historical data, which means that the older data drops out of  consideration for developing sets

Pros: Efficient use of resources, eliminates old and obsolete information
Cons: Large errors may be made because the information being eliminated may be important for the  rule to be effective

Sliding Window

Sliding favors new data (as in the Damped Window approaches), but does not completely eliminate  old data. Instead, it incorporates summarized versions of old data and data relations.

Pros:  Can incorporate past data and do so relatively quickly
Cons:  The assumptions made to create summaries of old data sets can be flawed

General Observations and Conclusions

At this point in time, the ability to collect data continues to expand and sometimes dramatically,  thanks to technological advances in both hardware and software. However, a review of the processes  and the literature make it clear that the algorithms use to process and make meaning of the data  batchs and streams differ widely. Consequently, the results and conclusions that are created using  data mining techniques (both collecting and in analyzing), can be highly variable. Thus, decisions  made through data mining need to be made carefully, and more than one analytical technique and set  of algorithms should be used.


References

Esling, P., & Agon, C. (2012). Time-Series Data Mining. ACM Computing Surveys, 45(1), 12:1-12:34.

Mala, A. A., & Dhanaseelan, F. (2011). Data Stream Mining Algorithms: A Review of Issues and  Existing Approaches. International Journal On Computer Science & Engineering, 3(7), 2726-2732.

Ramageri, B. M., & Desai, B. L. (2013). Role of data mining in retail sector. International  Journal On Computer Science & Engineering, 5(1), 47-50.


Wednesday, April 10, 2013

Review of Programming ArcGIS 10.1 with Python Cookbook


What makes much of Big Data extremely useful is the ability to integrate geospatial information, especially when tracked with time. To that end, ArcGIS is a "must have", and Python is a practical language that allows one to manipulate large data sets such as those found in databases, and that gathered via data acquisition module streams.

While the "cookbook" part of the title is a bit of a misnomer, Programming ArcGIS 10.1 with Python by Eric Pimpler and published by Packt Publishing does include very a total of 75 helpful recipes presented in a logical task-oriented sequence which take advantage of ArcGIS 10 features. It's useful for entrepreneurs who are coming up with innovative data mining solutions to help organizations and individuals in decision-making in many different fields and applications.



What I find most helpful is the fact that the organization of the book takes a building block approach which is helpful for someone who may need to get started, and equally so for someone who would like to simply pinpoint and extract what they want and need.

Here are some of the useful features:

* automated map production and printing: can automate the production of map production and printing (including exporting PDFs), which is helpful when creating a set of maps or map files.

* quickly using geoprocessing tools: this is a quick way to increase functionality and power without having to do everything separately; application-level environment settings are utilized quite helpfully as well.

* creating custom tools: the example shows how to filter the data for North American wildfires -- it's a useful example; I think it might be even more helpful to list some of the common sources of data and practice importing them and working with them by developing additional custom tools.

* working with attribute and spatial queries: I think it would have been good to go into a bit more detail about how / why syntax decisions are made, and to discuss the logic, the flow, and the structure; after all, mind and the mental processes are where clean code begins and ends. That said, the section discusses how Python interprets the queries and how / when it matters where a string is placed. The examples are clear, but I always need lots of examples, so I would have welcomed even more examples, but that would perhaps confuse some users, so I concluded that the book hit the right balance.

* for the more adventurous, the book includes how to use the add-in wizard. I have always been a bit leery of add-ins, believing (perhaps superstitiously) that they will create conflicts, and unleash a small troop of gremlins. This chapter shows how / where to place an add-in in a folder that is easily discoverable by ArcGIS Desktop. This is probably the key to having the thing work, and it solves a small mystery of why add-ins sometimes do not work.

In sum, I'd like to say that I find the book to be very clear, well-organized, and helpful. It's likely to have a nice, long shelf life as well.

I posted a version of this review on Amazon on the product page. Now you know I'm "Happy with Books."

Blog Archive