Professional Documents
Culture Documents
*
v.1
Streaming Data!
Problem
Definition and
Experimental
Design!
Evaluation!
!
Collection,
Storage and
Retrieval!
Visualization!
Analysis!
Data Curation
and Quality
Assurance!
Spreadsheets!
Common Charts!
Graphs!
Geospatial views!
Temporal views!
Spatiotemporal!
3D Views!
Spatiotemporal!
Simple/Common!
See Patterns/Trends!
Powerful analysis!
Intuitive maps!
Find patterns/trends!
Powerful analysis!
More immersive!
Intuitive / Powerful!
Linear!
Non-Linear!
Utility
Providing point estimates!
Pros
Cons
Graph!
Representation of data!
Limited inference,
computationally challenging!
Ethics. Consider:!
1.
2.
3.
4.
Tableau,
Centrifuge,
Visual
AnalyCcs
Analy(cs
Machine
learning
StaCsCcal
tools
Data Store
Oracle,
Marklogic,
YarcData,
Teradata
Ingest
TransformaCon
/
NormalizaCon
Ingest
/
Streams
Processing
Storm, Hadoop/MapReduce
CM,
Scheduling,
Monitoring
ApplicaCon
OperaCng
Systems
Computers,
Networks
AWS,
Azure,
Cloudera,
Red
Hat,
Rackspace,
vendor
specic
Infrastructure
(IaaS,
PaaS)
Data Breaches!
Data Loss!
Account Hijacking!
Insecure APIs!
Denial of Service!
Malicious Insiders!
Abuse and Nefarious Use!
Insufficient Due Diligence!
Shared Technology Issue!
Cloudera: 866-843-7207
Datastax: info@datastax.com
IBM: 800-333-6705
!License costs? O&M costs? License restrictions?! !Fast writes? Fast reads? Ease of use?!
How does this tool perform prediction?! Does the tool correct for, or enable one
!answers should list algorithms used (refer to table to correct for potential bias or
at left for descriptions/pros/cons)!
confounding variables?!
What types of data does the tool
!e.g. by introducing elements of chance or by
counting everything!
analyze?!
!e.g. structured, unstructured, hybrid (are these
right for your mission?). Petabytes of data?!
!
Supports more complex data,
Limited inference, high
Processing complex systems!
complex decisions!
supervision needed!
!
Representing highly complex,
Lower precision, seed value
Complex inference, messy data!
qualitative systems!
bias!
Probabilistic!
Team Skills!
Problem Definition!
Experimental Design!
Success/Evaluation Criteria!
Data, Curation & QA!
Solution Design!
Infrastructure!
Ingest & Storage!
Analytics!
Visualization!
Security!
Privacy & Ethics!
Budget & Schedule!
Not a linear process!
Try an agile approach!
Visualiza(on
User
Interface
Web-based
tools
!By dividing a problem into pieces and executing in parallel (e.g. MapReduce)!
!By building clever indexes of knowledge, so that you can search it quickly!
!By using high performance computers (HPCs) or other fast electronics (e.g. FPGAs, ASICs, Optics)!
!A mixture of the above (e.g. Netezza, YarcData, Next Generation Oracle)!
Pros
Used For
COTS
Examples
Utility
Stack Element
Open Source
Examples
Type
Data Science
etc.
Structured Data!
Data Visualization
Natural Language
Processing
Machine Learning
Data Mining
Statistical Analysis
Many definitions!
And by using:!
!Risk Management: ICD 503!
!Access controls, IDAM, biometrics, PKI,
physical security, cell-level security, smart
data, encryption!
!CND, Anti-Malware, anti-virus!
*Please email Dan Law if you have comments or suggest any improvements to this poster
Koverse: info@koverse.com
MarkLogic: info@marklogic.com
MongoDB: 866-237-8815
Oracle: 800-633-0738
Splunk: 866-438-7758
Yarcdata: 925-264-4700