Showing posts with label Data Sciences. Show all posts
Showing posts with label Data Sciences. Show all posts

Nov 22, 2019

Machine Learning (ML) helps in not wasting time in non responsive telemarketing calls

My last year post was more about how ML helped in identifying best time for telemarketing calls. This one on not wasting time in non responsive telemarketing calls.























Views expressed here are from author’s industry experience. Author trains on and blogs Machine (Deep) Learning applications; for further details, he will be available at mavuluri.pradeep@gmail.com for more details. Find more about author at http://in.linkedin.com/in/pradeepmavuluri

Jul 4, 2018

Data Summary in One Go

Data Description R Code

This function and package is long pending for publishing from my side, this time expecting soon to put as package for quick usage, before that thought releasing it for feedback.

Below function provides R code for getting data description details like missing, distinct, min, max, mean, median, mode in one go for ready to use and for quick interpretation purposes.

This provide regular data summary stats needed (as shown in below image) in *.csv format which can be copied an pasted to excel as per your needs.

To use it follow the syntax:

source("https://raw.githubusercontent.com/pradeepmav/data_description_function/master/data_description.R")
data_description("datasetname")

 
Happy R Programming!


Author trains & develops Machine Learning (AI) applications, and can be reached at info@tatvaai.com or besteconometrician@gmail.com for more details.
Find more about author at http://in.linkedin.com/in/pradeepmavuluri

Mar 26, 2018

Why Record Linkage needs a scalable computing power?


Though, “Record Linkage” is a popular word among statisticians, and epidemiologists - “the problem of matching/joining records from one data source to another which describe the same entity”; has a long historical attention from the time since data collection gained (1960s) and continues to gain attention as new methods of collection, formats and stacks of data being added to the existing. The other popular terms for the same are deduplication, data matching, entity/name resolution, record matching, etc. Please, refer to the following paper https://homes.cs.washington.edu/~pedrod/papers/icdm06.pdf, for one of the good works in this field. Also, one can look at the below google trends graph for the attention to this filed from 2014 to the present.



The purpose of this blog is to bring forth, why record linkage needs a scalable computing power, for which I present my observations with an simple example as show below:



Views expressed here are from his industry experience. He can be reached at mavuluri.pradeep@gmail or besteconometrician@gmail.com for more details.

Find more about author at http://in.linkedin.com/in/pradeepmavuluri

Sep 8, 2017

Hard-nosed Indian Data Scientist Gospel Series - Part 2 : Certificate (or Degree) Mania

This is second in series, first is here.

Again, whole past decade before & after recession seems to be & seeming to be revolving around a mania called certificate or degree’s around some topic / tool. Let it be subject / concept namely., Analytics or Machine Learning or Data Science etc. and tool / technology namely., SAS or SPSS or R or Python etc. (where price of such unequal to (s) ranged from 0,000’s to 000,000’s).



This always reminded and reminds me that most of marketers duped aspirants around data science by hiding its important characteristic namely., “multi-disciplinary one”, that led to ending up with partial learning or incomplete or incompetent learning which couldn’t cater industry needs.


Author undertook several projects, courses and programs in data sciences for more than a decade, views expressed here are from his industry experience. He can be reached at mavuluri.pradeep@gmail or besteconometrician@gmail.com for more details.

Find more about author at http://in.linkedin.com/in/pradeepmavuluri

Aug 24, 2017

Hard-nosed Indian Data Scientist Gospel Series - Part 1 : Incertitude around Tools and Technologies


Before recession a commercial tool was popular in the country, hence, uncertainty around tools and technology was not much; however, after recession, incertitude (i.e. uncertainty) around tools and technology have pre-occupied and occupying data science learning, delivery and deployment.

When python was continuing as general programming language, R was the left out best choice (became more popular with the advent of an IDE i.e. RStudio) and author still see its popularity among non-programming background (i.e. other than computer scientists) data scientists. Yet, author notices in local meet ups, panel discussions, webinars, still, a clarity on which is better from aspirants towards the data sicence as a everyday interest as shown in below image.

Author undertook several projects, courses and programs in data sciences for more than a decade, views expressed here are from his industry experience. He can be reached at mavuluri.pradeep@gmail or besteconometrician@gmail.com for more details.
Find more about author at http://in.linkedin.com/in/pradeepmavuluri