Friday, June 20, 2025

Digital Transformation: The Emperor’s New Clothes (Another Tale of Corporate Self-Deception)



“Whenever you find yourself on the side of the majority, it is time to pause and reflect.” — often attributed to Mark Twain, and definitely ignored by every boardroom chanting “Digital Transformation!”

Corporations Love Their Fancy Dress-Up Games

Remember being a kid and shouting, “The emperor has no clothes!”? Fast-forward to today and you’ll find grown-ups in expensive suits playing the same game. Only now the tailor calls the outfit Digital Transformation and charges seven figures for the fitting.

Evacuate the Building (Consultants Approaching)

First tip: post a lookout in reception. If someone whispers “DT” or flashes a Bain or McKinsey business card, activate the sprinkler system and escort them to a safe distance—say, five miles from any impressionable employee. Your culture will thank you.

Clark Kent, Phone Booths, and Other Myths

Technology is not Clark Kent sprinting into a phone booth to emerge as Superman. It’s closer to a marathon runner who never stops—always improving pace, occasionally tripping on a shoelace, but relentlessly moving forward. Calling this perpetual motion a transformation implies there’s a finish line. Spoiler: there isn’t.

Treating “digital” like a one-time renovation betrays a deeper issue: confusing tactics for strategy. Strategy, like technology, is alive, breathing, and evolving. If your PowerPoint declares, “Complete Digital Transformation by Q4 2026,” your actual strategy is, “Hope nothing important changes before then.” Good luck with that.

Tech Evolution Beats Tech Metamorphosis

Let’s retire transformation and embrace Tech Evolution instead. Evolution doesn’t hand you a certificate when you’re “done”; it keeps you alive and competitive. It also forces every tech decision to connect directly to measurable business goals—not to vendor roadmaps or Gartner hype cycles.

Need proof that buzzwords fossilize? Ask anyone who sank millions into an Enterprise Data Warehouse circa 2005, only to discover that business questions outpaced the schema changes. Many of those warehouses are now museum pieces—right next to the T-1 modem and the BlackBerry Pearl.

The Sensible Alternative (It’s Boring—That’s Why It Works)

  1. Set business strategy first. Pretend technology doesn’t exist (unless technology is your product). Clarify the value you deliver and to whom.
  2. Identify friction. Where does current tech slow, block, or bend that strategy?
  3. Apply targeted tech changes. Iterate, measure, and repeat—continuously.
  4. Celebrate quietly and keep going. Evolution is less glamorous than Transformation, but it costs less, delivers more, and won’t leave you shivering when the buzzword blanket slips.

Ten years from now, the phrase “Digital Transformation” will feel as dated as dial-up. Your CFO will glance at the depreciation schedule and wonder why anyone believed a big-bang makeover could replace steady, purposeful progress.

So skip the mythical makeover, embrace perpetual motion, and let your technology evolve at the speed of your business—not at the speed of the latest buzzword.

Monday, June 09, 2025

Startups, Vibe Coding, and the Unexpected Lifeline of AI



Startups, Vibe Coding, and the Unexpected Lifeline of AI

by William McCann

These are tough times to launch a startup.

Founders today are navigating a perfect storm—rising costs, political uncertainty, a tightening economy, and a venture landscape increasingly dominated by AI moonshots. The funding well is deep, but the water’s being siphoned by billion-dollar models.

Ironically, it’s AI that may end up saving you.

Over the past two years, AI-assisted development has reshaped what’s possible. Tools that didn’t exist 24 months ago are now making it feasible for lean teams—or even solo devs—to build in weeks what used to take quarters.

I saw it firsthand just last week. In an unscripted demo during a talk, I used bolt.new to generate a working multiplayer chess game—with a twist: every three turns, a random piece vanishes. That prototype was up and running in minutes. A year ago, that would have taken a team and a sprint.

The tech community calls this Vibe Coding. The name is catchy, but the shift is real. It’s the practice of building software through a fluid, back-and-forth interaction with AI tools. Sometimes it’s fully prompt-driven: “Build me a SaaS app for managing studio rentals.” Other times it’s more like having a tireless junior dev on standby: “Refactor this API,” “Write a test for this edge case,” “What’s the time complexity of this loop?”

It’s less like writing formal specs and more like a conversation. You describe what you want, and the code follows. Vibe coders are more like technical product managers than pure coders.

Of course, these tools aren’t magic. They require a skilled pilot—someone who can steer the AI, spot its blind spots, and shape its output into something robust, secure, and scalable. AI doesn’t replace expertise—it amplifies it.

The spectrum of tools is expanding fast. Builders now have access to platforms that scaffold entire apps from scratch, copilots that autocomplete entire functions, bots that write documentation, and assistants that lint, test, debug, and optimize in real-time.

As a founder, this isn’t just an interesting shift. It’s strategic. Your tech partner—whether a co-founder, an early hire, or a contractor—can now be 10x more productive if they’ve embraced the vibe.

The productivity gap between a developer who uses AI and one who doesn’t is no longer incremental—it’s a chasm. One is writing code with hand tools. The other is operating a semi-automated factory. Both can build a house, but not on the same budget or timeline.

So here’s my advice:

Find the partner who’s already vibe coding.
Someone who’s using these tools fluently. Who’s not afraid of AI, but curious about what it unlocks. Who understands that building fast isn’t the same as building sloppy—and knows how to do both when needed.

In this era of constrained resources, that kind of leverage isn’t optional.
It’s existential.

Sunday, September 13, 2020

Your Hiring Process is Broken



Hiring staff is hard. Hiring software developers is particularly challenging. Despite our best efforts, positions are frequently left open for weeks and months. Teams spend countless hours screening resumes, giving interviews, and reviewing coding samples. The process is time consuming and expensive. And here's the tough part to admit: it is practically certain that the person you hired was not the best candidate to come across your desk.

Your hiring process is broken.

The problem is we are chasing unicorns. When I was working with PowerInbox, we had adopted a widely published recruiting process called WHO. The process is built on the premise of finding and hiring "A Players". But I know first hand that the company did not even attract "A Players" let alone hire them. In tech, the top talent go to top tier firms, or here in New York top talent goes to Wall Street. That means other companies are looking for an ideal employee that they will never find and in the process will pass over very talented candidates.

But chasing unicorns is just part of the problem. It is likely that your process involves too many people, and gives too many of them veto power over candidates. I myself am guilty of setting up a hiring process with too many interviewers. The idea is to give the current team ownership on the hire. Unfortunately the more people involved in a hiring decision, the more likely a candidate is rejected. At issue is basic human nature; it is easier to rule someone out than rule someone in.

There is a conventional thinking that fuels these practices. The thinking is that it is very expensive and disruptive to make a "bad hire" and therefore every effort should be made to prevent bad hires. While I would never suggest that companies hire unqualified candidates, it is apparent that the typical hiring process goes well beyond the point of marginal return. In fact, I will put forth that many "bad hires" weren't bad hires at all, but instead are examples of bad management.

It's not hopeless. You can fix this.

Empower your hiring manager. Hiring is not a task suited for committees. And yet that is how most companies hire people. When I interviewed with Amazon, I spoke with a dozen of their team. In my experience it is common for candidates to speak with at least six individuals. It is very difficult for large hiring teams to come to consensus on a given candidate, especially when there are many choices. It is common for "analysis paralysis" to result. Interview by committee is also very taxing on both the company and the candidates.

Set a limit on resumes. Be realistic. If you receive one hundred applications for a senior software engineer, what is the probability that one of those candidates will be a productive employee for you. It's a near certainty, and yet we are tempted to keep screening resumes because someone else could be better. Someone else may be better, but we will never know. You will never know how any of the candidates you pass over will compare to the person you hired. So you must take that fear out of the equation. Instead of leaving a position open until the right person comes along, set a fixed size of the candidate pool that you will consider and stop accepting applications when that number is reached.

Narrow your finalists based on objective skills. A typical hiring process starts with a recruiter spending a half hour on the phone with candidates that look good on paper. If you limited your resumes to 100 you will likely want to have screening calls with at least the top twenty prospects. That's ten hours spent on the phone, and nearly as much time to summarize and rank the candidates. And it's not time well spent. Instead, test the individuals on the key skills first. The test should be difficult enough to truly separate the skilled from the un-skilled.

Here's the advantage of testing first: you will have a high degree of confidence in every candidate you speak with. Take the top ten and schedule a limited set of conversations with each of them (as described in the next paragraph).

Limit the number of conversations. You would think that the more time spent with the candidate, by more people, would supply better information. In the book Talking to Strangers Malcom Galdwell highlights research that demonstrates more time does not equal better judgement. On the contrary, more time with the candidate will often result in worse decisions. My suggestion is to hold no more than four conversations. One with the recruiter. One with the hiring manager. One with a peer, and one with a subordinate. If the job is not managerial, then hold two peer conversations.

Run a "Democratic Lottery". The democratic lottery is another concept borrowed from the research of Malcom Gladwell (and described here). For your process, let the hiring manager narrow the pool down to three or four candidates and then randomly pick one. That's right. Randomly select the finalist.

While a lottery is not a perfect system, it has advantages over the typical hiring process. First, employees selected by lottery would be more representative of the population as a whole, resulting in a diverse workforce. Second, you are unlikely to pass-over highly qualified candidates for those that "feel right". Finally, your process is faster, less expensive, and far less stressful on the organization.

So here it is. Review a specific number of resumes. Test their skills first. Hold a small number conversations with your top ten. Have the hiring manager pick three or four. Randomly select the finalist. I assure you that your hiring will be faster and your new hires as good or better than before.

 

 

Thursday, May 28, 2020

Using React Hooks with Firebase Authentication


If you are a React developer, and unless you have been living under a rock for the past year, you will be familiar with Context, Reducers, and Hooks. Sure, these are usually just called "Hooks" or "custom hooks". Hooks give the developer two advantages over the old-school method of class-based components. These are:

  • Ability to build an entire application using function components, and
  • Elimination of the need for Redux.

Functional components make the code more concise and readable. They also provide a fairly significant performance boost over classes. As for Redux. Well. Good riddance.

Armed with hooks I decided that my next React project would be written entirely with functions and hooks. And this approach worked great until I needed to wire up Firebase Authentication. I discovered that the tutorials available provided examples using class-based components. This article then, is my attempt to fill that gap with my notes on using Hooks with Firebase Authentication.

Here are the things I had to build.

  • A "context" with a "context provider",
  • A "reducer",
  • A custom hook that uses an Effect hook,
  • Helper functions to interact with Firebase, and
  • All the JSX to make it work.

The Reducer

Let's start by writing a Reducer. That might seem a bit backward as I typically begin with my Context. But since the Reducer is referenced by the Context, I'll start with the Reducer. The reducer is used to make changes to the state of the properties managed by the Context. If all this sounds like jargon (which it is), then I suggest reading up on React Hooks; especially useContext, useReducer, and useEffect. I'd love to explain all that here, but well, this article will be very long as it is.

The Reducer and the Context will be supporting Firebase Authentication. I should mention here that I am using the Authentication SDK and not the drop-in UI. I should also mention that I am only describing email based authentication. The properties of a "user" in Firebase are:

  • Name, which is the "display name" of the user. The display name is usually the user's full name.
  • Email, an email address the user is treating as her id.
  • PhotoURL, a URL to an image the user has uploaded for their avatar.
  • EmailVerified, a boolean value that indicates the user responded to a verification message.
  • Uid, which is a global unique identifier assigned by Firebase to the user.

Reducers contain logic for adding, removing, or manipulating the Context's data. In this case though, the user data is managed by Firebase via its' SDK. All our reducer needs to do is assure that the context is current. I created a file named SessionReducer.js and included the code below.

    export const SessionReducer = (state, action) => {
        switch (action.type) {
          case "UPDATE":
            return {
              name: action.session.name,
              email: action.session.email,
              photourl: action.session.photourl,
              emailVerified: action.session.emailVerified,
              uid: action.session.uid
            };
          default:
            return state;
        }
      };

The action.type property is part of the Reducer specification. So the Reducer is passed the state and an action object and returns the new value of the state. In this case, we will only take one action that I have set to "UPDATE". The value UPDATE, by-the-way, is a discresionary name set by the developer. Also note that I do not need any import statements here.

The Context

The Context is a bit bigger and contains a couple of functions. I'll show the code and then describe it. I created a file named SessionContext.js and included the code below.

    // React imports
    import React, { createContext, useReducer, useContext, useEffect } from "react";
    
    // Firebase imports
    import firebase from "../firebase";
    
    // My imports
    import { SessionReducer } from "../reducers/SessionReducer";
    
    // initial state values
    const initialState = {
      name: null,
      email: null,
      photourl: null,
      emailVerified: false,
      uid: null
    };
    
    // create the context
    export const SessionContext = createContext();
    
    // create the context provider
    const SessionContextProvider = props => {
      const [session, dispatch] = useReducer(SessionReducer, initialState);
      return (
        
          {props.children}
        
      );
    };
    
    // create the custom hook
    export const useSession = () => {
      const contextState = useContext(SessionContext);
      const { dispatch } = contextState;
      useEffect(() => {
        firebase.auth().onAuthStateChanged(user => {
          var currentUser = {};
          if (user) {
            currentUser = {
              name: user.displayName,
              email: user.email,
              photourl: user.photoURL,
              emailVerified: user.emailVerified,
              uid: user.uid
            };
          } else {
            currentUser = initialState;
          }
          dispatch({
            type: "UPDATE",
            session: currentUser
          });
        });
      }, [dispatch]);
      return contextState;
    };
    
    export default SessionContextProvider;    

Notice near the top of the code I have included an import of the Reducer shown earlier. I have not described the structure of my application but suffice to say that my Contexts and Reducers each reside in folders made for that purpose. You can organize your source files any way you like.

Following the imports, I create and export the SessionContext using React's API for that purpose. It's one line of code that requires no parameters. Next I have a few lines of code to create the "Context Provider". This is a small amount of code that includes JSX that ties the context to components in the application. React has two methods for this: Provider and Consumer, but since I am using functional components exclusively I must use the Provider method.

The Provider binds the Reducer to the Context. In this code I use the useReducer function to deconstruct its' return value into the "session" and the "dispatch". The session is the current state and the dispatch is the function I supplied when I wrote the Reducer. These values are passed as props (via HTML attributes) to the provider context. The {props.children} value assures that this context provider will be available to all children components to the component where it is used. I should also mention that the provider context is just another component. In my case, the context provider is my default exported function.

The Custom Hook

The last function creates the custom hook. By convention the names of hooks start with "use" in lower case; i.e. useContext, useReducer, or in my case useSession. The custom hook first retrieves data from this context, which is the state and the dispatch function, and stores it in contextState. The dispatch function is deconstructed out of the contextState. For my purpose here, I will not need the actual values of the state.

The work of this custom hook is accomplished within a useEffect function. If you’re familiar with React class lifecycle methods, you can think of useEffect Hook as componentDidMount, componentDidUpdate , and componentWillUnmount combined. In this case, whenever a component that uses our hook is rendered, a call is made to the Firebase onAuthStateChanged method. That method is passed a function that checks the state of the current Firebase user. If the user exists then the Firebase used is passed to the dispatch function, otherwise the initial state of the user (which is null) is passed to dispatch.

Wiring it up

The rest is simply writing the custom hook into the app. In this case I wanted to make the hook available to the entire application. To accomplish this I added my context provider to the app.js main code file.

    // React imports
    import React from "react";
    import { BrowserRouter, Switch, Route } from "react-router-dom";
    
    // Material UI imports
    import { MuiThemeProvider, useTheme } from "@material-ui/core/styles";
    import CssBaseline from "@material-ui/core/CssBaseline";
    import Box from "@material-ui/core/Box";
    import Container from "@material-ui/core/Container";
    
    // My imports
    import Dashboard from "./components/dashboard/Dashboard";
    import AppHeader from "./components/layout/AppHeader";
    import HomeDetail from "./components/houses/HomeDetail";
    import SignIn from "./components/auth/SignIn";
    import SignUp from "./components/auth/SignUp";
    import HouseContextProvider from "./contexts/HouseContext";
    import SessionContextProvider from "./contexts/SessionContext";
    import LaunchPage from "./components/dashboard/LaunchPage";
    
    const App = () => {
      const theme = useTheme();
    
      return (
        <BrowserRouter>
          <SessionContextProvider>
            <MuiThemeProvider theme={theme}>
              <HouseContextProvider>
                <CssBaseline />
                <AppHeader />
                <Container maxWidth="xl">
                  <Box m={3}>
                    <Switch>
                      <Route path="/" exact component={LaunchPage} />
                      <Route path="/dashboard" component={Dashboard} />
                      <Route path="/homes/:id" component={HomeDetail} />
                      <Route path="/signin" component={SignIn} />
                      <Route path="/signup" component={SignUp} />
                    </Switch>
                  </Box>
                </Container>
              </HouseContextProvider>
            </MuiThemeProvider>
          </SessionContextProvider>
        </BrowserRouter>
      );
    };
    
    export default App;

In this file there are just a couple of lines to point out. First is the import of the SessionContextProvider. And the others are the JSX SessionContextProvider tags.

To signup a new user bind the function below to your form submission handler.

...
import { signup } from "../../services/firebaseAuth";
...
// form actions
const handleSubmit = e => {
  e.preventDefault();
  enroll();
};

// firebase signup function
const enroll = () => {
  const results = validate(
    {
      email: email,
      password: password,
      firstName: firstName,
      lastName: lastName,
      confirmedPwd: confirmedPwd
    },
    {
      email: constraints.email,
      password: constraints.password,
      firstName: constraints.firstName,
      lastName: constraints.lastName,
      confirmedPwd: constraints.confirmedPwd
    }
  );

  if (results) {
    setErrors(results);
  } else {
    setErrors(null);
    signup(email, password);
  }
};

A couple of points about the code above. First, the validate function it its' constraints are out of the scope of this article. If anyone reads this, and someone asks about validate then I will consider writing another post to describe that function. And second, there is a helper function signup that I show a bit later.

Signing in is similar.

...
import { SessionContext, useSession } from "../../contexts/SessionContext";
import { signin } from "../../services/firebaseAuth";
    
const SignIn = () => {
const [email, setEmail] = useState("");
const [password, setPassword] = useState("");
const [errors, setErrors] = useState();
const { session } = useSession(SessionContext);
const handleSubmit = e => {
  e.preventDefault();
  authenticate();
};

// firebase signin function
const authenticate = () => {
const results = validate(
    {
      email: email,
      password: password
    },
    {
      email: constraints.email,
      password: constraints.password
    }
  );
 
  if (results) {
    setErrors(results);
  } else {
    setErrors(null);
    signin(email, password);
  }
};
...

Again, the signin function is shown later. And lastly I sign out with a link button on my AppBar that executes the signout fuction (and yes, the function must be imported).

    ...
    <button color="inherit" component="{Link}" onclick="{signout}" to="/">
        Sign Out
    </button>
    ...

Firebase

The final pieces of the puzzle are the supporting functions that call out to Firebase. These are the signup, signin, and signout functions mentioned above. I organized the functions into a single source code file. For purposes of this paper, these functions were copied directly from the Firebase documentation website and do not incude any of my application specific code. I created the file firebaseAuth.js for this. The code is below.

    // Firebase imports
    import firebase from "../firebase";
    
    export const signup = (email, password) => {
      firebase
        .auth()
        .createUserWithEmailAndPassword(email, password)
        .catch(error => {
          // Handle Errors here.
          alert("Error during sign up " + error.message); // delete this!
          var errorCode = error.code;
          var errorMessage = error.message;
          if (errorCode === "auth/weak-password") {
            alert("The password is too weak.");
          } else {
            alert(errorMessage);
          }
          console.log(error);
        });
    };
    
    export const signin = (email, password) => {
      firebase
        .auth()
        .signInWithEmailAndPassword(email, password)
        .catch(function(error) {
          // Handle Errors here.
          var errorCode = error.code;
          var errorMessage = error.message;
          if (errorCode === "auth/wrong-password") {
            alert("Wrong password.");
          } else {
            alert(errorMessage);
          }
          console.log(error);
        });
    };
    
    export const signout = () => {
      firebase
        .auth()
        .signOut()
        .then(function() {
          // Sign-out successful.
        })
        .catch(function(error) {
          console.log(error);
        });
    };

So the next person who needs to implement Firebase Authentication with React Hooks now has some reference material to start with. Questions and comments are welcome.

Thursday, May 21, 2020

Talking To Strangers



Don't judge a book by it's cover. Or, not by its' title either. Ok, I was guilty of both, so this book was not what I expected. I should mention that I am not a great conversationalist and often struggle to make small talk with people I meet. I am familiar with Malcom Gladwell and have read a handful of his books. I thought this book would provide the data and anecdotal based advice I could use to improve my ability to talk with others. That's not what I got.

 

I should be clear, Talking to Strangers is a great book. I would highly recommend it, especially for people who wish to gain insight into why communication between people breaks down. More specially, how communication between strangers breaks down ... tragically. The research is bookended by the story of Sandra Bland and discusses Bernie Madoff, Neville Chamberlain, Sylvia Plath, and Amanda Knox along the way.

 

Gladwell books have a common tone. The books hold advice, but they are not self-help tomes. He chooses his anecdotes carefully and backs them up with other data and research. I like that he names names. Books of this type often refer to "a successful fortune 500 company" or "a national leader" without indicating who it is. When that happens I am inclined to think one of two things … the material is fiction or the subject does not agree with the assessment. That does not happen here. In fact, if you listen via Audible, you will hear actual recording from many of the subjects.

 

So I did not come away better prepared to strike up a conversation. But I learned a lot about communication regardless.


https://www.amazon.com/Talking-to-Strangers-audiobook/dp/B07NJCG1XS/ref=sr_1_1?crid=TEZVMYU97T6I&dchild=1&keywords=talking+with+strangers+malcolm+gladwell&qid=1590086680&sprefix=talking+wi%2Caps%2C148&sr=8-1


Saturday, May 13, 2017

Why Data Science Can't Find the Needle in the Haystack

A colleague of mine wants a predictive model. He is trying to determine which people on a health insurance plan will visit the hospital in the next couple of months. He has pretty good data for making this type of prediction. He knows who visited the hospital in the past; and he knows they are more likely to revisit. He knows what illnesses these people have, and which illnesses likely result in hospital visits. He knows what drugs have been prescribed, and whether patients are taking their drugs.
Even with this data and even with very good models, he still complains that the predictions are not good enough. His problem is too many false positives. And he simply doesn’t have enough employees to review every patient the model predicts.

Data science is a great tool, but it is not perfect. And if you are going to weld the data science tool, you should be aware its’ shortcomings. Data science is simply not very good at finding a needle in a haystack.

This concept can be illustrated with an example. Say a banker is managing the mortgages of 500,000 homeowners. He knows from experience that roughly 1,000 of these homeowners will default on their loan. There is plenty of data to help zero in on these 1,000 people: zip code, income, payment history, and credit rating. He knows that if I can put the right people some assistance, they may not default on their loan.

We have the data of build a predictive model. But regardless of how good the model is, it will not be perfect. When the model is run, each homeowner will be classified as at-risk for default or not at-risk. In this scenario, there are four possible outcomes for each homeowner. The homeowner is at-risk and is properly identified by the model; the homeowner is not at-risk is properly identified by the model. These are the two accurate predictions.

Every predictor gets some wrong too. When the homeowner is not at-risk, but the model says he is, that is a false positive. If the homeowner is at-risk and was not identified by the model, that is a false negative. In statistics, false positives are also referred to as “type I errors”. False negatives are referred to as “type II errors”.

Now let’s say that we construct a model that is 80% accurate, which is a rule-of-thumb threshold for a good prediction. With 80% accuracy on 500,000 loans, 400,00 will be correctly predicted and 100,000 will not. At an 80% prediction rate, 800 of the 1,000 homeowners would be predicted correctly.

Of course, an 80% success rate means there is also a 20% failure rate. I mentioned that 100,000 are not predicted correctly. There are 200 false negatives that are the balance of the 1,000 target loans. Subtracting the 200 false negatives from the 100,000-people identified incorrectly leaves 99,800 false positives. That is the extraordinary 500 times as many false positives as correctly predicted at-risk loans.

Even if the model can predict at the incredible rate of 99% the numbers of false-positives will outnumber the correctly identified at-risk cases by nearly five to one.


The problem here isn’t with data science or prediction methods. It simple math. When trying to use statistics to find a very small number among a very large number the false-positives will always greatly outnumber the actual positive prediction. This type of problem is truly a needle in a haystack.

Wednesday, May 10, 2017

Data Science isn't Rocket Science

Science is intimidating. It conjures up images of lab coats, telescopes and microscopes, or petri dishes and chemicals. But Data Science isn’t rocket science. Science is about discovery and Data Science is discovery in data.

You don’t need an army of PhDs for successful data science. Instead you need a basic understanding of statistics and a little computing power. Then follow this recipe of five steps to put Data Science to work for your business.

Step 1. Decide what to predict

First and foremost, you must know what you want to predict. Data Science in business is about making predictions. That predicting might be finding customers that will buy a product; or which patients will be readmitted to a hospital; or when to buy shares of stock.

Everyone knows the story of Target sending coupons for pre-natal items to a teenage girl, only to surprise her father. But that case didn’t happen by chance. Instead, someone at Target decided to specifically focus on pregnant women. That person decided to predict which of their customers were pregnant.

Make your prediction on a single thing. That thing will be represented by a prediction variable. The variable might simply be “Yes” or “No”, such as in the Target pregnancy example. It could be an item in a list, such as a day of the week. It frequently will be a number, such as the price of a barrel of crude oil.

Step 2. Make your hypothesis

Your hypothesis is the key to the puzzle. And yes, the word hypothesis comes right out of the scientific method because that’s what we’re doing; we’re applying the scientific method to data to make predictions. You decided in the previous step what to predict, now it’s time to guess how to make that prediction.

Some insight into your problem is helpful here. For example, in the case of Target above, someone had made the presumption that items a person purchased could indicate whether they are pregnant. They very likely narrowed the items down to a specific list of items or types of items. The hypothesis may have been as simple as “someone who buys pre-natal vitamins is likely to be pregnant”. Or even more specific, “a woman between 18 and 45 who buys pre-natal vitamins is likely to be pregnant”.

Step 3. Get the data

Of course, this whole exercise assumes the data is available to make these predictions. You will need to collect data points on any attribute you are testing, as well as the values you wish to predict.
Again, going back to our example to finding customers who are expecting. Assuming we want to build a model on the hypothesis that “a woman between 18 and 45 who buys pre-natal vitamins is likely to be pregnant”. We will need the following data:
  • ·        A list of customers who bought products.
  • ·         A list of the products they bought.
  • ·         For each customer, we need their age and gender.
  • ·         For the products, we need to know which are pre-natal vitamins.
  • ·         And most importantly, we need to know who is pregnant.

In this age of big data, you may have all the data available. That is, every customer, and every product they purchased. If you have all the data, great, you can build your model based on the population, which is the statistician’s way of saying “all the data”. Otherwise we will use a sample, which is a way of saying some of the data.

Some data points may be difficult to obtain. In our example, the fact that a customer is expecting a child may not be readily available. In this case, special steps will be required to obtain the data. Target may have performed a customer survey or used some other means of gathering the information directly from the individual.

Step 4. Build a model

Building a model is where the fun starts. This is the statistical model, or predictive model. A basic understanding of statistics and knowledge of modeling software is necessary.

Before building a model, you should run some analysis to see if your hypothesis is worth pursuing. The typical first step is testing a null hypothesis for statistical significance. The null hypothesis checks that the predictor variable affects the prediction. In our case, the null hypothesis would be “knowing a person is a woman, between 18 and 45, who bought pre-natal vitamins has no impact on their being pregnant.”

In short, we compare the number of pregnancies in a random sample of customers against 18 to 45-year-old women buying pre-natal vitamins. If the difference in pregnancy rates between these two groups is greater than 5%, the null hypothesis is disproved, and our assumption is considered statistically significant.

Some caution is needed here, because a 5% difference could be attributed to improbable random samples. You can protect yourself against improbable results by repeating the test against additional random samples.

Once you have your data. And your hypothesis is sound. Building a predictive model is relatively simple. For example, the R code for creating a model looks like this:

modelFit = train(class ~ .,method="rf",
data=trainingCV, prox=TRUE)

And the code for making predictions looks like this:

prediction = predict(modelFit, testCV)

This code is illustrative to show that modelling does not require complicated commands.

Step 5. Check your results

Before creating your model, divide your data into two sets; a training set and a test set. The training data is used to create the model. The test set is used to demonstrate that it works. Typically, you would split the original data such that 80% of it is used for training with the remaining 20% used for testing.

In our sample, let’s say we have 1,000 customers in our data. We would randomly select 800 for training and 200 for testing. But there is nothing sacred about an 80/20 split. In fact, if your dataset is very large, say 100,000 or more, you could create multiple test sets using a 60/20/20 split.

When your data is prepared, create your model; if using R then run the train() function. Then take the created model and run it against the test data; if using R then run the predict() function. The predict() function will make prediction of whether the customer is pregnant or not.

The results of the prediction are compared against the original data to determine if our model makes reliable predictions. Using our sample, we would make predictions against 200 people. If the model correctly determines pregnancy in 150 of them, then our prediction was 75% accurate.

Your software should be able to provide you with an Area Under the Curve (AUC) analysis. AUC values close to 1 indicate a very good model, those near .5 are little better than flipping a coin. You should strive for AUCs of .80 or more.

And now we’re done

Well, maybe we’re not done. If you’re AUC is poor, then you should start the process over. But it’s not a complete loss, because even a bad model gives knowledge; knowing what doesn’t work is important too.


Of course this is a blog post. And I have over-simplified every step. Still, creating predictive model is not as intimidating as one might think. It’s most definitely not like putting a man on the moon.

Tuesday, May 02, 2017

Beware of Expert Blindness

More and more companies are looking to big data to help them market their products or improve their services. Of course that means more companies are seeking out data scientists and statisticians. But to truly take advantage of big data means the firm must commit to the principles of data science. That is often easier said than done.

Enter the Subject Matter Expert and the common trap of “Subject Expert Blindness”.
Yes. The Subject Matter Expert; the person who has spent a career building knowledge of their business. These are the people who drive a company’s offering; or whose stamp of approval is necessary on any significant project. They believe their experience and learning has given them special insight that others simply do not have.

If you are a specialist in data science, then it is unlikely that you have spent years earning experience in any particular industry. Automotive. Healthcare. Insurance. Finance. It doesn’t matter because your expertise is data. Data is data. And you tell your story with the data.

The expert does not rely on data. Or only needs it to confirm their preconceived insight. The expert, then, becomes blind to alternatives hidden in the data.

Take the case of a recent project of mine. I was approached by a firm looking to find groups of people who were likely to be the most expensive customers to service. The expert provided a list of twenty such groups and asked that we demonstrate that these are statistically more likely to consume services than an ”average” customer.

But the notion of pre-determined groups is silly in the world of data science. Why not run have the data tell us what the highest risk groups are? If the results match the expert’s groups then great, her hunches are confirmed. But the expert will never find the hidden gems that the data often exposes. The expert is simply blind to the alternatives.

The bottom-line: let the data tell the story. Don’t force the story onto the data. Resist the temptation to rely on personal experience to shape the story before the data is even crunched.

Tuesday, January 20, 2015

Ask them “What do you want to learn?”

You hear the question all the time, "What do you want to be?" I'm even guilty of asking it myself. This is how we start advising our youth when they are considering colleges.

It's the wrong question. And it's part of a pervasive thinking causes kids to spin through multiple majors and spend more time in school than is necessary. We are programmed to think of university study as job training. It's not. And if you think I'm wrong, ask the most successful people you know if they are working in their field of study (very possibly not). Then ask them if college was a waste of time (most definitely not).

Instead, college is where we go to broaden our knowledge. It's where we sharpen communication skills. It's where we learn how to work independently; it's where we learn to work with others (and no, those aren't mutually exclusive). It's an opportunity to explore topics in depth because it interests us, rather than because we have to. And more importantly, it shows future employers that we can set a long range goal, work hard, and finish it successfully.

The right question then, is "what do you want to learn?" If the person already knows what she wants to be, then she probably already knows what she wants to learn. More importantly, though, if the thought of learning a subject is distasteful, then that career choice is not wise.

Now here's the tricky part for us adults giving guidance: how do we respond when the young man answers our question with "Literature" or "Philosophy"? Typically, the thought is "what kind of job can you get with that?" That is wrong thinking.

Literature? What business or agency couldn't benefit from a person who has deep knowledge of communication?
Philosophy? What business or agency couldn't benefit from a person who has deep understanding of how people are motivated?

Most importantly, though, when a student considers what they want to learn; And when they spend time exploring that subject; they will find ways to apply that knowledge to other areas. Their new knowledge will guide them into an appropriate career. So as crazy as it may seem, the study of the oceans can help a person with a later career in sales. Or the study of music can help enrich one's later family life.

So next time you are congratulating a high school senior on their graduation, ask the right question. Ask them, "so now what do you want to learn."

Tuesday, June 17, 2014

It’s education, not job training.

Several years ago I found myself behind a podium in a packed gymnasium. Hundreds of people stared up at me and yet the cavernous room was almost silent. Half the crowd were high school graduates; the others their proud parents and families. I was giving the commencement address. Next to me, on the floor, I had a bag of props; a vinyl record album, Michael Jackson. A compact disk, Pearl Jam. A cell phone.

The topic of my speech was how fast the world changes. And as examples I displayed my props and explained how that, in only the time the students had been going to school, changing technology dramatically affected our lives. The tie-in to the day, the commencement, was that their education needed make them ready for this changing technology, rather than the common notion of learning the current technology itself. Because, as they could already see, their world was evolving fast.

It's one of the most common mistakes made by people in general; the mistaken belief that the purpose of schooling is to provide job training. It's why graduates are rarely asked what they want to study, instead they are asked what they want to be (or do). It's why philosophy majors are asked "what are you going to do with that?"

While this may seem very subtle, instead it's a huge leap to understand the greater purpose of education is to prepare us to have careers; not train us in a career. Education, especially college degrees, demonstrate the abilities to apply oneself toward a goal that takes years to achieve; often without direct guidance or intervention from parents. Imagine you are hiring for an entry level marketing position in an established corporation. All things being equal, would you pick an English student who finished in four years, or a business student who took six years and changed majors? Obviously it's not the material studied, but the work habits formed that is the important learning.

The thing is, careers (and life in general) require a strong foundation of skills that everyone needs to learn. Basic math. Proper writing. Fundamentals of science. Even history and art. You can even make an argument that it's important to understand demand curves and probabilities and logic. Providing these kinds skills should be the goal of every educational system. Then freed of the stress of "what job will I find", students can apply themselves in areas of interest. Then those that have an interest in health can learn to be doctors; and those that like to build things can learn to be engineers.

So this summer, when you are making small talk with recent graduates, don't ask them what they intend to do; instead ask them what they want to study. Then remind them that the real learning is finding their way to the finish.

Tuesday, June 26, 2012

What I am reading | "Savages"

Savages: A Novel
Don Winslow


Nothing makes a long flight more tolerable than a good read. So long along those lines, I found myself in the Phoenix airport looking for a novel to entertain me on my trip back to New Jersey. I settled on Savages because it looked like an engaging quick read.

Anyway, the book is certainly a quick read. Honestly, though, this is probably the dumbest story that I've read since, well, maybe since ever. The characters are all stereotypes. The plot is cliche'. The prose is minimalist. It felt like Winslow tried to channel Cormac McCarthy, but the writing doesn't engage in the same way. I felt I was reading a story written by a sixth grader. Speaking of which, Winslow's poetic license of English grammar is worse than a sixth grader; the writing really has no form.

After reading an action passage, my son summed it up nicely, "wow, he made that explosion sound boring." Oops.

Monday, May 14, 2012

What I am reading | "Ultra-Marathon Man"

Ultramarathon Man: Confessions of an All-Night Runner
Dean Karnazes



I finished this book recently. It's a pretty entertaining read. While Dean is obviously not a writer by trade, his stories draw you in. While I have no desire to run ultra-marathons, the stories inspire me to continue my running. I found the marathon to the South Pole to be particularly interesting. In any case, this is an easy read; I recommend it for runners at all levels.


Monday, May 07, 2012

What I am reading | "Don't Make Me Think"

Don't Make Me Think! A Common Sense Approach to Web Usability
Steve Krug


I'm slogging my way through this guide to designing easy to use web sites. I had bought it with the idea that it would provide insight into good (web) application design. However, after reading a couple of chapters, it is apparent that web sites, those that provide content, and web applications, those that do something, are different animals.

The book is fine, and is a recommended read for anyone new to web design. Unfortunately for me, Mr. Krug hasn't provided any new ideas that I haven't seen, tried, or evangelized at some point. It's easy to read with good examples; follow the guidelines in the book and your web site won't suck.



Thursday, May 03, 2012

Time to Kill the Relational Model




It’s time to kill the relational database model. Well not the relational model per se, but rather the practice of building systems by developing relational schemas first.


This is quite the revelation for me. When first I discovered relational database design over three decades ago I thought it was religion. The simplicity of tables and relationships represented freedom from the rigid hierarchical structures of the day. Since then, relational design and its’ normal form remained a constant foundation for the systems I constructed.


But during a recent project I tired of the burden that relational design placed on our development effort. In particular, a tight adherence to rules of normalization and relational philosophy stifled our opportunity for nimble software engineering. What on paper looked like a method of enforcing integrity became complex morass of surrogate keys and multiple joins.


Fundamentally my complaint is over the complexity that normalization brings; an odd turn considering that relational databases were popularized on simplicity of tables. Frankly, though, there is nothing intuitive about joining tables; and the more joins the more complexity.


While relational databases are not (and should not) go away, there are flaws with some relational theory that makes writing software difficult. For example…


  • Normalization produces a lot of tables. A lot of tables translate into a lot of joins. A lot of joins is a red-flag for complexity. When queries frequently have three or more joins, the schema is probably overly complex.
  • Relational theorists discourage use of null foreign keys. The only way to accomplish this goal is to introduce a link table for the one to many relationships in the model. This introduces two problems, first, the extra table adds an additional (and unnecessary join). And second, it is harder to determine if a row in one table is related to a row in the other.
  • Using data to determine state (or status) introduces complexity. While not necessarily a tenant of relational models, some designers prefer to compute the state of an item on-the-fly based on data in the tables. This is easy enough when all the data needed to determine the state are stored on the same table row. Unfortunately it is frequently necessary to scan additional tables on multiple rows to calculate the status.


A better method of system design starts with a design of objects. The database schema, then, is modeled to keep these objects intact. In this way, the database becomes nothing more than the repository for data at rest. The validity of the data is managed by methods within the objects and not by referential integrity or database triggers. Think of this as object first, data last design; there OFDL, I just coined a new acronym.

Friday, April 27, 2012

The Key To A Great Meeting Is Kicking Some People Out Of It | Fast Company


The small-group principle is deeply woven into the religion of simplicity. It’s key to Apple’s ongoing success and key to any organization that wants to nurture quality thinking. The idea is pretty basic: Everyone in the room should be there for a reason. There’s no such thing as a “mercy invitation.” Either you’re critical to the meeting or you’re not. It’s nothing personal, just business.
Steve Jobs actively resisted any behavior he believed representative of the way big companies think--even though Apple had been a big company for many years. He knew that small groups composed of the smartest and most creative people had propelled Apple to its amazing success, and he had no intention of ever changing that. When he called a meeting or reported to a meeting, his expectation was that everyone in the room would be an essential participant. Spectators were not welcome.
If big companies really feel compelled to put something on their walls, a better sign might read:
How to Have a Great Meeting
  1. Throw out the least necessary person at the table.
  2. Walk out of this meeting if it lasts more than 30 minutes.
  3. Do something productive today to make up for the time you spent here.


Keep reading:
The Key To A Great Meeting Is Kicking Some People Out Of It | Fast Company:

'via Blog this'

Wednesday, March 14, 2012

6 Time-Management Tips From Accelerator Programs | Fast Company

6 Time-Management Tips From Accelerator Programs | Fast Company: "


Sage advice from Alina Dizik...


1. Avoid the email time suck.
2. Choose your most important goal each week
3. Know your productivity limits.
4. Be like Dorsey: Take breaks to prevent burnout.
5. Skip some meetings.
6. Say "no" when you need to



'via Blog this'

Tuesday, March 13, 2012

Some music ages so well...




I'm up at 5:30am every morning for a 5K run. This morning my legs were stiff and my thighs burning from extra mileage I put in over the weekend. As I rounded the last bend and mustered strength for one more hill, The Who's Tommy pounded through my iPod;


Listening to you,
I get the music.
Gazing at you,
I get the heat.
Following you,
I climb the mountains.
I get excitement at your feet.

Right behind you,
I see the millions.
On you,
I see the glory.
From you,
I get opinions.
From you,
I get the story. 

The music made me come alive and I charged up the hill and sprinted to the finish as if I was on fresh legs. I was only nine years old when The Who recorded Tommy. I knew virtually nothing about it until a decade later in college. I liked it, but not enough to buy a copy. Now thirty years later, and more than forty since the album's release, I catching up with it again. Some music ages really well.

Wednesday, December 14, 2011

Microsoft Office 365 Cloud-Based Productivity Service Now Helps Customers Comply with HIPAA Privacy and Security Standards - Microsoft in Health - Site Home - MSDN Blogs

Microsoft Office 365 Cloud-Based Productivity Service Now Helps Customers Comply with HIPAA Privacy and Security Standards - Microsoft in Health - Site Home - MSDN Blogs:

With reimbursements falling and medical loss ratio minimums rising, hospitals, physicians, and health plans are under unprecedented pressure to drive down operating costs while still improving the quality and safety of patient care. The economic advantages of cloud-based productivity solutions to drive down operational costs and complexity are well understood, but for most health organizations, HIPAA security and privacy concerns have been a showstopping barrier to realizing the full anywhere, anytime productivity potential of cloud-based technologies.


That is, until now. Today ... Microsoft is helping remove that barrier by embedding privacy and security capabilities in Office 365, our next-generation cloud productivity service. This means that Office 365 is now a cloud-based platform that complies with leading information privacy and security standards for customers operating in the United States and European Union. As part of its contractual commitment to customers, Microsoft will now sign business associate agreements under the U.S.-mandated Health Insurance Portability and Accountability Act (HIPAA).


'via Blog this'

Thursday, October 27, 2011

Great anecdote from "Users, not Customers"

I just started the book "Users, not Customers: Who Really Determines the Success of Your Business." It starts with a great short anecdote about comparison shopping in our brave new world. If the rest of the book stays as intriguing, this will be a great read.


"My wife loves seltzer water. I can’t stand it, but she will hardly drink water if bubbles aren’t in it. So I thought it’d be great to buy her a soda maker. One afternoon, I passed by a Williams-Sonoma store and decided to stop in. Lo and behold, they had one sitting on the shelf: a SodaStream Genesis drinks maker for $150. But it seemed expensive. I could buy her a pantry full of 150 bottles of premade seltzer for that price. So I decided to shop around. 
I opened the RedLaser app on my iPhone and used it to scan the machine’s bar code to find out what other retailers charged. Bed Bath & Beyond carried the same thing for a hundred dollars. Success! Fifty dollars in savings. I waved down a sales clerk and showed her my findings. But she declined to match the price. 
So right there, in the middle of a beautiful Williams-Sonoma store in a high-rent location on the Upper East Side of Manhattan, I bought the SodaStream Genesis drinks maker—from Bed Bath & Beyond by using my mobile browser."


You might also like ...