Showing posts with label LabKitty primer. Show all posts
Showing posts with label LabKitty primer. Show all posts

Monday, December 4, 2017

A Primer on the General Linear Model

LabKitty Academy Greal Seal
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and mostly work only the simplest cases. I show mistakes and dead ends and things any normal person would try that don't work. Also, there's probably a swear word or two. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

Mathozoids describe the General Linear Model (GLM) as, well, a general linear model, one that combines regression and ANOVA into a unified whole. Those words aren't particularly helpful in explaining the thing, especially if you're coming to GLM for the first time. Heck, it took me a month just to understand there is also a GeneraliZED Linear Model which is something different (sort of) even though it generally goes by the same abbreviation.

But I had an epiphany the other midnight about a different way of explaining GLM, a way that has little to do with combining regression and ANOVA into a unified whole and which the beginner might find to be a more approachable inroad. At first glance, my way may seem like more work because I will ask you to think about other analysis tools that have little to do with statistics. I encourage you to resist this reaction. For we will not discuss the details of these other methods, only their character -- their "data philosophy," if you will. Once you see how GLM fits into the big picture, all else will follow. We'll construct a few simple models, consider the briefest derivation of the underlying equations, review statistical testing, and have a look at some Matlab code. As usual, I'll finish up with a few leads into the literature.

For your part, you'll need to recall a little matrix algebra and enough statistics to know what regression and ANOVA are. You'll also need to surrender your old statistics worldview for the new vista GLM offers. It's natural to resist taking such a leap -- your old worldview may have been acquired through much classroom effort and pain. However, embrace GLM and no longer will you look upon statistics as a collection of disjoint topics, each with its own ad hoc nomenclature, recipes, and purpose. Instead, look upon the One True Statistics, like some being provoked out of the absolute rock and set nameless and at no remove from its own loomings to wander the brutal wastes of Gondwanaland in a time before nomenclature was and each was all. GLM is just that cool.

As has been said of Mordor, a root canal, and puberty, the only way around it is through it. And so your journey begins now.

Tuesday, November 7, 2017

A Primer on Fourier Analysis -- Part II

LabKitty Academy Greal Seal
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and mostly work only the simplest cases. I show mistakes and dead ends and things any normal person would try that don't work. Also, there's probably a swear word or two. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

Previously on A LabKitty Primer on Fourier Analysis, we met the beast and poked it with a stick. We discovered -- glossing over the hard stuff that made Fourier famous -- a Fourier transform is just a fancy way of correlating sines and cosines with a given signal of interest.

That simple worldview works conceptually, but now we're going to look at how the pros do things. As I stated in Part I, my goal is to get you to understand the numbers Matlab's fft( ) function returns, and so we shall. We're still decomposing a signal into its frequency components. Nothing new there. However, Matlab is top-shelf engineering software and not just some lunatic ranting on the Internet. As such, there's a rash of new details we must encompass and eclipse. Some of those new details are unpleasant.

Yes, our tale today is long and fairly horrible. The good news is you have a sagacious and sympathetic guide, one who asks only for your patience and tolerance of occasional stilted prose (but would it kill you to purchase some swag from the LabKitty store in return for all of this free learning goodness? Answer: No, it would not). Still, if you wish to proceed, the gloves must come off. LabKitty degloved, we might say, which is an adjective I advise you do not Google.

It's like grandpap LabKitty used to say: Lace up your mukluks children, 'cause we're goin' to the slaughterhouse.

Let's press on.

Wednesday, April 12, 2017

A Primer on ANOVA

LabKitty Great Seal
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and mostly work only the simplest cases. I'll show mistakes and dead ends and things any normal person would try that don't work. Also, there's probably a swear word or two. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

Recently, I've been reading various and sundry explanations of ANOVA because reasons. These are bugging me and I can't quite put my finger on why. The people on the other end of the word processor all seem nice. They're competent enough, I guess. Good intentioned. The material is accurate. Nicely formatted. Pleasing fonts. Still, it's all so, I dunno, labor intensive. Too many equations. Too much effort.

The logic of ANOVA should just fall out of it. Effortlessly. Like a toddler out of Michelle Duggar. You shouldn't have to work for it, is my point. Instead, all I encounter are rats nests of variance partitions and degree of freedom Sudoku and something called R. Sure, you can memorize the formulas. That'll get you though your stats course. Probably. But wouldn't it be nice to Understand with a capital U? The logic of ANOVA should just fall out of it. Effortlessly. Like a toddler out of Michelle Duggar (that joke is bound to offend somebody, so I'm going to go ahead and double down on it up front. NO FEAR).

Let's see if we can't derive ANOVA using a minimal amount of thinking. Clearly, we are well on our way.

Thursday, December 10, 2015

A Primer on Fourier Analysis -- Part I

LabKitty Great Seal
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and mostly work only the simplest cases. I'll show mistakes and dead ends and things any normal person would try that don't work. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

The Fourier transform doesn't only transform data, it transformed civilization. In the guise of the FFT, it created a digital signal processing revolution that took the technique out of the classroom and into everything from GPS to Roombas. It is a staple of STEM education, and for good reason: There is almost no field, no application, no problem that does not yield to or benefit from application of Fourier analysis. If aliens have a Wiki on Earthlings, Fourier and those blankets you can wear may well be the sole entries.

That being said, Fourier analysis can be intimidating to the newcomer. The material has a fierce reputation, often taught as some holy relic impenetrable as death. That's where LabKitty comes in. Yes, there is much about Fourier analysis that only time and dedicated study can unravel. But the basic idea is surprisingly straightforward. And if there's anything textbook authors hate, it's admitting something is straightforward.

Wednesday, July 22, 2015

A Primer on Stochastic Differential Equations

LabKitty Academy Greal Seal
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and work only the simplest cases. I'll show mistakes and dead ends and things any normal person would try that don't work. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

The appearance of Margot Robbie in the recently leaked footage of Suicide Squad got me thinking about Margot Robbie in The Wolf of Wall Street, and that of course got me thinking about the Black Scholes financial model. (Aside: I don't understand what about Suicide Squad requires a 2016 release; the movie could be Harley Quinn reading the phone book for 90 minutes and nerds would flock to see it. But I digress.)

The Black Scholes model is an example of a stochastic differential equation (SDE). It's also an example of mathematics destroying the world, sorta like how an atom bomb is really just a Markov chain except this time we dropped it on ourselves. For Black Scholes is one of the cornerstones of the derivatives trading biz (the other three corners are cocaine, prostitutes, and suborned congressmen). Scholes won a Nobel prize for his work and the very next year his brokerage firm had to be bailed out for something like a $4 billion. If that doesn't make you hate math, I don't know what will.

Continuing my runaway train of thought, SDE are also an example of something I call Horace mathematics, after the Roman poet Horace who penned: Parturient montes, nascetur ridiculus mus (the mountains will be in labor and an absurd mouse will be born). SDE are an absurd mouse indeed. The gulf between their fearsome reputation and how simple they are to implement on a computer is wide enough to fit all of the planets in the solar system (which is approximately the distance from the Earth to the moon, by the way). Wider than the time between Cleopatra and the building of the pyramids (who lived closer to modern times than to that of the pharaohs). Wider than the gap between Tyrannosaurus and Stegosaurus (who lived closer to us than to Stegosaurus).

Seriously. I haven't been let down this hard since sea monkeys.

Wednesday, June 24, 2015

A Primer on Principal Component Analysis

LabKitty primer logo
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and mostly work only the simplest cases. I'll show mistakes and dead ends and things any normal person would try that don't work. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

Principal Component Analysis (PCA) is a black art of the statistical majicks having a reputation of equal parts usefulness and difficulty. It's a way of simplifying complex data, a way of extracting structure hiding in mess, sometimes even pointing you to an underlying physical process that explains what you measured. And it does so without any assumptions or restrictions, without even knowing what the data are or where they came from. When PCA works, it's a glorious thing to behold.

Yet, the method is cursed with an air of impenetrableness. It's a (relatively) new invention, made possible only with the rise of computers and computational statistics. PCA ain't a t-test or ANOVA -- you can't do it with a pocket calculator and a table at the back of the textbook. So it is we poor punters have been at the mercy of the mathematicians and their inscrutableness when it comes to building user-friendly programs, or explaining anything really.

Well, LabKitty is here to help. The what of PCA is fairly straightforward; it's the how that presents a problem, assuming you aren't fluent in singular value decomposition and eigenvalues and covariance matrices. I dare say PCA is like making a baby: easy to describe in pictures (see, e.g., the Internet) or in cold clinical detail (gametofusion, embryo implantation, growth and development, dilation and contraction, fetus expulsion), but inside the box, so to speak, is some rather unnerving grossocology. As such, I will describe PCA in pictures, and show you how to do it in code (it's just a couple of lines of Matlab) but not have much to say about what happens behind the curtain. If you're looking for mathematical details, you need to look elsewhere.

After having talked about the what and how of PCA, we'll then turn briefly to the why, although if you are reading this you probably already have an application in mind. I'll wrap up things with what can go wrong, and present a few leads into the literature which I have found helpful, although nothing that can't be bested with a tailored Google search.

So, then. I believe someone mentioned baby making. Let's get at it.

Wednesday, October 22, 2014

A Primer on Mathematical Epidemiology

LabKitty Great Seal
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and mostly work only the simplest cases. I show mistakes and dead ends and things any normal person would try that don't work. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

Note for the confused: This Primer makes reference to events from the recent (2015) Ebola outbreak, which featured some arguably questionable decisions on the part of various healthcare professionals. FYI.

Now that we're all going to die in a global apocalyptic Ebola apocalypse, it's a good time to learn some epidemiology. This will soon be all over the "for your health" segment of the news, so you best get yourself prepared. When survivors have devolved into samizdat armed camps, you'll need something to make yourself useful to the local pod of crazies. Demonstrating facility with this material could be your golden ticket to an MRE and a place in the repopulation efforts.

The bad news is it's differential equations. The good news is it's epidemiology differential equations, and if recent actions of the CDC are any indication, plenty of knuckleheads understand epidemiology. Or don't understand epidemiology, at least the "keep the infected from flying coach" portion of the program. Although, to be fair to the CDC, perhaps the larger problem here is common sense. Apparently we're no longer covering the germ theory of disease in our nursing schools.

As always, I blame the hip hop.

Wednesday, July 16, 2014

A Primer on Variational Calculus

labkitty seal
A LabKitty Primer is an introduction to some technical topic I am smitten with. They differ from most primers as to the degree of mathematical blasphemy on display. I gloss over details aplenty and mostly work only the simplest cases. I show the dead ends and things any normal person would try that don't work. Mostly I want to get you, too, excited about the topic, perhaps enough to get you motivated to learn it properly. And, just maybe, my little irreverent introduction might help you understand a more sober treatment when the time comes.

Variational calculus (VC) is the Cirque du Soleil of mathematics. It has a reputation of impenetrable weirdness. It can do things seemingly nothing else can. And even though you recognize the basic ingredients that go into it (juggling, tumbling, French), they are taken to such extremes you can't but help feel you don't know them at all.

The motivating idea is simple: find a function that maximizes or minimizes some expression. Often, that expression is an integral involving the function. The concept is not entirely alien to the student of calculus. You are taught how to find function extrema back in Calc-I: recall taking derivatives and setting them equal to zero. There the function was given and you sought special values of the independent variable(s). But finding a function that satisfies a given expression is also not new. I dare say that's what it means to solve a differential equation.

Variational calculus leverages the synergy of these two ideas. A mathematical peanut butter cup, as it were. And VC allows you to attack problems that can be solved in no other way. The type of problems that are at the heart of engineering and technology. Find the airfoil shape that minimizes drag. Find the transistor layout that maximizes chip speed. Find the protein conformation that maximizes drug uptake. Beyond the human realm, we find almost every fundamental law of nature expressed in the language of variational calculus. The motion of a particle minimizes the difference between its potential and kinetic energy. Soap bubbles assume the shape that has minimal surface area. Light propagates along the path that minimizes travel time. Variational calculus is like reading the mind of God.