Saturday, 24 December 2016

Tools are not (just) tools

Note to self

I asked someone today whether he still likes maths, and got the response that maths is just a tool.

It made me think whether tools are actually just tools, or slightly more than that.

Accidentally I was about to study innovation management, which reminded me of adoption curves.

The Rogers adoption curve starts with the innovators. These people are the first to get acquainted with an innovation, use a new gadget, try out a new service, etc. Their characteristics include trying out things l'art pour l'art, which means they don't care about any practical value of those attempts.
Just toy with them.

I had two things to note:

1. these people are those who exercise themselves with things without a particular purpose

2. this is very much an analogue of how ADD people (in my imagination) get distracted by literally anything

About 1.


Actually what purpose do we have in mind when we are kids playing with things? Some made up one I think normally drives things, but this can be so minimally targeted as 'joy', 'fun', etc.
Sounds like a good idea to me.

And just because they have the "thing" around, they can play around with it. The presence of the item allows for another degree of freedom.
And people, like a gas, fill the space they are given.

So the new toy gets played with, new routine gets added to what was there before.

Tools as toys, become doors, and so is maths a door, doors get entered, and so will maths likely become part of your future path, once you've played with it.

About 2.


This is just a corollary - I would guess being overloaded with options just as well as overloaded with topics on the internet created its new addict group. Innovators are the ADD of the market.
Positively mad people. See 1. - allowing themselves options, they allow for creativity, by heading against discipline.

Now is a good time to ask - is discipline good or bad then?
Obviously, it depends.

Discipline and minimalism are just two sides of the same coin.
Capitalism and competition made us put things perhaps too much on the minimalist, the efficiency side.

You'll choose the offer for 999 coins but not for 1000 unless the cheaper is noticeably worse.
It's only the resolution of your value perception that needs to be tricked and the quality rot begins, a worse product sold for almost as high a price as the better one.

Discipline is a tool, not an objective. And an option, a freedom to choose. A tool, but once not a must, once you played with it without a need, more than a tool, too.

You can design bottom up as well as top down, and so you can create something without knowing the final design. Just because you didn't slap yourself in the face to get back in line before you'd have found another path.

We shouldn't forget to embrace freedom. And ADD makes people less controllable. A good thing, in some times.

Thursday, 15 December 2016

Is education going out of fashion?

Apparently people have finally found their ways to the schools ... or what's going on? They don't seem to search for institutions that much anymore.



It could be interesting to find out what alternative terms people have started to search for.

MOOC's are a very natural first thought. And they have clearly gained interest over time...



Are they all that significant?



Indeed.

So perhaps the above does not mean education is getting out of fashion, maybe as expected, it is actually getting to play a more and more significant role - but who knows.

The interest in academic degrees seems to nearly stagnate, although there is a likely shift towards shorter degrees:

Sunday, 16 October 2016

Idle priority: data project contributions

I will mention two here, so that I won't forget them. One is new, that sounds more interesting (of course, new things, always ...)

Amnesty Decoders (update: false trail...)

"Join a global network of digital volunteers helping us research and expose human rights violations."


This is actually something that deceived me big time:


They want you! To click on sections of satellite imagery where villages (artificial structures) are present. Hm... I would have thought people (with some proficiency) can contribute to the image analysis with code :(
I wonder why these projects don't end up on Kaggle. Or do they?
This task is absolutely crying out for automatisation ...



... and another one, which I actually started doing something with, from the past:

Open Corporates



It is a large database, commercially available for companies - that's how they make their living. It is backed by paid coders as well as enthusiasts, their meetup if someone wants a bit of hacking for good (they do/did provide free means of accessing the data as well, especially for contribution) is at:


I never found the time again, so my bot is still 'booting', but would have been nice to. Should.

The task is to create bots, in Python, which do the web scraping (in my days, it happened with BeautifulSoup), and crack up the data (through multiple steps) to extract relevant attributes.

Their stuff is helpful for those doing data journalism, and there are hopes it can help to cut back and/or eliminate corruption hidden behind the ownership graphs.

Friday, 7 October 2016

Why or why not Go [draft]

Gnack, language is a matter of choice, but the choice is a matter of the market, etc.; so this year as usual:
-- Away from me, Pascal!

So what language?


Starts to be apparent that some companies just do choose newer technologies, despite the safe player masses, although it's not always obvious when looking at the overall trend. Workforce performance (I mean code * quality / dev_hours or something similar) increase? I wonder. Secrets, never told.

One example: I've been trying to find the reasons pro and con for learning Go programming, as I'm not a massive C++ fan, but I know I like Python and, after a quick assessment, it didn't seem to be adding a lot over it. Still I doubted my judgement on this (being a master of neither of the two), so I looked at Google Trends and didn't find anything promising. Even Google didn't market it too well :)



However, just today I got reminded by a lecture of ItJobsWatch's charts, and see these UK-wide statistics! (Numbers at the front are approximately correct as of 8/10/2016)

1%: Go on ITJobsWatch

That's a small outbreak :) Mind that, my earlier investments, R and Python have gone way way up, too.

1.5% R on ITJobsWatch

14%: Python on ITJobsWatch

This is how it works! Or rather when it works.

Then add that Docker is written in Go, that CloudFoundry and RabbitMQ and that things people write for people to use, have started to recognize the power of this language, the result starts to get much nicer.

UPDATE:
On Quora they found that the Golang expression statistics show a nice, steep upward Google trend. I'd add that it's never guaranteed that it's a sign of success, while people realize that this expression works, the apparent explosion is possibly just a manifestation of the overtaking of term search numbers from other, related expressions (sort of a cannibalisation). "Golang" to date is backed by fraction of the searches for "Go (programming language)" and "Google go".

The Bug #1: Windows not that much love Go


However, there's this little bug, which still prevents (at the time writing) building Windows DLL's...

So then it doesn't seem like a "platform-neutral" development attempt, but on the other hand, something that has a brand starting "accidentally" like that of Google, forever, and not supporting full scale Windows development. Hm :) Me? Not suspecting a thing. Good question: who knows when?

TODO: Would be nice to have a time vs. number of comments/total length of comments :) when the hell is it going to get closed?

Update: The Bug #2 - unfriendly again, Python would love GoLang if...


So as "Bug #1" says, you'll have to think before building Windows DLL's with Go. Although from the epic talk on the bug page it seems it's only affecting multi-threaded code via the Windows TLS support. However, something like that should work, I guess ... especially if you use Go which prides itself of its multi-threading support.

So (or without noticing why) on SO they compromise on building .so's for Linux.
http://stackoverflow.com/questions/12443203/writing-a-python-extension-in-go-golang

But who wants to really create a *nix-only Python package, or one that may not be extensible at some point - further from internal use? (Think of Anaconda - I guess it's a blocker for more official python distros.)

Then there's gopy also mentioned, for making importing trivial. However, it's still not compatible with go >= 1.6. Even if it's a very good start to creep in to commonplace use as an extension language first.

And Linux is still not even nearly everything.

So seems like there's a little longer while to wait for the ecosystem to get ready for broader market penetration. Exciting moments anyway.

Surely more promising already on the server side!

TODO: Mention fun language - intrinsic motivation - creativity association, weakness in analytics.
TODO: Mention recently created Go on GitHub charts? Create new ones?

Thursday, 6 October 2016

An entropy paradox

Entropy/diversity is personally one of my favourite brain-wasting topics (I have my reasons for that, good ones :) ). Here's a short line of thoughts that illustrates the tricky nature of these notions. At the end, I'll probably not connect this back to naive everyday thinking which everyone is taking for granted, as it would be either embarrassing for many or worse even (in case I'm wrong), very embarrassing, but only for me alone :) But beware I think I could (make a fool out of myself)...

So the opening thought is that when we are young, we are more similar to both of our parents, and as we turn older, we will exhibit stuff related to the matching gender parent, i.e. specialize.

Now let's describe this by something that behaves like the Herfindahl index (any entropy measure does the job in one way or another) over a pair of similarity metrics. We'll find that from a pair: (mom_similarity, dad_similarity) closer to (0.5, 0.5) the individual's stats move closer to either (1, 0.0) or (0, 1.0), and that this means this index will increase, suggesting a less diverse individual.

Now let's take a look at this on the macro level, multiply up the aforementioned individual so that it becomes a population (of roughly indetically aged, random gender people which is growing up)!

From a series similar to [(0.5, 0.5), ..., (0.5, 0.5)] we observe a transition towards (assuming equal probabilities for the genders) a series that is more like [(1.0, 0.0), (0.0, 1.0), ...] etc.

What we then find is that the diversity on the macro level did exactly the opposite - a population of randomly grown ups is more diverse than that of babies. Actually that was quite trivial without the numbers already, but the joy and the words with the weird spelling ... :)

So yes, almost paradoxically, micro and macro level entropy may work against each other - the level of abstraction does matter a lot!



P.S.: Don't calm down. I look forward to distribute similarly useless thoughts in the future.

Tuesday, 28 June 2016

R: a quick permutation sightseeing

Listing permutations of a number of distinct elements is surely doable, as long as the number of elements is not too big. But with only a little pickiness, such an innocent task can raise many interesting (but small) challenges to deal with once everyday expectations (memory, heap, CPU efficiency) are respected.

I found myself facing this task which involved a list of permutations of 5 elements (related to the Draper Satellite Image Chronology competition on Kaggle), and while it's truly simple, I thought I would just have a peek at how far the common sense has progressed with it.

Answers to a StackOverflow question should tell something about it.

There were
  1. the brute force exhaustive search approach, probably the simplest of all (which is probably the most suited for my original motivation)
  2. the classic recursive implementation
  3. Museful's much more vectorized ("matrixized") solution, a variant of which had too crossed my mind (unfortunately the SO one is just as good if not better than mine would have been, so I think I shouldn't waste many words on it)
and some others which seemed conceptually equivalent to either of the above.

The problem was pretty much solved, but my brain kept ticking... so there is one more I would like to take a note of here.

Firstly, I did deal with this before, a very long time ago, and I had a solution which I annoyingly forgot about, but it was easy to recreate something similar.
That key concept was that (over the same elements) a permutation is only different from another one by (a number of) exchanges between pairs of elements.

Actually, that's almost too easy to see, as just by always exchanging a misplaced element with the one with the right number, and repeating it, in no more than n exchanges (actually n-1 is sharper, and probably the average number of minimum required steps is much lower) we should get there.

E.g. have 12345 want 54321

exchange 1 & 5, get:

52341

exchange 2 & 4, get:

54321

done.

In other words, just by exchanging elements, we can go from one permutation to the other.

So this is probably an option to list all the permutations just as well.

Going from 12345 to 12354 is a simple move: replace the last two. To get to the next meaningful step, a third element surely needs to be involved. Then we may get back to varying the first two. Then the elements in the most frequently varying pair need some fresh blood again.

To see what I mean, the pair variations over a given 'symbol set' may be:
  1. 12345 and 12354
  2. 12435 and 12453
  3. 12534 and 12543
Giving the first 6 permutations. The generalisable recipe of items for the above starters:
  1. don't do anything,
  2. replace the 3rd with one element from the pair,
  3. replace the 3rd with the other element from the pair.
The flow goes as:

This may still feel a bit unclear, but technically from the permutations of (n - 1) elements we get those of n elements by placing the new element into different positions within all n slots, which means in the first case it's not involved in the last (n - 1), then it substitutes a different one of those (n - 1) in each "big" step. The last (n - 1) elements are just shuffled around as usual.

Slightly modifying it (reverse order) for the ease of coding, we get the below implementation.

I was really happy to sketch this out, although it's not much of an improvement. If I really want to be nice to myself then it is, as it does not have to be recursive, surely has a light heap footprint (the momentarily state of the algo is stored in 2 * n integers), but otherwise isn't too interesting, at least in R, where vectorization over tons of short loops is much preferred. It could possibly be a good choice in C(++) if someone's bored enough.

Then I Googled permutations with pair switches and got to a really interesting PDF.
Their solution assigns a 'momentum' (direction) to each number at each step, and I don't fully understand it (yet). Powers beyond my control :)
But surely is a simple one and is more efficient in that this one doesn't put back the items from where they came from, i.e. only applies half of the switches. However, it does have to find the largest movable element.

So do I win or what? :) Maybe only as long as I resist the temptation to read beyond the first page of their stuff... I'd say, for the temporary illusion, not today! Sure they have something better.. :)