Wednesday, 26 September 2018

6 page memos

Just have been reading an article about Jeff Bezos's approach to keeping a company as dynamic as a startup is or should be.

I'll take the liberty to (disagree and commit lol) almost just bullet-point list out the keywords ... just becasue this is how much time I have :) - the below only makes sense together with the above Forbes link really)

First, "senior executives start meetings at Amazon in silence, with everyone reading six-page narrative memos about the topic they are gathered to discuss, for up to 30 minutes"
Note that here quality preparations are brought into the game - when appropriate at least.

Then, it is important to make a distinction between decisions: "Type 1 decisions can't be reversed and as such require great care. Type 2 decisions can be easily reversed."

"Make Decisions With 70% Of The Info You Wish You Had"

"Disagree and commit"

"You have to somehow make high-quality, high-velocity decisions"

Well, that's it, build your own Amazon!

Saturday, 4 August 2018

Easter eggs for experts in MongoDB: beat the 8

^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:38Z")
> ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:38Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
> ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
> ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
> ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
> ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:39Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:40Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:40Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:40Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:40Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:40Z")
^[ObjectId().getTimestamp()
ISODate("2018-08-05T00:16:40Z")

Yes, that's eight of a kind...!

Monday, 25 June 2018

EBS init hell

I should remind myself that there is some initial penalty for creating an EBS volume on AWS - right after creation it is really slow to access for a good while.

It's a gp2 volume and iotop reports a 5 MB/s access speed, in agreement with the CloudWatch metrics.

I am wondering for how long - the stuff on this link suggests straight after first access it starts feeling fine.

However, I can see it being slow on the second run of the recommended
dd if=/dev/xvda of=/dev/null bs=1M

okay I only partially ran it at first, but I'd expect walking the blocks to be a very deterministic process for dd. So maybe it's rather the first complete access? Or a few hours of initialization, such as the case is a large enough chunks when growing a volume or e.g. with a complete drive type change (i.e. gp2-io1)?

Well, I give up on that for today/night but best remember this caveat ...

Update: really a little googling confirms this... that it needs a complete read through:
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ebs-initialize.html


Monday, 4 June 2018

Processing close: filter early

After all, it's just a bit of extension on the previous post, and is a generally utilizable intuitive augmentation concept in IT.

+1) Consider filtering/reducing the information or data in transport early


If need be for performance optimization, and there is some filtering involved at a late stage in data processing (e.g. with databases - a where clause), it's worth considering whether it improves the overall process if that filtering is done at an early stage.

Of course, the most radical filtering steps, i.e. those with the best selectivity are often the best to evaluate first, etc..

This can again be useful when designing map-reduce algorithms, or - more generally - data processing flows. Obvious factors to consider could be the network transportation costs, the temporary persistence in a file system, reducing on these can easily improve the overall performance of the system.

One explicit case: for symmetric graphs, where the symmetry is preserved in intermediate results, it can be a good thought to transfer only a half of it, e.g. in and above the diagonal, and only "double" the output in the very last step, if that is required at all.

Tuesday, 22 May 2018

Processing close

People sometimes write code that is very assertive about this or that piece of a system. In a more suspicious case, many parts.
This typically isn't going to yield performant code - something will often try to slow your job down.

A specific advise:

0) Keep (simple) processing close to the database.


What I mean by simple is something that easily could be translated to a few machine level instructions (e.g.: sum) in each iteration, is done in a few iterations (e.g. 1), doesn't operate on a large dataset in each iteration.

Let's generalize it:

1) Keep avoiding to cross boundaries between system parts unnecessarily.


Or: try to consider alternatives where fewer communication routes are involved.

Such means of system boundary crossing to be avoid could be anything really, depending on the scale of operation, and at what magnification we consider the system:
  • Querying data from a server
  • Calling into a DLL/library
  • Exchanging information with a service/daemon (inter-process - but also, probably with a smaller overhead, synchronized inter-thread communication).
  • Looking up data from another table (one implication: look for unnecessary JOINs, maybe consider denormalization, esp. in a NoSQL setting)
  • Calling Java bytecode from a compiled binary
  • Calling another function (remember there's a cost of leaving return information on top of the stack)
  • Blocking on-screen confirmation with the user

 

A major exception: if the overall complexity is high.


In this case it may be worth turning to some hard/software dedicated to the processing - using a GPU, an SSD, scaling up, a cluster etc. a few terms to think of as opposed to CPU/FPU, HDD, your regular hardware, single compute node. And this will all involve moving data from one part of the system, less suited for the processing task, to another, more feasible, possibly more dedicated, potentially a temporarily allocated resource from a cloud.

However, it's also worth noting:

 

faulty and or critical system components may cry for redundancy


Such as - people. And then there come the desirable boundaries - but also come peer insight. Whether the superior efficiency of the one-man teams is a myth... well, while we'd like to believe in myths, we do know, that may not always be the best that can happen. You may need very proficient and disciplined people for that to work out - these days, with an expanding IT industry, years of experience on average is going plummeting.
Stumbling upon them could be way more the exception than the norm.

Thursday, 17 May 2018

A quick reminder to self about error handling in R (baby steps #1)

I guess the below code tells pretty much it all:

tryCatch({
  x = function() {
    stop("hahaha now you blew the code!")
  }
  y = function() {
    x()
  }
  y()
}, error=function(e){
  print(e$call)
  print(e$message)
})


And then sourcing it gives:

x()
[1] "hahaha now you blew the code!"


I guess that's all that there is in practice...
... except maybe that there's a finally parameter included which I never notice :

function (expr, ..., finally) 

So... worth a second look :)

(To be continued...)

Monday, 7 May 2018

JavaScript versus integer sequence

Looks like the ever feared, popularly hated JS really has its drawbacks ... you need functional programming to create an integer sequence? you need a loop? and wouldn't anybody want to just do something about it? :)

E.g.
https://stackoverflow.com/questions/3895478/does-javascript-have-a-method-like-range-to-generate-a-range-within-the-supp

Jeez :) time to realize how good Python and R are... well, R as a language only in certain aspects, but still.