# Random Talks and Thoughts

**URL:** <https://discourse.pijul.org/t/random-talks-and-thoughts/29>\
**Category:** Development\
**Created:** [June 2, 2017, 4:27pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29 "2017-06-02T16:27:13Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![lthms](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/lthms/32/124_2.png) [@lthms](https://discourse.pijul.org/u/lthms)\
**Post date:** [June 2, 2017, 4:27pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/1 "2017-06-02T16:27:13Z")

</div>

If you want to share something with the pijul contributors, yet you don’ feel it deserves its own thread, here is the perfect place. The main idea is to exchange random thoughts or ideas in an informal way.

For instance: I have just pushed a patch to add a [`NotImplementedYet` error](https://nest.pijul.com/pijul_org/pijul/patches/Ae3ai9gtUMaxOtPv5IIBZ1Q3HufvPJQhk78iHci6Z6q-LP9obAkM5ZmFHGn5J_3mew1WfwfuW-YK38B7PvdJVzI) so that we can be lazy sometimes if the feature we are working on is too big d:.

---

<div class="post-metadata">

**Author:** ![lthms](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/lthms/32/124_2.png) [@lthms](https://discourse.pijul.org/u/lthms)\
**Post date:** [June 5, 2017, 1:59pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/2 "2017-06-05T13:59:34Z")

</div>

[I can’t create a proper PR in the Nest](https://nest.pijul.com/pijul_org/nest/issues/28) so this is a good occasion to use this thread. I’ve tried to tweak a little the output of `pijul -V` to give more information to the user. It says if pijul has been compiled as `debug` or `release`.

```auto
# release
λ pijul-dev -V
pijul 0.6.0
2017-06-05 12:28:33.040423785 +02:00
# debug
λ target/debug/pijul -V
pijul 0.6.0 -- development version
2017-06-05 12:27:07.925990949 +02:00

```

The patch is [on my repo](https://nest.pijul.com/lthms/pijul/patches/Acfe4SjAixp0FgfC5DMQa7ZLw1JWoc2Dw5N3Mu3rHfWHUhOMMbiFkiEYUrGhI7RDXqpVvTliMzBJSX0oHPh3Quw) and I’d like your feedback!

---

<div class="post-metadata">

**Author:** ![thedward](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/thedward/32/17_2.png) [@thedward](https://discourse.pijul.org/u/thedward)\
**Post date:** [June 8, 2017, 5:41pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/3 "2017-06-08T17:41:57Z")

</div>

`bincode` is a bad choice for serialization for two reasons — ⓐ it doesn’t produce self describing output, ⓑ it is too closely tied to Rust; ⓐ is not a huge deal, but ⓑ could seriously limit future adoption. I am not questioning your choice of implementation language, in fact based on what I know about Rust it sounds like a great fit. However, once Pijul is more mature it would help tremendously with adoption if it was practical to read and write Pijul repositories with tools written in other languages (without having to depend on parsing executable output or interfacing to a foreign library). The alternate implementations of git for example have made a lot of interesting software possible.

I want to acknowledge that I am late to the game here — I just now read the April 2nd blog post. Furthermore, I want to acknowledge that my conclusions may be misinformed¹ or simply wrong. I just wanted to make sure these concerns were taken into account, if only to be dismissed.

At the very least, it seems like a (machine readable) specification of the binary format² would be a huge win for testing and maintaining compatibility across versions? Does `bincode` include support for versioning your data?

¹ I didn’t dig deep into `bincode`, so my understanding of it is somewhat superficial. Neither am I competent with Rust, so maybe the data structures are so well defined that this is not an issue.

² Even if _post hoc_

p.s. None of this is going to prevent me from trying it out; In fact, I’m in the process of getting Rust up and going on my system just for that purpose.

---

<div class="post-metadata">

**Author:** ![pmeunier](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/pmeunier/32/4_2.png) [@pmeunier](https://discourse.pijul.org/u/pmeunier)\
**Post date:** [June 9, 2017, 6:09am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/4 "2017-06-09T06:09:19Z")

</div>

Thanks for the comments. In the beginning, we were using cbor, but the lack of unique representation forced us to change: different implementations of cbor in Rust were using different features.

Bincode _is_ a great fit for this: it is extremely simple, super easy to define, and has unique encodings. It is not self-described, but you probably don’t really care, since Pijul patches store binary information, which you would probably not make sense of by yourself.

Bincode is also extremely fast. Being tied to Rust is also not a big issue, since you’ll probably want to interact with Pijul via libpijul, itself written in Rust.

---

<div class="post-metadata">

**Author:** ![XWinkel](https://avatars.discourse-cdn.com/v4/letter/x/5daacb/32.png) [@XWinkel](https://discourse.pijul.org/u/XWinkel)\
**Post date:** [August 10, 2017, 8:19pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/5 "2017-08-10T20:19:22Z")

</div>

I very much like the idea of patch based version control; I am not a dumb person but find it hard to imagine what git is doing, when I enter git commands such as found from the web (non-intuitive commands). My question here, is even more basic as the difference between patch based and snapshot based version control : Suppose one has a source code file that contains a loop and a counter in that loop has been forgotten to be incremented, and it does not really matter where in the loop body that counter gets incremented when that gets corrected, as long as, say it is the second half of that loop body (for sake of this discussion). Now suppose person A develops a patch AA in which counter gets incremented right in middle of loop body and person B develops a patch BB in which counter gets incremented at end of loop body. Now _as there is no conflict of lines_ it appears as if both patches can be merged without conflict. However, then, the counter gets incremented twice. This is a fundamental problem of automatic merging.

---

<div class="post-metadata">

**Author:** ![pmeunier](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/pmeunier/32/4_2.png) [@pmeunier](https://discourse.pijul.org/u/pmeunier)\
**Post date:** [August 10, 2017, 8:44pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/6 "2017-08-10T20:44:40Z")

</div>

Sure, but even if you managed to get a formal definition of the problem you would like to solve, that problem is likely to be undecidable. Instead, what Pijul does is to guarantee at least a few axioms to rely on: associativity, commutativity, inverses.

We’re not claiming much more than that, but it’s already much better than others (including Git).

---

<div class="post-metadata">

**Author:** ![XWinkel](https://avatars.discourse-cdn.com/v4/letter/x/5daacb/32.png) [@XWinkel](https://discourse.pijul.org/u/XWinkel)\
**Post date:** [August 10, 2017, 9:11pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/7 "2017-08-10T21:11:08Z")

</div>

Another “random idea” concerns the definition of a “patch”. Such a definition, in order to pin down the location of an edit, can use either line numbers, or context (lines before/after the edit, sufficient number of them such that they are unique in the file). However, myself I think it would help enormously if each text file line would have a (hidden) unique ID number go with it (that a text editor would hide). In that case, insertion or deletion of text at other locations (through other patches) that would affect line numbers, would leave such an ID untouched. During copy paste actions, the editor would have to insert new ID numbers for the pasted lines, to keep the IDs unique at all times. The downside of this is that it needs a slight extension of the text editor people would use to a) hide these IDs b) upon paste actions replace existing IDs by new IDs. However, it would seem to me such unique ID’s per line would enable a more powerful definition of the concept of a “patch”. [https://stackoverflow.com/questions/45258674/would-it-help-source-code-version-control-systems-to-start-each-line-with-a-64-b](https://stackoverflow.com/questions/45258674/would-it-help-source-code-version-control-systems-to-start-each-line-with-a-64-b)

---

<div class="post-metadata">

**Author:** ![pmeunier](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/pmeunier/32/4_2.png) [@pmeunier](https://discourse.pijul.org/u/pmeunier)\
**Post date:** [August 10, 2017, 9:26pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/8 "2017-08-10T21:26:49Z")

</div>

We have something like that in Pijul, but we also have a fairly robust theory proving what patches can and cannot do if we want to keep some guarantees. Moving text around while keeping line ids is sometimes possible, sometimes not.  
In particular, it needs to work when several persons do it in parallel.

---

<div class="post-metadata">

**Author:** ![pointfree](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/pointfree/32/13_2.png) [@pointfree](https://discourse.pijul.org/u/pointfree)\
**Post date:** [August 12, 2017, 7:41pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/9 "2017-08-12T19:41:49Z")

</div>

With Victor Grishchenko’s [ctre](https://github.com/gritzko/ctre) every character has identity spanning versions. This makes following changes to a text over time much easier because the identity of a string is not bound to whatever position it happens to be at in the current state of the repo.

[https://github.com/gritzko/ctre/blob/master/doc/ws10.pdf](https://github.com/gritzko/ctre/blob/master/doc/ws10.pdf)

---

<div class="post-metadata">

**Author:** ![lthms](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/lthms/32/124_2.png) [@lthms](https://discourse.pijul.org/u/lthms)\
**Post date:** [August 13, 2017, 5:54pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/10 "2017-08-13T17:54:23Z")

</div>

Eh all, I have wrote and submit a little patch for pijul in order to give a little more information with `pijul -V`. Would love to get some feedback!

[https://nest.pijul.com/pijul\_org/pijul/discussions/140](https://nest.pijul.com/pijul_org/pijul/discussions/140)

---

<div class="post-metadata">

**Author:** ![felix91gr](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/felix91gr/32/87_2.png) [@felix91gr](https://discourse.pijul.org/u/felix91gr)\
**Post date:** [April 16, 2018, 6:26am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/11 "2018-04-16T06:26:35Z")

</div>

> [@XWinkel](#):
>
> However, then, the counter gets incremented twice.

This is why we use tests, I guess. CI would maybe detect this error.

On the other side of things, if you were using a dependently typed programming language, you could state the property “this counter gets incremented only once per step”, and something like an extension to Pijul that could query your compiler could say: “Hey, I know this merge looks fine, but it actually breaks this property”. But then again, if you were using such a language, you could introduce other kinds of bugs of a kind inexpressible in the language itself. So I don’t really know.

Tests for the win! Fight it with tests 🙂

---

<div class="post-metadata">

**Author:** ![lthms](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/lthms/32/124_2.png) [@lthms](https://discourse.pijul.org/u/lthms)\
**Post date:** [April 23, 2018, 6:12am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/12 "2018-04-23T06:12:34Z")

</div>

I’ve run into this very interesting reddit comment: [https://www.reddit.com/r/rust/comments/8dx08d/new\_release\_of\_pijul\_010\_more\_stable\_than\_ever/dxsuabt](https://www.reddit.com/r/rust/comments/8dx08d/new_release_of_pijul_010_more_stable_than_ever/dxsuabt)

It looks like interesting benchmarks.

---

<div class="post-metadata">

**Author:** ![yory](https://avatars.discourse-cdn.com/v4/letter/y/da6949/32.png) [@yory](https://discourse.pijul.org/u/yory)\
**Post date:** [April 23, 2018, 6:55am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/13 "2018-04-23T06:55:20Z")

</div>

this is exactly my issue, I reported it on this forum a while back (title “bad performance”). If you read through it, you’ll find two python tests that deal with the issue. They show:

1. the .pijul folder becomes quickly ENORMOUS, while .git’s remains very compact
2. pijul quickly becomes VERY slow, until to a point where it freezes the computer by sucking away all ram it can find.

Pmeunier told he expects the issue to strongly mitigate when we move to myers diff

---

<div class="post-metadata">

**Author:** ![lthms](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/lthms/32/124_2.png) [@lthms](https://discourse.pijul.org/u/lthms)\
**Post date:** [April 23, 2018, 7:29am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/14 "2018-04-23T07:29:39Z")

</div>

I wasn’t aware of this discussion, thanks for pointing me out!

I wonder if changing for a better diff algorithm will tackle both 1. and 2., or only 2.

---

<div class="post-metadata">

**Author:** ![yory](https://avatars.discourse-cdn.com/v4/letter/y/da6949/32.png) [@yory](https://discourse.pijul.org/u/yory)\
**Post date:** [April 23, 2018, 7:36am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/15 "2018-04-23T07:36:37Z")

</div>

I don’t know either, it’s a bit too low level for my skill.

---

<div class="post-metadata">

**Author:** ![pmeunier](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/pmeunier/32/4_2.png) [@pmeunier](https://discourse.pijul.org/u/pmeunier)\
**Post date:** [April 24, 2018, 5:46pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/16 "2018-04-24T17:46:49Z")

</div>

I can obviously answer only from a theoretical point of view, since I’ve not implemented the fix to diff. I would definitely expect Pijul to take up more disk space than Git, although I don’t know in which proportions. One issue is that there is some redundancy in the storage format, which we could get rid of in the future. Some work on the backend could also help. I chose to use copy-on-write B trees with skiplists in the nodes, which I believe to be optimal in terms of speed. But maybe there’s a better trade-off to be found!

Another problem is that line deletions tend to rewrite the database a lot, because of a naive implementation of something in the apply algorithm.

Since the only benchmark I have about Sanakirja is in its own tests, I cannot tell for sure that Sanakirja doesn’t leak, especially when it is used in Pijul (Pijul has some unsafe code). I should definitely try to write some more benchmarks in Pijul, by exposing the leak detection functions in the Sanakirja API.

---

<div class="post-metadata">

**Author:** ![lthms](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/lthms/32/124_2.png) [@lthms](https://discourse.pijul.org/u/lthms)\
**Post date:** [September 5, 2018, 6:52am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/17 "2018-09-05T06:52:22Z")

</div>

@flobec, could you give me moderation privilege for the `#pijul` channel on IRC?

The poor thing is actually flooded by regular spammer, and we need to do something about it if we want to regain some proper discussion medium. Also, it gives a poor image of the project, for those who want to talk with us and only find… strange messages, to say the least.

---

<div class="post-metadata">

**Author:** ![flobec](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/flobec/32/23_2.png) [@flobec](https://discourse.pijul.org/u/flobec)\
**Post date:** [September 5, 2018, 12:01pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/18 "2018-09-05T12:01:36Z")

</div>

I think you have the privilege now, but you should check.

---

<div class="post-metadata">

**Author:** ![lthms](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/lthms/32/124_2.png) [@lthms](https://discourse.pijul.org/u/lthms)\
**Post date:** [September 5, 2018, 12:45pm UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/19 "2018-09-05T12:45:59Z")

</div>

It worked, thanks! I set the channel mode to `+r`, as it looks like a recommended mode to deal with spam.

---

<div class="post-metadata">

**Author:** ![no\_identd](https://yyz1.discourse-cdn.com/flex031/user_avatar/discourse.pijul.org/no_identd/32/126_2.png) [@no\_identd](https://discourse.pijul.org/u/no_identd)\
**Post date:** [December 24, 2019, 10:34am UTC](https://discourse.pijul.org/t/random-talks-and-thoughts/29/20 "2019-12-24T10:34:01Z")

</div>

@pmeunier @lthms: You might wish to consider switching the serialization format to bencode, and more specifically, bencode using the following **EXCEEDINGLY** excellent Rust library, which runs circles around pretty much all other implementations, for reasons you can read at the link below:

> <https://github.com/P3KI/bendy/blob/master/README.md#why-should-i-use-it>
