• 0 Posts
  • 7 Comments
Joined 3 years ago
cake
Cake day: September 2nd, 2023

help-circle

  • If you read the conclusion, it is very clear that you should not ever break convention. You should not apply this if it would affect any other program.

    That is, if your input is in u8, your output should be in u8 too. Otherwise you are “exposing” your different quantization.

    I don’t think this post is telling you that you should use the 256 method. I think it is pretty clear that it is just explaining it, not advocating for it. At various points of the article it is written something like “I struggle to see a situation where it would be beneficial”.




  • The random number generation is not the source of the complaint. That is just made to easily prove the point in the article.

    The argument is not to use another method when randomly sampling from [0,1). The argument is that f32 color channels are already in the [0,1) range, and following the alternate method results in a “fair-er” representation of the range.

    The plot that was generated with the random sampling is just to show that the standard method is not “fair”.


  • The correct way to scale a number from one range to another is …

    Citation needed.

    It is just as valid to quantize the following way:

    q = (x+0.5)/256

    The inverse would be:

    x = min(truncate((q*256)), 255)

    In fact, you could argue that this method is technically more correct. The only issue being the exact value of 1.0, which we have to handle explicitly with min(..., 255).

    This method is idempotent. Try it with x=90 for example. It does destroy data, but that is true for all quantization methods.

    We have to remember that quantization is the process of representing a bigger set of numbers (usually infinite) with a smaller one. We sort a range of values from the bigger set into a single bin in the smaller set.

    The argument of the article is not about “resolution”. The argument is about using the 8bits available as efficiently as possible.

    As the article points out, if you follow the naive approach of:

    q = x/255

    x = round(q * 255)

    You are losing a tiny bit of the 8bit range. Since the ranges represented by 0 and 255 are half as big as the other ones. The naive approach is really quantizing in the range of [0-0.5/255, 1+0.5/255) instead of [0,1). This is because round() maps approaching values to x from both sides, so some negative values map to 0, but there are no negative values in our set.

    The conclusion at the end of the article is mostly the correct one though. If there is an industry standard, it is best to follow that standard. Since mixing a quantization function from one method with the dequantization function from another is just wrong. You can only choose your own method if both your input and your output is the small range.

    EDIT:

    The “naive” approach described in the article does x = truncate(q * 255 + 0.5) which is equivalent to the x = round(q*255)